Will it run? / Qwen3.8 flash next 125B (A6B) / Mac Studio M5 Ultra (256GB)

Can the Mac Studio M5 Ultra (256GB) run Qwen3.8 flash next 125B (A6B)?

Fast Yes, at about 84 tokens/sec (estimated). Replies stream faster than you can read.

84estimated tokens/sec
121.5 GBmemory needed
192 GBavailable to the model
1228.8 GB/smemory bandwidth
ollama run qwen3.8-flash-next:125b-a6b

Check the Mac Studio M5 Ultra (256GB) price →

How we got this number

Each generated token reads the active part of the model from memory once: 5.8 GB. The Mac Studio M5 Ultra (256GB) moves 1228.8 GB/s, so the ceiling is 120+ tokens/sec. Real machines reach a fraction of that ceiling; we use 80%, calibrated against a measured Mac run. Send us your real numbers and we'll replace the estimate.

Faster machines for Qwen3.8 flash next 125B (A6B)

NVIDIA DGX Spark (128GB)9.8 tok/sCheck price →
Ryzen AI Max+ 395 mini PC (128GB)9.2 tok/sCheck price →

Other models on the Mac Studio M5 Ultra (256GB)

Nemotron 3.5 lightning 30B (A3B)Fast120+ tok/s
gpt-oss 120BFast120+ tok/s
mixtral 8X7BFast82 tok/s
Qwen3.5 122B (A10B)Fast81 tok/s
Nemotron 3 super 120B (A12B)Fast63 tok/s
Nemotron 3 33BFast35 tok/s