Will it run? / Qwen3.8 flash next 125B (A6B) / NVIDIA DGX Spark (128GB)

Can the NVIDIA DGX Spark (128GB) run Qwen3.8 flash next 125B (A6B)?

Usable, but slow Yes, at about 9.8 tokens/sec (estimated). It works, but you'll wait on longer answers. It only fits if you raise the GPU memory limit, and it'll be slower than usual.

9.8estimated tokens/sec
121.5 GBmemory needed
108.8 GBavailable to the model
273 GB/smemory bandwidth
ollama run qwen3.8-flash-next:125b-a6b

Check the NVIDIA DGX Spark (128GB) price →

How we got this number

Each generated token reads the active part of the model from memory once: 5.8 GB. The NVIDIA DGX Spark (128GB) moves 273 GB/s, so the ceiling is 47 tokens/sec. Real machines reach a fraction of that ceiling; we use 70%, an assumption until we get measured runs for this kind of machine. Send us your real numbers and we'll replace the estimate.

Faster machines for Qwen3.8 flash next 125B (A6B)

Mac Studio M5 Ultra (256GB)84 tok/s$9,499Check price →
Ryzen AI Max+ 395 mini PC (128GB)9.2 tok/sCheck price →

Other models on the NVIDIA DGX Spark (128GB)

Nemotron 3.5 lightning 30B (A3B)Fast43 tok/s
gpt-oss 120BFast33 tok/s
mixtral 8X7BComfortable16 tok/s
Qwen3.5 122B (A10B)Comfortable16 tok/s
Nemotron 3 super 120B (A12B)Comfortable12 tok/s
Nemotron 3 33BUsable, but slow6.8 tok/s