Will it run? / Qwen3.8 flash next 125B (A6B)

Can I run Qwen3.8 flash next 125B (A6B) on a Mac or mini PC?

Qwen3.8 flash next 125B (A6B) needs about 121.5 GB of memory. It fits on 3 of the 15 machines we track. The cheapest current machine that runs it at a comfortable 10+ tokens/sec is the Mac Studio M5 Ultra (256GB) ($9,499).

120 GBdownload (Q4_K_M)
~121.5 GBmemory needed
5.8 GBread per token (MoE)
3/15machines it fits on
ollama run qwen3.8-flash-next:125b-a6b
MachineMemoryFits?Tokens/secFrom
Mac Studio M5 Ultra (256GB) 256 GBYes84
$9,499Check price →
NVIDIA DGX Spark (128GB) 128 GBBarely9.8
Check price →
Ryzen AI Max+ 395 mini PC (128GB) 128 GBBarely9.2
Check price →
Mac mini M6 (16GB) 16 GBNo–
$899Check price →
Mac mini M6 (32GB) 32 GBNo–
Check price →
Mac mini M5 Pro (24GB) 24 GBNo–
$1,699Check price →
Mac mini M5 Pro (64GB) 64 GBNo–
Check price →
Mac Studio M5 Max (40-core GPU, 64GB) 64 GBNo–
$3,799Check price →
Mac Studio M5 Ultra (96GB) 96 GBNo–
$5,499Check price →
Mac mini M4 (16GB) older 16 GBNo–
Check price →
Mac mini M4 Pro (64GB) older 64 GBNo–
Check price →
Mac mini M2 (16GB) older 16 GBNo–
Check price →
GEEKOM A6 (Ryzen 7 6800H, 32GB) 32 GBNo–
Check price →
Beelink SER7 (Ryzen 7 7840HS, 32GB) 32 GBNo–
Check price →
Minisforum UM790 Pro (Ryzen 9 7940HS, 32GB) 32 GBNo–
Check price →

This is a mixture-of-experts model: it only reads part of itself for each token, so it runs much faster than its size suggests, but the whole file still has to fit in memory. Around 10 tokens/sec reads comfortably; under 5 feels slow. How we estimate.

Embed this table

Free for blogs, READMEs and model cards. Updates automatically.

<iframe src="https://vsmacs.com/embed/qwen3-8-flash-next-125b-a6b" width="100%" height="420" style="border:0" loading="lazy" title="Qwen3.8 flash next 125B (A6B) speed by machine"></iframe>