This is a mixture-of-experts model: it only reads part of itself for each token, so it runs much faster than its size suggests, but the whole file still has to fit in memory. Around 10 tokens/sec reads comfortably; under 5 feels slow. How we estimate.
Embed this table
Free for blogs, READMEs and model cards. Updates automatically.
<iframe src="https://vsmacs.com/embed/qwen3-8-flash-next-125b-a6b" width="100%" height="420" style="border:0" loading="lazy" title="Qwen3.8 flash next 125B (A6B) speed by machine"></iframe>