This is a mixture-of-experts model: it only reads part of itself for each token, so it runs much faster than its size suggests, but the whole file still has to fit in memory. Around 10 tokens/sec reads comfortably; under 5 feels slow. How we estimate.
Embed this table
Free for blogs, READMEs and model cards. Updates automatically.
<iframe src="https://vsmacs.com/embed/nemotron-3-super-120b-a12b" width="100%" height="420" style="border:0" loading="lazy" title="Nemotron 3 super 120B (A12B) speed by machine"></iframe>