Will it run? / Nemotron 3 super 120B (A12B)

Can I run Nemotron 3 super 120B (A12B) on a Mac or mini PC?

Nemotron 3 super 120B (A12B) needs about 88.5 GB of memory. It fits on 4 of the 15 machines we track. The cheapest current machine that runs it at a comfortable 10+ tokens/sec is the Mac Studio M5 Ultra (96GB) ($5,499).

87 GBdownload (Q4_K_M)
~88.5 GBmemory needed
8.7 GBread per token (MoE)
4/15machines it fits on
ollama run nemotron-3-super:120b-a12b
MachineMemoryFits?Tokens/secFrom
Mac Studio M5 Ultra (256GB) 256 GBYes63
$9,499Check price →
Mac Studio M5 Ultra (96GB) 96 GBBarely38
$5,499Check price →
NVIDIA DGX Spark (128GB) 128 GBYes12
Check price →
Ryzen AI Max+ 395 mini PC (128GB) 128 GBYes11
Check price →
Mac mini M6 (16GB) 16 GBNo–
$899Check price →
Mac mini M6 (32GB) 32 GBNo–
Check price →
Mac mini M5 Pro (24GB) 24 GBNo–
$1,699Check price →
Mac mini M5 Pro (64GB) 64 GBNo–
Check price →
Mac Studio M5 Max (40-core GPU, 64GB) 64 GBNo–
$3,799Check price →
Mac mini M4 (16GB) older 16 GBNo–
Check price →
Mac mini M4 Pro (64GB) older 64 GBNo–
Check price →
Mac mini M2 (16GB) older 16 GBNo–
Check price →
GEEKOM A6 (Ryzen 7 6800H, 32GB) 32 GBNo–
Check price →
Beelink SER7 (Ryzen 7 7840HS, 32GB) 32 GBNo–
Check price →
Minisforum UM790 Pro (Ryzen 9 7940HS, 32GB) 32 GBNo–
Check price →

This is a mixture-of-experts model: it only reads part of itself for each token, so it runs much faster than its size suggests, but the whole file still has to fit in memory. Around 10 tokens/sec reads comfortably; under 5 feels slow. How we estimate.

Embed this table

Free for blogs, READMEs and model cards. Updates automatically.

<iframe src="https://vsmacs.com/embed/nemotron-3-super-120b-a12b" width="100%" height="420" style="border:0" loading="lazy" title="Nemotron 3 super 120B (A12B) speed by machine"></iframe>