Will it run? / Llama 3.3 70B

Can I run Llama 3.3 70B on a Mac or mini PC?

Llama 3.3 70B needs about 44.5 GB of memory. It fits on 7 of the 15 machines we track. The cheapest current machine that runs it at a comfortable 10+ tokens/sec is the Mac Studio M5 Max (40-core GPU, 64GB) ($3,799).

43 GBdownload (Q4_K_M)
~44.5 GBmemory needed
7/15machines it fits on
ollama run llama3.3:70b
MachineMemoryFits?Tokens/secFrom
Mac Studio M5 Ultra (96GB) 96 GBYes23
$5,499Check price →
Mac Studio M5 Ultra (256GB) 256 GBYes23
$9,499Check price →
Mac Studio M5 Max (40-core GPU, 64GB) 64 GBYes11
$3,799Check price →
Mac mini M5 Pro (64GB) 64 GBYes5.7
Check price →
Mac mini M4 Pro (64GB) older 64 GBYes5.1
Check price →
NVIDIA DGX Spark (128GB) 128 GBYes4.4
Check price →
Ryzen AI Max+ 395 mini PC (128GB) 128 GBYes4.2
Check price →
Mac mini M6 (16GB) 16 GBNo–
$899Check price →
Mac mini M6 (32GB) 32 GBNo–
Check price →
Mac mini M5 Pro (24GB) 24 GBNo–
$1,699Check price →
Mac mini M4 (16GB) older 16 GBNo–
Check price →
Mac mini M2 (16GB) older 16 GBNo–
Check price →
GEEKOM A6 (Ryzen 7 6800H, 32GB) 32 GBNo–
Check price →
Beelink SER7 (Ryzen 7 7840HS, 32GB) 32 GBNo–
Check price →
Minisforum UM790 Pro (Ryzen 9 7940HS, 32GB) 32 GBNo–
Check price →

Around 10 tokens/sec reads comfortably; under 5 feels slow. How we estimate.

Embed this table

Free for blogs, READMEs and model cards. Updates automatically.

<iframe src="https://vsmacs.com/embed/llama3-3-70b" width="100%" height="420" style="border:0" loading="lazy" title="Llama 3.3 70B speed by machine"></iframe>