Will it run? / Llama 3.1 70B / Mac mini M4 (16GB)

Can the Mac mini M4 (16GB) run Llama 3.1 70B?

Won't fit No. Llama 3.1 70B needs about 44.5 GB of memory, and the Mac mini M4 (16GB) can give a model about 10.7 GB. You'd need a machine with more memory, or a smaller model.

–estimated tokens/sec
44.5 GBmemory needed
10.7 GBavailable to the model
120 GB/smemory bandwidth

Check the Mac mini M4 (16GB) price →

How we got this number

Each generated token reads the whole model from memory once: 43 GB. The Mac mini M4 (16GB) moves 120 GB/s, so the ceiling is 2.8 tokens/sec. Real machines reach a fraction of that ceiling; we use 80%, calibrated against a measured Mac run. Send us your real numbers and we'll replace the estimate.

Faster machines for Llama 3.1 70B

Mac Studio M5 Ultra (96GB)23 tok/s$5,499Check price →
Mac Studio M5 Ultra (256GB)23 tok/s$9,499Check price →
Mac Studio M5 Max (40-core GPU, 64GB)11 tok/s$3,799Check price →
Mac mini M5 Pro (64GB)5.7 tok/sCheck price →

Other models on the Mac mini M4 (16GB)

DeepSeek Coder v2 16BComfortable11 tok/s
DeepSeek-R1 14BComfortable11 tok/s
Qwen2.5 14BComfortable11 tok/s
Qwen2.5 Coder 14BComfortable11 tok/s
Phi 4 14BComfortable11 tok/s
Qwen 3 14BUsable, but slow6.2 tok/s