Will it run? / Code Llama 70B / Mac mini M5 Pro (24GB)

Can the Mac mini M5 Pro (24GB) run Code Llama 70B?

Won't fit No. Code Llama 70B needs about 40.5 GB of memory, and the Mac mini M5 Pro (24GB) can give a model about 16.1 GB. You'd need a machine with more memory, or a smaller model.

–estimated tokens/sec
40.5 GBmemory needed
16.1 GBavailable to the model
307 GB/smemory bandwidth

Check the Mac mini M5 Pro (24GB) price →

How we got this number

Each generated token reads the whole model from memory once: 39 GB. The Mac mini M5 Pro (24GB) moves 307 GB/s, so the ceiling is 7.9 tokens/sec. Real machines reach a fraction of that ceiling; we use 80%, calibrated against a measured Mac run. Send us your real numbers and we'll replace the estimate.

Faster machines for Code Llama 70B

Mac Studio M5 Ultra (96GB)25 tok/s$5,499Check price →
Mac Studio M5 Ultra (256GB)25 tok/s$9,499Check price →
Mac Studio M5 Max (40-core GPU, 64GB)13 tok/s$3,799Check price →
Mac mini M5 Pro (64GB)6.3 tok/sCheck price →

Other models on the Mac mini M5 Pro (24GB)

gpt-oss 20BFast60 tok/s
Qwen 3 30B (A3B)Fast43 tok/s
Qwen3 Coder 30B (A3B)Fast43 tok/s
Llama 3.2 Vision 11BFast32 tok/s
Phi 3 14BFast31 tok/s
orca mini 13BFast31 tok/s