Will it run? / mixtral 8X7B

Can I run mixtral 8X7B on a Mac or mini PC?

mixtral 8X7B needs about 27.5 GB of memory. It fits on 11 of the 15 machines we track. The cheapest current machine that runs it at a comfortable 10+ tokens/sec is the Mac Studio M5 Max (40-core GPU, 64GB) ($3,799).

26 GBdownload (Q4_K_M)
~27.5 GBmemory needed
7.2 GBread per token (MoE)
11/15machines it fits on
ollama run mixtral:8x7b
MachineMemoryFits?Tokens/secFrom
Mac Studio M5 Ultra (96GB) 96 GBYes82
$5,499Check price →
Mac Studio M5 Ultra (256GB) 256 GBYes82
$9,499Check price →
Mac Studio M5 Max (40-core GPU, 64GB) 64 GBYes41
$3,799Check price →
Mac mini M5 Pro (64GB) 64 GBYes20
Check price →
Mac mini M4 Pro (64GB) older 64 GBYes18
Check price →
NVIDIA DGX Spark (128GB) 128 GBYes16
Check price →
Ryzen AI Max+ 395 mini PC (128GB) 128 GBYes15
Check price →
Mac mini M6 (32GB) 32 GBBarely6.8
Check price →
Beelink SER7 (Ryzen 7 7840HS, 32GB) 32 GBBarely2.7
Check price →
Minisforum UM790 Pro (Ryzen 9 7940HS, 32GB) 32 GBBarely2.7
Check price →
GEEKOM A6 (Ryzen 7 6800H, 32GB) 32 GBBarely2.3
Check price →
Mac mini M6 (16GB) 16 GBNo–
$899Check price →
Mac mini M5 Pro (24GB) 24 GBNo–
$1,699Check price →
Mac mini M4 (16GB) older 16 GBNo–
Check price →
Mac mini M2 (16GB) older 16 GBNo–
Check price →

This is a mixture-of-experts model: it only reads part of itself for each token, so it runs much faster than its size suggests, but the whole file still has to fit in memory. Around 10 tokens/sec reads comfortably; under 5 feels slow. How we estimate.

Other mixtral sizes

Embed this table

Free for blogs, READMEs and model cards. Updates automatically.

<iframe src="https://vsmacs.com/embed/mixtral-8x7b" width="100%" height="420" style="border:0" loading="lazy" title="mixtral 8X7B speed by machine"></iframe>