| Machine | Fits? | Tokens/sec | |
|---|---|---|---|
| Mac Studio M5 Ultra (96GB) | Yes | 82 | Check price → |
| Mac Studio M5 Ultra (256GB) | Yes | 82 | Check price → |
| Mac Studio M5 Max (40-core GPU, 64GB) | Yes | 41 | Check price → |
| Mac mini M5 Pro (64GB) | Yes | 20 | Check price → |
| Mac mini M4 Pro (64GB) older | Yes | 18 | Check price → |
| NVIDIA DGX Spark (128GB) | Yes | 16 | Check price → |
| Ryzen AI Max+ 395 mini PC (128GB) | Yes | 15 | Check price → |
| Mac mini M6 (32GB) | Barely | 6.8 | Check price → |
| Beelink SER7 (Ryzen 7 7840HS, 32GB) | Barely | 2.7 | Check price → |
| Minisforum UM790 Pro (Ryzen 9 7940HS, 32GB) | Barely | 2.7 | Check price → |
| GEEKOM A6 (Ryzen 7 6800H, 32GB) | Barely | 2.3 | Check price → |
| Mac mini M6 (16GB) | No | – | Check price → |
| Mac mini M5 Pro (24GB) | No | – | Check price → |
| Mac mini M4 (16GB) older | No | – | Check price → |
| Mac mini M2 (16GB) older | No | – | Check price → |
This is a mixture-of-experts model: it only reads part of itself for each token, so it runs much faster than its size suggests, but the whole file still has to fit in memory. Around 10 tokens/sec reads comfortably; under 5 feels slow. How we estimate.
Other mixtral sizes
Embed this table
Free for blogs, READMEs and model cards. Updates automatically.
<iframe src="https://vsmacs.com/embed/mixtral-8x7b" width="100%" height="420" style="border:0" loading="lazy" title="mixtral 8X7B speed by machine"></iframe>