How we got this number
Each generated token reads the whole model from memory once: 7.8 GB. The Ryzen AI Max+ 395 mini PC (128GB) moves 256 GB/s, so the ceiling is 33 tokens/sec. Real machines reach a fraction of that ceiling; we use 70%, an assumption until we get measured runs for this kind of machine. Send us your real numbers and we'll replace the estimate.
Faster machines for Llama 3.2 Vision 11B
Other models on the Ryzen AI Max+ 395 mini PC (128GB)
| lfm 2.5 8B (A1B) | Fast | 120+ tok/s |
| minicpm v4.6 1B | Fast | 112 tok/s |
| Gemma 2 2B | Fast | 112 tok/s |
| codeGemma 2B | Fast | 112 tok/s |
| Gemma 2B | Fast | 105 tok/s |
| smollm 2 1.7B | Fast | 100 tok/s |