How we got this number
Each generated token reads the whole model from memory once: 7.9 GB. The GEEKOM A6 (Ryzen 7 6800H, 32GB) moves 76.8 GB/s, so the ceiling is 9.7 tokens/sec. Real machines reach a fraction of that ceiling; we use 60%, an assumption until we get measured runs for this kind of machine. Send us your real numbers and we'll replace the estimate.
Faster machines for orca mini 13B
Other models on the GEEKOM A6 (Ryzen 7 6800H, 32GB)
| lfm 2.5 8B (A1B) | Fast | 38 tok/s |
| minicpm v4.6 1B | Fast | 29 tok/s |
| Gemma 2 2B | Fast | 29 tok/s |
| codeGemma 2B | Fast | 29 tok/s |
| Gemma 2B | Fast | 27 tok/s |
| smollm 2 1.7B | Fast | 26 tok/s |