How we got this number
Each generated token reads the whole model from memory once: 2.3 GB. The Minisforum UM790 Pro (Ryzen 9 7940HS, 32GB) moves 89.6 GB/s, so the ceiling is 39 tokens/sec. Real machines reach a fraction of that ceiling; we use 60%, an assumption until we get measured runs for this kind of machine. Send us your real numbers and we'll replace the estimate.
Faster machines for Qwen 4B
Other models on the Minisforum UM790 Pro (Ryzen 9 7940HS, 32GB)
| Qwen3.5 0.8B | Fast | 54 tok/s |
| DeepSeek-R1 1.5B | Fast | 49 tok/s |
| Qwen 1.8B | Fast | 49 tok/s |
| Llama 3.2 1B | Fast | 41 tok/s |
| Qwen 3 1.7B | Fast | 38 tok/s |
| Granite 3.1 moe 1B | Fast | 38 tok/s |