How we got this number
Each generated token reads the active part of the model from memory once: 1.9 GB. The Beelink SER7 (Ryzen 7 7840HS, 32GB) moves 89.6 GB/s, so the ceiling is 47 tokens/sec. Real machines reach a fraction of that ceiling; we use 60%, an assumption until we get measured runs for this kind of machine. Send us your real numbers and we'll replace the estimate.
Faster machines for Qwen3 Coder 30B (A3B)
Other models on the Beelink SER7 (Ryzen 7 7840HS, 32GB)
| lfm 2.5 8B (A1B) | Fast | 44 tok/s |
| Qwen 3 30B (A3B) | Comfortable | 16 tok/s |
| Qwen 3-VL 30B (A3B) | Comfortable | 15 tok/s |
| Llama 2 7B | Comfortable | 14 tok/s |
| Code Llama 7B | Comfortable | 14 tok/s |
| DeepSeek Coder 6.7B | Comfortable | 14 tok/s |