How we got this number
Each generated token reads the whole model from memory once: 47 GB. The NVIDIA DGX Spark (128GB) moves 273 GB/s, so the ceiling is 5.8 tokens/sec. Real machines reach a fraction of that ceiling; we use 70%, an assumption until we get measured runs for this kind of machine. Send us your real numbers and we'll replace the estimate.
Faster machines for Qwen2.5 72B
Other models on the NVIDIA DGX Spark (128GB)
| Qwen 3 30B (A3B) | Fast | 56 tok/s |
| Qwen3 Coder 30B (A3B) | Fast | 56 tok/s |
| Qwen 3-VL 30B (A3B) | Fast | 53 tok/s |
| Qwen3.5 35B (A3B) | Fast | 50 tok/s |
| Qwen3.6 35B (A3B) | Fast | 50 tok/s |
| gpt-oss 20B | Fast | 46 tok/s |