Will it run? / Qwen3 Coder 30B (A3B) / NVIDIA DGX Spark (128GB)

Can the NVIDIA DGX Spark (128GB) run Qwen3 Coder 30B (A3B)?

Fast Yes, at about 56 tokens/sec (estimated). Replies stream faster than you can read.

56estimated tokens/sec
20.5 GBmemory needed
108.8 GBavailable to the model
273 GB/smemory bandwidth
ollama run qwen3-coder:30b-a3b

Check the NVIDIA DGX Spark (128GB) price →

How we got this number

Each generated token reads the active part of the model from memory once: 1.9 GB. The NVIDIA DGX Spark (128GB) moves 273 GB/s, so the ceiling is 120+ tokens/sec. Real machines reach a fraction of that ceiling; we use 70%, an assumption until we get measured runs for this kind of machine. Send us your real numbers and we'll replace the estimate.

Faster machines for Qwen3 Coder 30B (A3B)

Mac Studio M5 Max (40-core GPU, 64GB)120+ tok/s$3,799Check price →
Mac Studio M5 Ultra (96GB)120+ tok/s$5,499Check price →
Mac Studio M5 Ultra (256GB)120+ tok/s$9,499Check price →
Mac mini M5 Pro (64GB)72 tok/sCheck price →

Other models on the NVIDIA DGX Spark (128GB)

lfm 2.5 8B (A1B)Fast120+ tok/s
Qwen 3 30B (A3B)Fast56 tok/s
Qwen 3-VL 30B (A3B)Fast53 tok/s
Llama 2 7BFast50 tok/s
Code Llama 7BFast50 tok/s
DeepSeek Coder 6.7BFast50 tok/s