Will it run? / Llama 3.2 Vision 90B / NVIDIA DGX Spark (128GB)

Can the NVIDIA DGX Spark (128GB) run Llama 3.2 Vision 90B?

Too slow to enjoy Yes, at about 3.5 tokens/sec (estimated). It technically runs, but slowly enough to be frustrating.

3.5estimated tokens/sec
56.5 GBmemory needed
108.8 GBavailable to the model
273 GB/smemory bandwidth
ollama run llama3.2-vision:90b

Check the NVIDIA DGX Spark (128GB) price →

How we got this number

Each generated token reads the whole model from memory once: 55 GB. The NVIDIA DGX Spark (128GB) moves 273 GB/s, so the ceiling is 5.0 tokens/sec. Real machines reach a fraction of that ceiling; we use 70%, an assumption until we get measured runs for this kind of machine. Send us your real numbers and we'll replace the estimate.

Faster machines for Llama 3.2 Vision 90B

Mac Studio M5 Ultra (96GB)18 tok/s$5,499Check price →
Mac Studio M5 Ultra (256GB)18 tok/s$9,499Check price →
Mac Studio M5 Max (40-core GPU, 64GB)5.4 tok/s$3,799Check price →
Ryzen AI Max+ 395 mini PC (128GB)3.3 tok/sCheck price →

Other models on the NVIDIA DGX Spark (128GB)

Qwen 3 30B (A3B)Fast56 tok/s
Qwen3 Coder 30B (A3B)Fast56 tok/s
Qwen 3-VL 30B (A3B)Fast53 tok/s
Qwen3.5 35B (A3B)Fast50 tok/s
Qwen3.6 35B (A3B)Fast50 tok/s
gpt-oss 20BFast46 tok/s