Will it run? / Llama 3.2 Vision 11B / NVIDIA DGX Spark (128GB)

Can the NVIDIA DGX Spark (128GB) run Llama 3.2 Vision 11B?

Fast Yes, at about 25 tokens/sec (estimated). Replies stream faster than you can read.

25estimated tokens/sec
9.3 GBmemory needed
108.8 GBavailable to the model
273 GB/smemory bandwidth
ollama run llama3.2-vision:11b

Check the NVIDIA DGX Spark (128GB) price →

How we got this number

Each generated token reads the whole model from memory once: 7.8 GB. The NVIDIA DGX Spark (128GB) moves 273 GB/s, so the ceiling is 35 tokens/sec. Real machines reach a fraction of that ceiling; we use 70%, an assumption until we get measured runs for this kind of machine. Send us your real numbers and we'll replace the estimate.

Faster machines for Llama 3.2 Vision 11B

Mac Studio M5 Ultra (96GB)120+ tok/s$5,499Check price →
Mac Studio M5 Ultra (256GB)120+ tok/s$9,499Check price →
Mac Studio M5 Max (40-core GPU, 64GB)63 tok/s$3,799Check price →
Mac mini M5 Pro (24GB)32 tok/s$1,699Check price →

Other models on the NVIDIA DGX Spark (128GB)

lfm 2.5 8B (A1B)Fast120+ tok/s
minicpm v4.6 1BFast119 tok/s
Gemma 2 2BFast119 tok/s
codeGemma 2BFast119 tok/s
Gemma 2BFast112 tok/s
smollm 2 1.7BFast106 tok/s