Will it run? / Llama 3.2 Vision 90B / Mac Studio M5 Max (40-core GPU, 64GB)
Can the Mac Studio M5 Max (40-core GPU, 64GB) run Llama 3.2 Vision 90B?
Usable, but slow Yes, at about 5.4 tokens/sec (estimated). It works, but you'll wait on longer answers. It only fits if you raise the GPU memory limit, and it'll be slower than usual.
5.4estimated tokens/sec
56.5 GBmemory needed
48 GBavailable to the model
614 GB/smemory bandwidth
ollama run llama3.2-vision:90b
Check the Mac Studio M5 Max (40-core GPU, 64GB) price →