AI & Local
What LLM Can My Mac Run?
Choose your Apple Silicon chip and unified memory and see, at a glance, which open LLMs will run comfortably on your Mac, which fit with a tighter context, and which are too large — with the quantization and memory headroom for each.
Free · runs in your browser · no sign-up to use.
M3 Pro · 36 GB → about 25 GB usable for a model after the OS and app overhead.
Fits comfortably
10Llama 3.2 1B
~0.8 GB at Q4 · 3% of usable
Gemma 3 1B
~0.9 GB at Q4 · 4% of usable
Llama 3.2 3B
~2 GB at Q4 · 8% of usable
Phi-3 mini (3.8B)
~2.3 GB at Q4 · 9% of usable
Qwen2.5 7B
~4.4 GB at Q4 · 17% of usable
Mistral 7B
~4.4 GB at Q4 · 17% of usable
Llama 3.1 8B
~4.9 GB at Q4 · 19% of usable
Gemma 2 9B
~5.4 GB at Q4 · 21% of usable
Phi-3 medium (14B)
~8.6 GB at Q4 · 34% of usable
Qwen2.5 14B
~9 GB at Q4 · 36% of usable
Fits with constraints
2Gemma 2 27B
~16 GB at Q4 · 63% of usable
Qwen2.5 32B
~19 GB at Q4 · 75% of usable
Too large
4Llama 3.3 70B
~40 GB at Q4 · 159% of usable
Qwen2.5 72B
~42 GB at Q4 · 167% of usable
Mixtral 8x7B
~26 GB at Q4 · 103% of usable
Llama 3.1 405B
~230 GB at Q4 · 913% of usable
For chat, the largest model in the “fits comfortably” list is usually the sweet spot.
Next step
Run these on your Mac with OptiQ
OptiQ ships Apple-Silicon-optimised quants of these models, so you can go from “it fits” to actually running it in a couple of commands.
Estimates only. Real fit depends on the exact quant, context length, and what else is running. Sizes shown are for ~Q4_K_M weights.