mSmeraSaaS
AI & Local

What LLM Can My Mac Run?

Choose your Apple Silicon chip and unified memory and see, at a glance, which open LLMs will run comfortably on your Mac, which fit with a tighter context, and which are too large — with the quantization and memory headroom for each.

Free · runs in your browser · no sign-up to use.

M3 Pro · 36 GB → about 25 GB usable for a model after the OS and app overhead.

Fits comfortably

13
MiniCPM5 1BOptiq
~0.84 GB at Q4 · 3% of usable
sub-gigabyte, with a reasoning mode
LFM2.5-VL 3BOptiq
~1.82 GB at Q4 · 7% of usable
vision
Qwen3.5 4BOptiq
~3.05 GB at Q4 · 12% of usable
Qwen3.5 9BOptiq
~6.61 GB at Q4 · 26% of usable
LFM2.5 1.2B
~0.61 GB at Q4 · 2% of usable
on-device series
LFM2.5 2.6B
~1.41 GB at Q4 · 6% of usable
Qwen3.5 2B
~1.6 GB at Q4 · 6% of usable
Gemma 4 E2B
~3.31 GB at Q4 · 13% of usable
multimodal
LFM2.5 8B-A1B
~4.44 GB at Q4 · 18% of usable
MoE, 1B active
Gemma 4 12B
~10.23 GB at Q4 · 41% of usable
QAT
Devstral Small 2 24B
~14.07 GB at Q4 · 56% of usable
agentic coding
Gemma 4 26B-A4B
~14.29 GB at Q4 · 57% of usable
MoE, 4B active
Qwen3.5 27B
~14.95 GB at Q4 · 59% of usable

Fits with constraints

3
Nemotron 3.5 Lightning 30B-A3BOptiq
~20.55 GB at Q4 · 82% of usable
MoE, 3B active
Gemma 4 31B
~17.15 GB at Q4 · 68% of usable
multimodal
Qwen3.5 35B-A3B
~18.99 GB at Q4 · 75% of usable
MoE, 3B active

Too large

1
Qwen3.5 122B-A10B
~64.81 GB at Q4 · 257% of usable
MoE, 10B active

For chat, the largest model in the “fits comfortably” list is usually the sweet spot.

Next step

Run these on your Mac with Optiq

Optiq ships Apple-Silicon-optimised quants of these models, so you can go from “it fits” to actually running it in a couple of commands.

Estimates only. Real fit depends on the exact quant, context length, and what else is running. Sizes are the measured 4-bit MLX weights on Hugging Face; models badged Optiq have a build published for Apple silicon.

Related tools

Free, and they run in your browser too.