mSmeraSaaS
AI & Local

LLM Memory Calculator

Two estimates in one place. Running a model: weights at your chosen quantisation, the KV cache for your context length, and runtime overhead. Fine-tuning it: frozen base, gradients, optimizer state and activations for full, LoRA or QLoRA training, so you can see why a model that runs fine will not fine-tune on the same machine.

Free · runs in your browser · no sign-up to use.

Assumes 0.5 bytes/param for weights and a 2-byte KV cache, plus ~18% runtime overhead. KV is estimated from the parameter count.

Estimated memory to run
Weights3.73 GB
KV cache4.25 GB
Runtime overhead1.44 GB
Total9.42 GB

Estimate, not a guarantee. Actual use varies with the runtime, batching, and quant format.

Next step

Find a matching Apple-Silicon quant to run locally.

Browse Optiq models

Related tools

Free, and they run in your browser too.