60 models at Q4 · 16 GB budget

Local AI models for 16GB VRAM

16 GB is the comfortable laptop and mid-range desktop budget — including many Apple Silicon Macs. 20B–27B dense models fit at Q4, and 14B models can run at higher quality.

Typical fit: 20B–27B at Q4, comfortable 14B at higher quality. · Example:Apple M4

These sizes assume Q4_K_M plus a little runtime overhead.Grade the full catalog on your machineorpick a specific GPU.

Common questions

Is 16GB VRAM enough for a local LLM?
Yes. It covers the models most people actually want to run daily: strong 14B–27B chat and coding checkpoints, plus mid-size image generation.
Can a 16GB Mac run local AI?
An M-series Mac with 16 GB unified memory can run the same class of models as a 16 GB GPU, with speed depending on the chip (base vs Pro vs Max).