45 models at Q4 · 12 GB budget

Local AI models for 12GB VRAM

12 GB unlocks the 12B–14B class at Q4 and leaves headroom for context. It is the step up from an 8 GB card without jumping to 16 GB workstation memory.

Typical fit: 12B–14B dense models, or smaller mixture-of-experts. · Example:RTX 4070

These sizes assume Q4_K_M plus a little runtime overhead.Grade the full catalog on your machineorpick a specific GPU.

Common questions

What can I run on a 12GB GPU?
Most 12B–14B instruct models at Q4, plus 7B–9B at higher quality. Compact image models run well. Video still wants more memory.