74 models at Q4 · 32 GB budget

Local AI models for 32GB+ VRAM

32 GB and above is where large open mixture-of-experts and high-quality image or video models stop feeling like a stretch. Headroom also means longer context without spilling to system RAM.

Typical fit: Larger open MoE and high-quality image generation. · Example:RTX 5090

These sizes assume Q4_K_M plus a little runtime overhead.Grade the full catalog on your machineorpick a specific GPU.

Common questions

Do I need 32GB to run local AI well?
No. Most daily chat and coding fits in 8–16 GB. 32 GB matters when you want large MoE, long context, or local video at a usable resolution.