73 models at Q4 路 24 GB budget

Local AI models for 24GB VRAM

24 GB is the enthusiast sweet spot. 30B dense models fit, mid-size mixture-of-experts become usable, and local image or short video generation is realistic.

Typical fit: 30B dense models, mid-size MoE, or local video.Example:RTX 4090

These sizes assume Q4_K_M plus a little runtime overhead.Grade the full catalog on your machineorpick a specific GPU.

Common questions

What models can an RTX 4090 run locally?
Almost every popular open chat and coding model at high quality, plus FLUX-class image generation and the lighter video checkpoints.