GLM-4.6

MIT

Z.ai Β· 357B (32B active) Β· Mixture of Experts

Large GLM MoE with strong coding and 200K context Check if your GPU or Mac can run GLM-4.6 locally β€” 199.5 GB min, 332.5 GB recommended.

2025-09195K context

Mixture of Experts

Total experts: 160
Active experts: 8
Active params: 32.0B

Quantization Options

QuantBitsVRAMQualityStatus
Q2_K2114.8 GBlowβ€”
Q3_K_M3160.5 GBmoderateβ€”
Q4_K_M4183.4 GBgoodβ€”
Q5_K_M5229.1 GBgoodβ€”
Q6_K6274.8 GBexcellentβ€”
Q8_08366.2 GBexcellentβ€”
F1616732 GBlosslessβ€”

Can I run GLM-4.6 locally?

Can I run GLM-4.6 locally?
GLM-4.6 needs about 199.5 GB of memory at a minimum and 332.5 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does GLM-4.6 need?
At Q4_K_M, GLM-4.6 uses about 183.4 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.