Mistral Nemo 12B

Apache 2.0

Mistral AI Β· 12B Β· Dense

Multilingual 12B with 128K context Check if your GPU or Mac can run Mistral Nemo 12B locally β€” 6.7 GB min, 11.2 GB recommended.

2024-07128K context

Quantization Options

QuantBitsVRAMQualityStatus
Q2_K24.3 GBlowβ€”
Q3_K_M35.9 GBmoderateβ€”
Q4_K_M46.6 GBgoodβ€”
Q5_K_M58.2 GBgoodβ€”
Q6_K69.7 GBexcellentβ€”
Q8_0812.8 GBexcellentβ€”
F161625.1 GBlosslessβ€”

About this model

Mistral NeMo is a 12B model built in collaboration with NVIDIA. Mistral NeMo offers a large context window of up to 128k tokens. Its reasoning, world knowledge, and coding accuracy are state-of-the-art in its size category. As it relies on standard architecture, Mistral NeMo is easy to use and a drop-in replacement in any system using Mistral 7B.

nemo-base-performance.png

Reference

Blog

Hugging Face

Can I run Mistral Nemo 12B locally?

Can I run Mistral Nemo 12B locally?
Mistral Nemo 12B needs about 6.7 GB of memory at a minimum and 11.2 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does Mistral Nemo 12B need?
At Q4_K_M, Mistral Nemo 12B uses about 6.6 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.