Mixtral 8x7B

Apache 2.0

Mistral AI Β· 47B (12.9B active) Β· Mixture of Experts

MoE with 12.9B active params Check if your GPU or Mac can run Mixtral 8x7B locally β€” 26.3 GB min, 43.8 GB recommended.

2023-1232K context

Mixture of Experts

Total experts: 8
Active experts: 2
Active params: 12.9B

Quantization Options

QuantBitsVRAMQualityStatus
Q2_K215.5 GBlowβ€”
Q3_K_M321.6 GBmoderateβ€”
Q4_K_M424.6 GBgoodβ€”
Q5_K_M530.6 GBgoodβ€”
Q6_K636.6 GBexcellentβ€”
Q8_0848.6 GBexcellentβ€”
F161696.8 GBlosslessβ€”

About this model

The Mixtral large Language Models (LLM) are a set of pretrained generative Sparse Mixture of Experts.

Sizes

  • mixtral:8x22b
  • mixtral:8x7b

Mixtral 8x22b

ollama run mixtral:8x22b

Mixtral 8x22B sets a new standard for performance and efficiency within the AI community. It is a sparse Mixture-of-Experts (SMoE) model that uses only 39B active parameters out of 141B, offering unparalleled cost efficiency for its size.

Mixtral 8x22B comes with the following strengths:

  • It is fluent in English, French, Italian, German, and Spanish
  • It has strong maths and coding capabilities
  • It is natively capable of function calling
  • 64K tokens context window allows precise information recall from large documents

References

Announcement

HuggingFace

Can I run Mixtral 8x7B locally?

Can I run Mixtral 8x7B locally?
Mixtral 8x7B needs about 26.3 GB of memory at a minimum and 43.8 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does Mixtral 8x7B need?
At Q4_K_M, Mixtral 8x7B uses about 24.6 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.