Meta · 8B · Dense

Meta's versatile 8B — great quality/speed ratio Check if your GPU or Mac can run Llama 3.1 8B locally — 4.5 GB min, 7.5 GB recommended.

2024-07128K context

Quantization Options

QuantBitsVRAMQualityStatus
Q2_K23.1 GBlow—
Q3_K_M34.1 GBmoderate—
Q4_K_M44.6 GBgood—
Q5_K_M55.6 GBgood—
Q6_K66.6 GBexcellent—
Q8_088.7 GBexcellent—
F161616.9 GBlossless—

About this model

Meta Llama 3.1

image.png

Llama 3.1 family of models available:

  • 8B
  • 70B
  • 405B

Llama 3.1 405B is the first openly available model that rivals the top AI models when it comes to state-of-the-art capabilities in general knowledge, steerability, math, tool use, and multilingual translation.

The upgraded versions of the 8B and 70B models are multilingual and have a significantly longer context length of 128K, state-of-the-art tool use, and overall stronger reasoning capabilities. This enables Meta’s latest models to support advanced use cases, such as long-form text summarization, multilingual conversational agents, and coding assistants.

Meta also has made changes to their license, allowing developers to use the outputs from Llama models, including the 405B model, to improve other models.

Model evaluations

For this release, Meta has evaluation the performance on over 150 benchmark datasets that span a wide range of languages. In addition, Meta performed extensive human evaluations that compare Llama 3.1 with competing models in real-world scenarios. Meta’s experimental evaluation suggests that our flagship model is competitive with leading foundation models across a range of tasks, including GPT-4, GPT-4o, and Claude 3.5 Sonnet. Additionally, Meta’s smaller models are competitive with closed and open models that have a similar number of parameters.

image.png

image.png

image.png

References

Can I run Llama 3.1 8B locally?

Can I run Llama 3.1 8B locally?
Llama 3.1 8B needs about 4.5 GB of memory at a minimum and 7.5 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does Llama 3.1 8B need?
At Q4_K_M, Llama 3.1 8B uses about 4.6 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.