Command R 35B

CC BY-NC 4.0

Cohere Β· 35B Β· Dense

Optimized for retrieval-augmented generation Check if your GPU or Mac can run Command R 35B locally β€” 19.6 GB min, 32.6 GB recommended.

2024-03128K context

Quantization Options

QuantBitsVRAMQualityStatus
Q2_K211.7 GBlowβ€”
Q3_K_M316.2 GBmoderateβ€”
Q4_K_M418.4 GBgoodβ€”
Q5_K_M522.9 GBgoodβ€”
Q6_K627.4 GBexcellentβ€”
Q8_0836.4 GBexcellentβ€”
F161672.2 GBlosslessβ€”

About this model

Command R is a generative model optimized for long context tasks such as retrieval-augmented generation (RAG) and using external APIs and tools. As a model built for companies to implement at scale, Command R boasts:

  • Strong accuracy on RAG and Tool Use
  • Low latency, and high throughput
  • Longer 128k context
  • Strong capabilities across 10 key languages

There are currently two versions of Command R:

  • Original release tagged v0.1
  • August 2024 update tagged 08-2024

References

Can I run Command R 35B locally?

Can I run Command R 35B locally?
Command R 35B needs about 19.6 GB of memory at a minimum and 32.6 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does Command R 35B need?
At Q4_K_M, Command R 35B uses about 18.4 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.