Liquid AI · 24B (2.3B active) · Mezcla de expertos
Hybrid MoE with convolution+attention layers — 2.3B active Comprueba si tu GPU o Mac puede ejecutar LFM2 24B localmente — 13.4 GB mínimo, 22.4 GB recomendado.
2025-1132K contexto
Mezcla de expertos
Expertos totales: 64
Expertos activos: 4
Parámetros activos: 2.3B
| Cuant | Bits | VRAM | Calidad | Estado |
|---|---|---|---|---|
| Q2_K | 2 | 8.2 GB | low | — |
| Q3_K_M | 3 | 11.3 GB | moderate | — |
| Q4_K_M | 4 | 12.8 GB | good | — |
| Q5_K_M | 5 | 15.9 GB | good | — |
| Q6_K | 6 | 18.9 GB | excellent | — |
| Q8_0 | 8 | 25.1 GB | excellent | — |
| F16 | 16 | 49.7 GB | lossless | — |
Sobre este modelo
LFM2 is a family of hybrid models designed for on-device deployment. LFM2-24B-A2B is the largest model in the family, scaling the architecture to 24 billion parameters while keeping inference efficient.
- Best-in-class efficiency: A 24B MoE model with only 2B active parameters per token, fitting in 32 GB of RAM for deployment on consumer laptops and desktops.
- Fast edge inference: 112 tok/s decode on AMD CPU, 293 tok/s on H100. Fits in 32B GB of RAM.
- Predictable scaling: Quality improves log-linearly from 350M to 24B total parameters, confirming the LFM2 hybrid architecture scales reliably across nearly two orders of magnitude.
¿Puedo ejecutar LFM2 24B localmente?
- ¿Puedo ejecutar LFM2 24B localmente?
- LFM2 24B necesita alrededor de 13.4 GB de memoria como mínimo y 22.4 GB recomendados. Abre esta página para evaluarlo con tu GPU o Mac, y luego ejecútalo con runai, Ollama o LM Studio.
- ¿Cuánta VRAM necesita LFM2 24B?
- En Q4_K_M, LFM2 24B usa aproximadamente 12.8 GB de VRAM. Cuantizaciones más altas necesitan más memoria; las más bajas caben en tarjetas más ajustadas con una pérdida de calidad.