17 open models

Small on-device LLMs

Models at 4B and under. They fit 8 GB cards, many laptops and some phones. Quality is lower than a 27B, but they start instantly and leave memory for other apps.

GPUVRAMBWRAMCores

FLUX.2 Klein 4B

7mo ago

Black Forest Labs · 4B · Apache 2.0

Fastest open FLUX.2 — sub-second text-to-image and multi-reference editing on consumer GPUs

2.5GB·32K ctx·

Wan 2.2 TI2V 5B

1 year ago

Alibaba · 5B · Apache 2.0

Unified text/image-to-video — the local sweet spot under Apache 2.0

3.1GB·4K ctx·

Z-Image Turbo

9mo ago

Alibaba · 6B · Apache 2.0

8-step distilled image model — photorealism and bilingual text on 16GB cards

3.6GB·4K ctx·

HunyuanVideo 1.5

9mo ago

Tencent · 8.3B · Tencent Hunyuan Community

Compact cinematic video model — strong faces and motion on a single 4090

4.8GB·4K ctx·

Qwen3-VL 8B

11mo ago

Alibaba · 8.8B · Apache 2.0

The community-favourite local VLM — superb OCR, receipts & captioning

5GB·256K ctx·

Qwen 3.5 9B

6mo ago

Alibaba · 9B · Apache 2.0

Multimodal Qwen 3.5 mid-size

22 AA·5.1GB·32K ctx·

LTX 2.3

5mo ago

Lightricks · 19B · LTX-2 Community

Open 4K video with native stereo audio — text, image and video-to-video

10.2GB·4K ctx·

Qwen Image 2512

8mo ago

Alibaba · 20B · Apache 2.0

Open text-to-image with strong English and Chinese typography

10.7GB·4K ctx·

GPT-OSS 20B

1 year ago

OpenAI · 21B · Apache 2.0

OpenAI's open-weight MoE with configurable reasoning

15 AA·11.3GB·128K ctx·

Gemma 4 26B-A4B IT

4mo ago

Google · 27B · Gemma

Gemma 4 MoE instruct model (official)

26 AA·14.3GB·256K ctx·

Qwen 3.8 27B

26d ago

Alibaba · 27B · Apache 2.0

Flagship dense Qwen 3.8 — native multimodal all-rounder with video understanding

52 AA·14.3GB·256K ctx·

Wan 2.2 T2V A14B

1 year ago

Alibaba · 27B · Apache 2.0

Flagship open Wan 2.2 — 14B-active MoE for photoreal text-to-video

14.3GB·4K ctx·

Muse Glimmer 30B

26d ago

Meta · 30B · Apache 2.0

Open agentic 30B distilled from Muse Spark — tool use, vision and local recovery on a single GPU

35 AA·15.9GB·128K ctx·

Qwen3-VL 30B-A3B

1 year ago

Alibaba · 31B · Apache 2.0

Efficient vision MoE — 3B active, strong temporal & document understanding

16.4GB·256K ctx·

FLUX.2 Dev

9mo ago

Black Forest Labs · 32B · FLUX Non-Commercial

Flagship open-weight FLUX.2 — text-to-image and multi-reference editing up to 4MP

16.9GB·32K ctx·

MiniMax H3

26d ago

MiniMax · 33B · MiniMax Community

Open video generation — text/image to 2K video with native stereo audio

17.4GB·32K ctx·

Agents-A1 35B-A3B

2mo ago

InternScience · 35B · Apache 2.0

Efficient multimodal agentic MoE for long-horizon search, engineering and scientific research

18.4GB·256K ctx·

Ornith 1.0 35B-A3B

2mo ago

DeepReinforce · 35B · MIT

Agentic coding MoE with a 3B active working set and self-improving training

18.4GB·256K ctx·

Qwen 3.6 35B-A3B

4mo ago

Alibaba · 36B · Apache 2.0

Big-model quality at 3B-active speed — the mid-hardware sweet spot

32 AA·18.9GB·256K ctx·

HunyuanImage 3.0 Instruct

7mo ago

Tencent · 80B · Tencent Hunyuan Community

Reasoning image model — prompt rewrite, chain-of-thought and image-to-image editing

41.5GB·4K ctx·

Llama 4 Scout 17B

1 year ago

Meta · 109B · Llama 4 Community

MoE with 16 experts, 17B active params

10 AA·56.3GB·128K ctx·

GPT-OSS 120B

1 year ago

OpenAI · 117B · Apache 2.0

OpenAI's flagship open-weight MoE — 52.6% SWE-bench

24 AA·60.4GB·128K ctx·

Mistral Small 4 119B

5mo ago

Mistral AI · 119B · Apache 2.0

Sparse Mistral Small 4 — 6.5B active, strong local all-rounder

20 AA·61.5GB·256K ctx·

Qwen 3 VL 235B-A22B

9mo ago

Alibaba · 235B · Apache 2.0

Flagship vision-language MoE — frontier multimodal reasoning and agentic GUI control

120.9GB·256K ctx·

Hy3

1 month ago

Tencent · 295B · Apache 2.0

Production-focused agentic MoE with strong coding, tool use and long-context reasoning

42 AA·151.6GB·256K ctx·

MiniMax M3

2mo ago

MiniMax · 428B · MiniMax Community

Native multimodal MoE — understands text, image and long video with 1M context

45 AA·219.7GB·1024K ctx·

GLM-5.3

26d ago

Z.ai · 753B · MIT

Same 753B / 40B-active base as GLM-5.2 — post-training lifts coding and long-horizon agents, 1M context

386.2GB·1024K ctx·

LongCat 2.0

1 month ago

Meituan · 1.6T · MIT

Frontier-scale agentic and coding MoE with sparse attention and native 1M context

34 AA·820.1GB·1024K ctx·

DeepSeek V4 Pro

4mo ago

DeepSeek · 1.6T · MIT

Flagship V4 MoE — 49B active, 1M context

820.1GB·1024K ctx·

Qwen 3.8 2.4T-A95B

26d ago

Alibaba · 2.4T · Qwen

Frontier Qwen 3.8 MoE — 95B active, 1M context

58 AA·1229.8GB·1024K ctx·

Kimi K3

1 month ago

Moonshot AI · 2.8T · Kimi

Frontier 2.8T multimodal MoE — 104B active, native video understanding, 1M context

60 AA·1424.5GB·1024K ctx·
All models

Qwen 3 0.6B

1 year ago

Alibaba · 0.6B · Apache 2.0

Ultra-light Qwen 3 model for constrained devices

0.8GB·32K ctx·

Qwen 3.5 0.8B

6mo ago

Alibaba · 0.8B · Apache 2.0

Ultra-tiny model for embedded and edge

5* AA·0.9GB·32K ctx·

Llama 3.2 1B

2y ago

Meta · 1B · Llama 3.2 Community

Meta's smallest Llama for edge devices

1GB·128K ctx·

Gemma 3 1B

1 year ago

Google · 1B · Gemma

Google's tiny Gemma for on-device

1GB·32K ctx·

Wan 2.1 T2V 1.3B

1 year ago

Alibaba · 1.3B · Apache 2.0

Tiny open text-to-video — 480p clips on 8GB consumer GPUs

1.2GB·4K ctx·

Qwen 2.5 Coder 1.5B

1 year ago

Alibaba · 1.5B · Apache 2.0

Ultra-lightweight coding model

1.3GB·32K ctx·

DeepSeek R1 1.5B

1 year ago

DeepSeek · 1.5B · MIT

Tiny reasoning model distilled from R1

1.3GB·64K ctx·

Qwen 3 1.7B

1 year ago

Alibaba · 1.7B · Apache 2.0

Compact multilingual Qwen 3

1.4GB·32K ctx·

Qwen 3.5 2B

6mo ago

Alibaba · 2B · Apache 2.0

Small multimodal Qwen 3.5

7* AA·1.5GB·32K ctx·

Llama 3.2 3B

2y ago

Meta · 3B · Llama 3.2 Community

Lightweight Llama for mobile and edge

2GB·128K ctx·

SmolLM3 3B

1 year ago

HuggingFace · 3B · Apache 2.0

Lightweight multilingual reasoning

2GB·128K ctx·

Granite 4.1 3B

4mo ago

IBM · 3B · Apache 2.0

Compact enterprise model for edge and constrained environments

2GB·128K ctx·

Ministral 3 3B

8mo ago

Mistral AI · 3B · Apache 2.0

Current-gen tiny Ministral — edge chat with 256K context

7 AA·2GB·256K ctx·

Phi-4 Mini Reasoning

1 year ago

Microsoft · 3.8B · MIT

Lightweight reasoning model

2.4GB·16K ctx·

Gemma 3 4B

1 year ago

Google · 4B · Gemma

Multimodal Gemma with 128K context

2.5GB·128K ctx·

Qwen 3.5 4B

6mo ago

Alibaba · 4B · Apache 2.0

Small multimodal Qwen 3.5

20* AA·2.5GB·32K ctx·

Qwen3-VL 4B

11mo ago

Alibaba · 4.4B · Apache 2.0

Compact dedicated vision-language model — OCR & image chat on edge

2.8GB·256K ctx·

Gemma 4 E2B IT

4mo ago

Google · 5B · Gemma

Gemma 4 efficient instruct model (official)

10* AA·3.1GB·256K ctx·

Qwen 2.5 Coder 7B

1 year ago

Alibaba · 7B · Apache 2.0

Dedicated coding model

4.1GB·128K ctx·

DeepSeek R1 Distill 7B

1 year ago

DeepSeek · 7B · MIT

R1 reasoning distilled into Qwen 7B

4.1GB·64K ctx·

Gemma 4 E4B IT

4mo ago

Google · 8B · Gemma

Gemma 4 balanced instruct model (official)

12* AA·4.6GB·256K ctx·

Llama 3.1 8B

2y ago

Meta · 8B · Llama 3.1 Community

Meta's versatile 8B — great quality/speed ratio

4.6GB·128K ctx·

Qwen 3 8B

1 year ago

Alibaba · 8B · Apache 2.0

Qwen 3 with thinking mode support

4.6GB·128K ctx·

Granite 4.1 8B

4mo ago

IBM · 8B · Apache 2.0

Balanced general-purpose enterprise model

4.6GB·128K ctx·

Ministral 8B

1 year ago

Mistral AI · 8B · MRL

Mistral's efficient 8B model

4.6GB·32K ctx·

GLM-4 9B

2y ago

Zhipu AI · 9B · GLM-4

Multilingual model supporting 26 languages with 128K context

5.1GB·128K ctx·

Nemotron Nano 9B v2

1 year ago

NVIDIA · 9B · NVIDIA Open

Hybrid Mamba2 architecture for reasoning

9* AA·5.1GB·128K ctx·

Ornith 1.0 9B

2mo ago

DeepReinforce · 9B · MIT

Self-improving agentic coding model optimized for terminal and software engineering tasks

5.1GB·256K ctx·

FLUX.2 Klein 9B

7mo ago

Black Forest Labs · 9B · FLUX Non-Commercial

Higher-quality distilled FLUX.2 — sub-second generation and multi-reference editing

5.1GB·32K ctx·

Gemma 3 12B

1 year ago

Google · 12B · Gemma

Multimodal Gemma with 128K context

6.6GB·128K ctx·

Mistral Nemo 12B

2y ago

Mistral AI · 12B · Apache 2.0

Multilingual 12B with 128K context

6.6GB·128K ctx·

Gemma 4 12B IT

4mo ago

Google · 12B · Apache 2.0

Gemma 4 mid-size instruct — multimodal any-to-any

22* AA·6.6GB·256K ctx·

Phi-4 14B

1 year ago

Microsoft · 14B · MIT

Microsoft's reasoning-focused model

5* AA·7.7GB·16K ctx·

Qwen 3 14B

1 year ago

Alibaba · 14B · Apache 2.0

Strong all-rounder with thinking mode

7.7GB·128K ctx·

DeepSeek R1 Distill 14B

1 year ago

DeepSeek · 14B · MIT

R1 reasoning distilled into Qwen 14B

7.7GB·64K ctx·

Ministral 3 14B

8mo ago

Mistral AI · 14B · Apache 2.0

Current-gen Ministral mid-size — local assistant with 256K context

11 AA·7.7GB·256K ctx·

LFM2 24B

9mo ago

Liquid AI · 24B · Liquid AI

Hybrid MoE with convolution+attention layers — 2.3B active

5* AA·12.8GB·32K ctx·

Devstral Small 2 24B

8mo ago

Mistral AI · 24B · Apache 2.0

Coding-focused model with 256K context — 68% SWE-bench

12.8GB·256K ctx·

Mistral Small 3.1 24B

1 year ago

Mistral AI · 24B · Apache 2.0

Multimodal Mistral with vision support

12.8GB·128K ctx·

DiffusionGemma 26B-A4B IT

2mo ago

Google · 26B · Apache 2.0

Discrete diffusion MoE — 1100+ tok/s on H100, multimodal (text/image/video)

13* AA·13.8GB·256K ctx·

Qwen 3.5 27B

6mo ago

Alibaba · 27.8B · Apache 2.0

Flagship native multimodal Qwen 3.5

14.7GB·256K ctx·

Qwen 3.6 27B

4mo ago

Alibaba · 27.8B · Apache 2.0

Flagship dense Qwen 3.6 — native multimodal all-rounder

38 AA·14.7GB·256K ctx·

Qwen 3 30B-A3B

1 year ago

Alibaba · 30B · Apache 2.0

MoE with only 3.3B active — extremely efficient

15.9GB·128K ctx·

Nemotron 3 Nano 30B

1 year ago

NVIDIA · 30B · NVIDIA Open

MoE with 1M context and 3B active

15 AA·15.9GB·1024K ctx·

Granite 4.1 30B

4mo ago

IBM · 30B · Apache 2.0

High-capacity enterprise model for complex reasoning and tool use

15.9GB·128K ctx·

North Mini Code

2mo ago

Cohere · 30B · Apache 2.0

Open agentic coding MoE with 3B active — built for software engineering and terminal tasks

15.9GB·256K ctx·

Qwen 3 Coder 30B-A3B

1 year ago

Alibaba · 30B · Apache 2.0

Efficient agentic coding MoE — 3B active, 256K context

15.9GB·256K ctx·

Qwen 3 32B

1 year ago

Alibaba · 32B · Apache 2.0

Qwen 3 flagship dense model

16.9GB·128K ctx·

DeepSeek R1 Distill 32B

1 year ago

DeepSeek · 32B · MIT

R1 reasoning distilled into Qwen 32B — sweet spot

16.9GB·64K ctx·

OLMo 2 32B

1 year ago

Allen AI · 32B · Apache 2.0

Fully open research model by Allen AI

16.9GB·4K ctx·

Gemma 4 31B IT

4mo ago

Google · 33B · Gemma

Gemma 4 flagship instruct model (official)

30 AA·17.4GB·256K ctx·

Gemma 4 31B

4mo ago

Google · 33B · Gemma

Gemma 4 flagship base model (official)

17.4GB·256K ctx·

Command R 35B

2y ago

Cohere · 35B · CC BY-NC 4.0

Optimized for retrieval-augmented generation

18.4GB·128K ctx·

Qwen 3.5 35B-A3B

6mo ago

Alibaba · 35B · Apache 2.0

Efficient multimodal MoE with 3B active

18.4GB·256K ctx·

Mixtral 8x7B

2y ago

Mistral AI · 47B · Apache 2.0

MoE with 12.9B active params

24.6GB·32K ctx·

Llama 3.3 70B

1 year ago

Meta · 70B · Llama 3.3 Community

Best open model at 70B class

9* AA·36.4GB·128K ctx·

Qwen 3 Next 80B-A3B

8mo ago

Alibaba · 80B · Apache 2.0

High-sparsity MoE — extreme low activation ratio for fast inference at 80B scale

41.5GB·256K ctx·

Qwen 3 Coder Next 80B-A3B

6mo ago

Alibaba · 80B · Apache 2.0

Ultra-efficient agentic coding MoE optimized for tool-calling coding agents

41.5GB·256K ctx·

HunyuanImage 3.0

1 year ago

Tencent · 80B · Tencent Hunyuan Community

Largest open image MoE — 13B active, strong long-prompt generation

41.5GB·4K ctx·

GLM-4.5 Air

1 year ago

Z.ai · 106B · MIT

Consumer-friendly GLM MoE — 12B active, strong agentic & tool use

54.8GB·128K ctx·

Qwen 3.5 122B-A10B

6mo ago

Alibaba · 122B · Apache 2.0

Large multimodal MoE

33 AA·63GB·256K ctx·

DeepSeek V4 Flash

4mo ago

DeepSeek · 158B · MIT

Efficient long-context V4 — 13B active, 1M context

81.4GB·1024K ctx·

Qwen 3 235B-A22B

1 year ago

Alibaba · 235B · Apache 2.0

Massive MoE with 22B active — frontier quality

120.9GB·128K ctx·

GLM-4.6

1 year ago

Z.ai · 357B · MIT

Large GLM MoE with strong coding and 200K context

183.4GB·195K ctx·

Qwen 3.5 397B-A17B

6mo ago

Alibaba · 397B · Apache 2.0

Largest multimodal Qwen 3.5 MoE

34 AA·203.9GB·256K ctx·

Llama 4 Maverick 17B-128E

1 year ago

Meta · 400B · Llama 4 Community

Multimodal MoE with 128 experts — 17B active, 1M context

14 AA·205.4GB·1024K ctx·

Qwen 3 Coder 480B

1 year ago

Alibaba · 480B · Apache 2.0

Largest open coding MoE — 35B active

246.4GB·256K ctx·

DeepSeek R1

1 year ago

DeepSeek · 671B · MIT

Massive MoE reasoning model — 37B active

344.2GB·64K ctx·

DeepSeek V3.2

8mo ago

DeepSeek · 685B · MIT

State-of-the-art MoE — 37B active params

351.4GB·128K ctx·

GLM-5

6mo ago

Zhipu AI · 744B · MIT

MoE with 256 experts, 40B active — frontier-class agentic coding

381.6GB·128K ctx·

GLM-5.2

2mo ago

Z.ai · 753B · MIT

Frontier open-weight coder — top SWE-bench, 1M context

53 AA·386.2GB·1024K ctx·

GLM-5.1

4mo ago

Zhipu AI · 754B · MIT

Improved agentic coding — SOTA SWE-bench Pro, long-horizon tasks

386.7GB·128K ctx·

Kimi K2.6

4mo ago

Moonshot AI · 1.06T · Kimi

Natively multimodal 1T MoE — 32B active, frontier agentic

542.4GB·256K ctx·

Common questions

What is a small language model?
A small language model has only a few billion parameters — here, 4B or fewer. They run on modest hardware and are built for edge, mobile and always-on assistants.
Can I run an LLM on 8 GB of RAM?
Yes. Most models on this page are designed for that budget. Q4 quantization keeps them well under 8 GB of VRAM or unified memory.
Are tiny models good enough for chat?
They are strong for drafts, classification and simple coding. For hard reasoning or long documents, step up to a 9B–27B if your machine has the memory.

All product names, logos, and brands are property of their respective owners. Apple, NVIDIA, AMD, Intel, Qualcomm, and all AI model names mentioned on this site are trademarks or registered trademarks of their respective holders. This site is not affiliated with or endorsed by any of these companies.