Meta
Muse models you can run locally
Meta develops Muse and Llama, from efficient edge Llama variants to open agentic Muse models designed to run locally.
Available models
Ordered from the lightest model to the largest.
Llama 3.2 1B
1B · Dense
Meta's smallest Llama for edge devices
128K ctx
Llama 3.2 3B
3B · Dense
Lightweight Llama for mobile and edge
128K ctx
Llama 3.1 8B
8B · Dense
Meta's versatile 8B — great quality/speed ratio
128K ctx
Muse Glimmer 30B
30B · Dense
Open agentic 30B distilled from Muse Spark — tool use, vision and local recovery on a single GPU
128K ctx
Llama 3.3 70B
70B · Dense
Best open model at 70B class
128K ctx
Llama 4 Scout 17B
109B · 17B active · MoE
MoE with 16 experts, 17B active params
128K ctx
Llama 4 Maverick 17B-128E
400B · 17B active · MoE
Multimodal MoE with 128 experts — 17B active, 1M context
1024K ctx