6 open models

Local video generation models

Open video models are the heaviest thing you can run locally. Start with the smallest Wan or LTX checkpoints unless you have 24 GB or more.

GPUVRAMBWRAMCoresβ€”

FLUX.2 Klein 4B

7mo ago

Black Forest Labs Β· 4B Β· Apache 2.0

Fastest open FLUX.2 β€” sub-second text-to-image and multi-reference editing on consumer GPUs

2.5GBΒ·32K ctxΒ·

Wan 2.2 TI2V 5B

1 year ago

Alibaba Β· 5B Β· Apache 2.0

Unified text/image-to-video β€” the local sweet spot under Apache 2.0

3.1GBΒ·4K ctxΒ·

Z-Image Turbo

9mo ago

Alibaba Β· 6B Β· Apache 2.0

8-step distilled image model β€” photorealism and bilingual text on 16GB cards

3.6GBΒ·4K ctxΒ·

HunyuanVideo 1.5

9mo ago

Tencent Β· 8.3B Β· Tencent Hunyuan Community

Compact cinematic video model β€” strong faces and motion on a single 4090

4.8GBΒ·4K ctxΒ·

Qwen3-VL 8B

11mo ago

Alibaba Β· 8.8B Β· Apache 2.0

The community-favourite local VLM β€” superb OCR, receipts & captioning

5GBΒ·256K ctxΒ·

Qwen 3.5 9B

6mo ago

Alibaba Β· 9B Β· Apache 2.0

Multimodal Qwen 3.5 mid-size

22 AAΒ·5.1GBΒ·32K ctxΒ·

LTX 2.3

5mo ago

Lightricks Β· 19B Β· LTX-2 Community

Open 4K video with native stereo audio β€” text, image and video-to-video

10.2GBΒ·4K ctxΒ·

Qwen Image 2512

8mo ago

Alibaba Β· 20B Β· Apache 2.0

Open text-to-image with strong English and Chinese typography

10.7GBΒ·4K ctxΒ·

GPT-OSS 20B

1 year ago

OpenAI Β· 21B Β· Apache 2.0

OpenAI's open-weight MoE with configurable reasoning

15 AAΒ·11.3GBΒ·128K ctxΒ·

Gemma 4 26B-A4B IT

4mo ago

Google Β· 27B Β· Gemma

Gemma 4 MoE instruct model (official)

26 AAΒ·14.3GBΒ·256K ctxΒ·

Qwen 3.8 27B

26d ago

Alibaba Β· 27B Β· Apache 2.0

Flagship dense Qwen 3.8 β€” native multimodal all-rounder with video understanding

52 AAΒ·14.3GBΒ·256K ctxΒ·

Wan 2.2 T2V A14B

1 year ago

Alibaba Β· 27B Β· Apache 2.0

Flagship open Wan 2.2 β€” 14B-active MoE for photoreal text-to-video

14.3GBΒ·4K ctxΒ·

Muse Glimmer 30B

26d ago

Meta Β· 30B Β· Apache 2.0

Open agentic 30B distilled from Muse Spark β€” tool use, vision and local recovery on a single GPU

35 AAΒ·15.9GBΒ·128K ctxΒ·

Qwen3-VL 30B-A3B

1 year ago

Alibaba Β· 31B Β· Apache 2.0

Efficient vision MoE β€” 3B active, strong temporal & document understanding

16.4GBΒ·256K ctxΒ·

FLUX.2 Dev

9mo ago

Black Forest Labs Β· 32B Β· FLUX Non-Commercial

Flagship open-weight FLUX.2 β€” text-to-image and multi-reference editing up to 4MP

16.9GBΒ·32K ctxΒ·

MiniMax H3

26d ago

MiniMax Β· 33B Β· MiniMax Community

Open video generation β€” text/image to 2K video with native stereo audio

17.4GBΒ·32K ctxΒ·

Agents-A1 35B-A3B

2mo ago

InternScience Β· 35B Β· Apache 2.0

Efficient multimodal agentic MoE for long-horizon search, engineering and scientific research

18.4GBΒ·256K ctxΒ·

Ornith 1.0 35B-A3B

2mo ago

DeepReinforce Β· 35B Β· MIT

Agentic coding MoE with a 3B active working set and self-improving training

18.4GBΒ·256K ctxΒ·

Qwen 3.6 35B-A3B

4mo ago

Alibaba Β· 36B Β· Apache 2.0

Big-model quality at 3B-active speed β€” the mid-hardware sweet spot

32 AAΒ·18.9GBΒ·256K ctxΒ·

HunyuanImage 3.0 Instruct

7mo ago

Tencent Β· 80B Β· Tencent Hunyuan Community

Reasoning image model β€” prompt rewrite, chain-of-thought and image-to-image editing

41.5GBΒ·4K ctxΒ·

Llama 4 Scout 17B

1 year ago

Meta Β· 109B Β· Llama 4 Community

MoE with 16 experts, 17B active params

10 AAΒ·56.3GBΒ·128K ctxΒ·

GPT-OSS 120B

1 year ago

OpenAI Β· 117B Β· Apache 2.0

OpenAI's flagship open-weight MoE β€” 52.6% SWE-bench

24 AAΒ·60.4GBΒ·128K ctxΒ·

Mistral Small 4 119B

5mo ago

Mistral AI Β· 119B Β· Apache 2.0

Sparse Mistral Small 4 β€” 6.5B active, strong local all-rounder

20 AAΒ·61.5GBΒ·256K ctxΒ·

Qwen 3 VL 235B-A22B

9mo ago

Alibaba Β· 235B Β· Apache 2.0

Flagship vision-language MoE β€” frontier multimodal reasoning and agentic GUI control

120.9GBΒ·256K ctxΒ·

Hy3

1 month ago

Tencent Β· 295B Β· Apache 2.0

Production-focused agentic MoE with strong coding, tool use and long-context reasoning

42 AAΒ·151.6GBΒ·256K ctxΒ·

MiniMax M3

2mo ago

MiniMax Β· 428B Β· MiniMax Community

Native multimodal MoE β€” understands text, image and long video with 1M context

45 AAΒ·219.7GBΒ·1024K ctxΒ·

GLM-5.3

26d ago

Z.ai Β· 753B Β· MIT

Same 753B / 40B-active base as GLM-5.2 β€” post-training lifts coding and long-horizon agents, 1M context

386.2GBΒ·1024K ctxΒ·

LongCat 2.0

1 month ago

Meituan Β· 1.6T Β· MIT

Frontier-scale agentic and coding MoE with sparse attention and native 1M context

34 AAΒ·820.1GBΒ·1024K ctxΒ·

DeepSeek V4 Pro

4mo ago

DeepSeek Β· 1.6T Β· MIT

Flagship V4 MoE β€” 49B active, 1M context

820.1GBΒ·1024K ctxΒ·

Qwen 3.8 2.4T-A95B

26d ago

Alibaba Β· 2.4T Β· Qwen

Frontier Qwen 3.8 MoE β€” 95B active, 1M context

58 AAΒ·1229.8GBΒ·1024K ctxΒ·

Kimi K3

1 month ago

Moonshot AI Β· 2.8T Β· Kimi

Frontier 2.8T multimodal MoE β€” 104B active, native video understanding, 1M context

60 AAΒ·1424.5GBΒ·1024K ctxΒ·
All models

Qwen 3 0.6B

1 year ago

Alibaba Β· 0.6B Β· Apache 2.0

Ultra-light Qwen 3 model for constrained devices

0.8GBΒ·32K ctxΒ·

Qwen 3.5 0.8B

6mo ago

Alibaba Β· 0.8B Β· Apache 2.0

Ultra-tiny model for embedded and edge

5* AAΒ·0.9GBΒ·32K ctxΒ·

Llama 3.2 1B

2y ago

Meta Β· 1B Β· Llama 3.2 Community

Meta's smallest Llama for edge devices

1GBΒ·128K ctxΒ·

Gemma 3 1B

1 year ago

Google Β· 1B Β· Gemma

Google's tiny Gemma for on-device

1GBΒ·32K ctxΒ·

Wan 2.1 T2V 1.3B

1 year ago

Alibaba Β· 1.3B Β· Apache 2.0

Tiny open text-to-video β€” 480p clips on 8GB consumer GPUs

1.2GBΒ·4K ctxΒ·

Qwen 2.5 Coder 1.5B

1 year ago

Alibaba Β· 1.5B Β· Apache 2.0

Ultra-lightweight coding model

1.3GBΒ·32K ctxΒ·

DeepSeek R1 1.5B

1 year ago

DeepSeek Β· 1.5B Β· MIT

Tiny reasoning model distilled from R1

1.3GBΒ·64K ctxΒ·

Qwen 3 1.7B

1 year ago

Alibaba Β· 1.7B Β· Apache 2.0

Compact multilingual Qwen 3

1.4GBΒ·32K ctxΒ·

Qwen 3.5 2B

6mo ago

Alibaba Β· 2B Β· Apache 2.0

Small multimodal Qwen 3.5

7* AAΒ·1.5GBΒ·32K ctxΒ·

Llama 3.2 3B

2y ago

Meta Β· 3B Β· Llama 3.2 Community

Lightweight Llama for mobile and edge

2GBΒ·128K ctxΒ·

SmolLM3 3B

1 year ago

HuggingFace Β· 3B Β· Apache 2.0

Lightweight multilingual reasoning

2GBΒ·128K ctxΒ·

Granite 4.1 3B

4mo ago

IBM Β· 3B Β· Apache 2.0

Compact enterprise model for edge and constrained environments

2GBΒ·128K ctxΒ·

Ministral 3 3B

8mo ago

Mistral AI Β· 3B Β· Apache 2.0

Current-gen tiny Ministral β€” edge chat with 256K context

7 AAΒ·2GBΒ·256K ctxΒ·

Phi-4 Mini Reasoning

1 year ago

Microsoft Β· 3.8B Β· MIT

Lightweight reasoning model

2.4GBΒ·16K ctxΒ·

Gemma 3 4B

1 year ago

Google Β· 4B Β· Gemma

Multimodal Gemma with 128K context

2.5GBΒ·128K ctxΒ·

Qwen 3.5 4B

6mo ago

Alibaba Β· 4B Β· Apache 2.0

Small multimodal Qwen 3.5

20* AAΒ·2.5GBΒ·32K ctxΒ·

Qwen3-VL 4B

11mo ago

Alibaba Β· 4.4B Β· Apache 2.0

Compact dedicated vision-language model β€” OCR & image chat on edge

2.8GBΒ·256K ctxΒ·

Gemma 4 E2B IT

4mo ago

Google Β· 5B Β· Gemma

Gemma 4 efficient instruct model (official)

10* AAΒ·3.1GBΒ·256K ctxΒ·

Qwen 2.5 Coder 7B

1 year ago

Alibaba Β· 7B Β· Apache 2.0

Dedicated coding model

4.1GBΒ·128K ctxΒ·

DeepSeek R1 Distill 7B

1 year ago

DeepSeek Β· 7B Β· MIT

R1 reasoning distilled into Qwen 7B

4.1GBΒ·64K ctxΒ·

Gemma 4 E4B IT

4mo ago

Google Β· 8B Β· Gemma

Gemma 4 balanced instruct model (official)

12* AAΒ·4.6GBΒ·256K ctxΒ·

Llama 3.1 8B

2y ago

Meta Β· 8B Β· Llama 3.1 Community

Meta's versatile 8B β€” great quality/speed ratio

4.6GBΒ·128K ctxΒ·

Qwen 3 8B

1 year ago

Alibaba Β· 8B Β· Apache 2.0

Qwen 3 with thinking mode support

4.6GBΒ·128K ctxΒ·

Granite 4.1 8B

4mo ago

IBM Β· 8B Β· Apache 2.0

Balanced general-purpose enterprise model

4.6GBΒ·128K ctxΒ·

Ministral 8B

1 year ago

Mistral AI Β· 8B Β· MRL

Mistral's efficient 8B model

4.6GBΒ·32K ctxΒ·

GLM-4 9B

2y ago

Zhipu AI Β· 9B Β· GLM-4

Multilingual model supporting 26 languages with 128K context

5.1GBΒ·128K ctxΒ·

Nemotron Nano 9B v2

1 year ago

NVIDIA Β· 9B Β· NVIDIA Open

Hybrid Mamba2 architecture for reasoning

9* AAΒ·5.1GBΒ·128K ctxΒ·

Ornith 1.0 9B

2mo ago

DeepReinforce Β· 9B Β· MIT

Self-improving agentic coding model optimized for terminal and software engineering tasks

5.1GBΒ·256K ctxΒ·

FLUX.2 Klein 9B

7mo ago

Black Forest Labs Β· 9B Β· FLUX Non-Commercial

Higher-quality distilled FLUX.2 β€” sub-second generation and multi-reference editing

5.1GBΒ·32K ctxΒ·

Gemma 3 12B

1 year ago

Google Β· 12B Β· Gemma

Multimodal Gemma with 128K context

6.6GBΒ·128K ctxΒ·

Mistral Nemo 12B

2y ago

Mistral AI Β· 12B Β· Apache 2.0

Multilingual 12B with 128K context

6.6GBΒ·128K ctxΒ·

Gemma 4 12B IT

4mo ago

Google Β· 12B Β· Apache 2.0

Gemma 4 mid-size instruct β€” multimodal any-to-any

22* AAΒ·6.6GBΒ·256K ctxΒ·

Phi-4 14B

1 year ago

Microsoft Β· 14B Β· MIT

Microsoft's reasoning-focused model

5* AAΒ·7.7GBΒ·16K ctxΒ·

Qwen 3 14B

1 year ago

Alibaba Β· 14B Β· Apache 2.0

Strong all-rounder with thinking mode

7.7GBΒ·128K ctxΒ·

DeepSeek R1 Distill 14B

1 year ago

DeepSeek Β· 14B Β· MIT

R1 reasoning distilled into Qwen 14B

7.7GBΒ·64K ctxΒ·

Ministral 3 14B

8mo ago

Mistral AI Β· 14B Β· Apache 2.0

Current-gen Ministral mid-size β€” local assistant with 256K context

11 AAΒ·7.7GBΒ·256K ctxΒ·

LFM2 24B

9mo ago

Liquid AI Β· 24B Β· Liquid AI

Hybrid MoE with convolution+attention layers β€” 2.3B active

5* AAΒ·12.8GBΒ·32K ctxΒ·

Devstral Small 2 24B

8mo ago

Mistral AI Β· 24B Β· Apache 2.0

Coding-focused model with 256K context β€” 68% SWE-bench

12.8GBΒ·256K ctxΒ·

Mistral Small 3.1 24B

1 year ago

Mistral AI Β· 24B Β· Apache 2.0

Multimodal Mistral with vision support

12.8GBΒ·128K ctxΒ·

DiffusionGemma 26B-A4B IT

2mo ago

Google Β· 26B Β· Apache 2.0

Discrete diffusion MoE β€” 1100+ tok/s on H100, multimodal (text/image/video)

13* AAΒ·13.8GBΒ·256K ctxΒ·

Qwen 3.5 27B

6mo ago

Alibaba Β· 27.8B Β· Apache 2.0

Flagship native multimodal Qwen 3.5

14.7GBΒ·256K ctxΒ·

Qwen 3.6 27B

4mo ago

Alibaba Β· 27.8B Β· Apache 2.0

Flagship dense Qwen 3.6 β€” native multimodal all-rounder

38 AAΒ·14.7GBΒ·256K ctxΒ·

Qwen 3 30B-A3B

1 year ago

Alibaba Β· 30B Β· Apache 2.0

MoE with only 3.3B active β€” extremely efficient

15.9GBΒ·128K ctxΒ·

Nemotron 3 Nano 30B

1 year ago

NVIDIA Β· 30B Β· NVIDIA Open

MoE with 1M context and 3B active

15 AAΒ·15.9GBΒ·1024K ctxΒ·

Granite 4.1 30B

4mo ago

IBM Β· 30B Β· Apache 2.0

High-capacity enterprise model for complex reasoning and tool use

15.9GBΒ·128K ctxΒ·

North Mini Code

2mo ago

Cohere Β· 30B Β· Apache 2.0

Open agentic coding MoE with 3B active β€” built for software engineering and terminal tasks

15.9GBΒ·256K ctxΒ·

Qwen 3 Coder 30B-A3B

1 year ago

Alibaba Β· 30B Β· Apache 2.0

Efficient agentic coding MoE β€” 3B active, 256K context

15.9GBΒ·256K ctxΒ·

Qwen 3 32B

1 year ago

Alibaba Β· 32B Β· Apache 2.0

Qwen 3 flagship dense model

16.9GBΒ·128K ctxΒ·

DeepSeek R1 Distill 32B

1 year ago

DeepSeek Β· 32B Β· MIT

R1 reasoning distilled into Qwen 32B β€” sweet spot

16.9GBΒ·64K ctxΒ·

OLMo 2 32B

1 year ago

Allen AI Β· 32B Β· Apache 2.0

Fully open research model by Allen AI

16.9GBΒ·4K ctxΒ·

Gemma 4 31B IT

4mo ago

Google Β· 33B Β· Gemma

Gemma 4 flagship instruct model (official)

30 AAΒ·17.4GBΒ·256K ctxΒ·

Gemma 4 31B

4mo ago

Google Β· 33B Β· Gemma

Gemma 4 flagship base model (official)

17.4GBΒ·256K ctxΒ·

Command R 35B

2y ago

Cohere Β· 35B Β· CC BY-NC 4.0

Optimized for retrieval-augmented generation

18.4GBΒ·128K ctxΒ·

Qwen 3.5 35B-A3B

6mo ago

Alibaba Β· 35B Β· Apache 2.0

Efficient multimodal MoE with 3B active

18.4GBΒ·256K ctxΒ·

Mixtral 8x7B

2y ago

Mistral AI Β· 47B Β· Apache 2.0

MoE with 12.9B active params

24.6GBΒ·32K ctxΒ·

Llama 3.3 70B

1 year ago

Meta Β· 70B Β· Llama 3.3 Community

Best open model at 70B class

9* AAΒ·36.4GBΒ·128K ctxΒ·

Qwen 3 Next 80B-A3B

8mo ago

Alibaba Β· 80B Β· Apache 2.0

High-sparsity MoE β€” extreme low activation ratio for fast inference at 80B scale

41.5GBΒ·256K ctxΒ·

Qwen 3 Coder Next 80B-A3B

6mo ago

Alibaba Β· 80B Β· Apache 2.0

Ultra-efficient agentic coding MoE optimized for tool-calling coding agents

41.5GBΒ·256K ctxΒ·

HunyuanImage 3.0

1 year ago

Tencent Β· 80B Β· Tencent Hunyuan Community

Largest open image MoE β€” 13B active, strong long-prompt generation

41.5GBΒ·4K ctxΒ·

GLM-4.5 Air

1 year ago

Z.ai Β· 106B Β· MIT

Consumer-friendly GLM MoE β€” 12B active, strong agentic & tool use

54.8GBΒ·128K ctxΒ·

Qwen 3.5 122B-A10B

6mo ago

Alibaba Β· 122B Β· Apache 2.0

Large multimodal MoE

33 AAΒ·63GBΒ·256K ctxΒ·

DeepSeek V4 Flash

4mo ago

DeepSeek Β· 158B Β· MIT

Efficient long-context V4 β€” 13B active, 1M context

81.4GBΒ·1024K ctxΒ·

Qwen 3 235B-A22B

1 year ago

Alibaba Β· 235B Β· Apache 2.0

Massive MoE with 22B active β€” frontier quality

120.9GBΒ·128K ctxΒ·

GLM-4.6

1 year ago

Z.ai Β· 357B Β· MIT

Large GLM MoE with strong coding and 200K context

183.4GBΒ·195K ctxΒ·

Qwen 3.5 397B-A17B

6mo ago

Alibaba Β· 397B Β· Apache 2.0

Largest multimodal Qwen 3.5 MoE

34 AAΒ·203.9GBΒ·256K ctxΒ·

Llama 4 Maverick 17B-128E

1 year ago

Meta Β· 400B Β· Llama 4 Community

Multimodal MoE with 128 experts β€” 17B active, 1M context

14 AAΒ·205.4GBΒ·1024K ctxΒ·

Qwen 3 Coder 480B

1 year ago

Alibaba Β· 480B Β· Apache 2.0

Largest open coding MoE β€” 35B active

246.4GBΒ·256K ctxΒ·

DeepSeek R1

1 year ago

DeepSeek Β· 671B Β· MIT

Massive MoE reasoning model β€” 37B active

344.2GBΒ·64K ctxΒ·

DeepSeek V3.2

8mo ago

DeepSeek Β· 685B Β· MIT

State-of-the-art MoE β€” 37B active params

351.4GBΒ·128K ctxΒ·

GLM-5

6mo ago

Zhipu AI Β· 744B Β· MIT

MoE with 256 experts, 40B active β€” frontier-class agentic coding

381.6GBΒ·128K ctxΒ·

GLM-5.2

2mo ago

Z.ai Β· 753B Β· MIT

Frontier open-weight coder β€” top SWE-bench, 1M context

53 AAΒ·386.2GBΒ·1024K ctxΒ·

GLM-5.1

4mo ago

Zhipu AI Β· 754B Β· MIT

Improved agentic coding β€” SOTA SWE-bench Pro, long-horizon tasks

386.7GBΒ·128K ctxΒ·

Kimi K2.6

4mo ago

Moonshot AI Β· 1.06T Β· Kimi

Natively multimodal 1T MoE β€” 32B active, frontier agentic

542.4GBΒ·256K ctxΒ·

Common questions

Can I generate video locally?
Yes, with open models such as Wan, LTX and Hunyuan Video. They need more VRAM and are slower than chat models, but they run offline once downloaded.
How much VRAM for local video AI?
The smallest video models can start around 8–12 GB. Comfortable 720p-class generation usually wants 16–24 GB. Longer clips and audio-native models go higher.
What is the easiest local video model to try?
Pick the lightest Wan or LTX checkpoint on this page that grades S–B on your device. That is the fastest way to see if local video is usable on your GPU.

All product names, logos, and brands are property of their respective owners. Apple, NVIDIA, AMD, Intel, Qualcomm, and all AI model names mentioned on this site are trademarks or registered trademarks of their respective holders. This site is not affiliated with or endorsed by any of these companies.