Video · Released Aug 2026
MiniMax H3 Local
MiniMax's H3 video model running offline, with synchronised audio. The heaviest local entry by a wide margin: its text encoder alone is a 32B vision-language model.
MiniMax's H3 Local is an open-weight video-generation model. It produces short video clips from a text prompt. It runs entirely on your own hardware through GGML: no API key and no per-use fees.
Supported OS
macOSWindowsLinux
Minimum machine configuration
Local video (experimental) · 48 GB+ RAM · 40 GB disk · impractically slow without a high-end CUDA GPU
Download size · 33 GB
Specs
- Released
- 1 month ago (Aug 2026)
- Pricing
- Free, runs on your hardware
About the creator
MiniMax
MiniMax is a Chinese AI lab whose speech and music models are widely used for multilingual voiceover and generative music.
www.minimax.io ↗More from MiniMax
Speech 2 Turbo
AudioCloud
MiniMax fast speech synthesis with natural prosody and multi-language support.
Music 2.6
AudioCloud
MiniMax music model for full-arrangement AI compositions with rich instrumentation.
Speech 2.8
AudioCloud
MiniMax's studio-grade speech model with 332 voices, emotion control, voice cloning, and a 50,000-character input limit.
H3
VideoCloud
MiniMax's Hailuo 3 video model with native synchronized audio, multi-reference consistency, and prompt-driven editing of a finished clip.
H3 Max
VideoCloud
MiniMax's throughput-tuned H3, built for faster clips with native synchronized audio and first-to-last frame control. It trades H3's multi-reference conditioning and 2K tier for speed.