MiniMax H3 Fast
MiniMax's fastest H3 tier: 480p clips of 4 to 15 seconds with native audio from text, first and last frames, or reference images, videos and audio.
H3 Fast is a cloud video-generation model from MiniMax. It produces short video clips from a text prompt, a starting image, or by interpolating between a first and last frame. Clips run up to 15s. It also supports reference images. It runs through Runware using your own API key, from about $0.046 per second of output.
- Released
- Sep 2026
- Pricing
- $0.046 per second
- Aspect ratios
- 21:9, 16:9, 4:3 +3 more
- Clip length
- Up to 15s
- Image-to-video
- Start & end frame + reference images
Examples
Generated with H3 Fast via Runware. The same two prompts run against every video model in the catalog (one landscape, one portrait), so the only thing that changes between two models’ clips is the model.
Camera motion, scene coherence over time
Fluid motion, fine detail, temporal stability
MiniMax
MiniMax is a Chinese AI lab whose speech and music models are widely used for multilingual voiceover and generative music.
www.minimax.io ↗MiniMax music model for full-arrangement AI compositions with rich instrumentation.
MiniMax's studio-grade speech model with 332 voices, emotion control, voice cloning, and a 50,000-character input limit.
MiniMax's Hailuo 3 video model with native synchronized audio, multi-reference consistency, and prompt-driven editing of a finished clip.
MiniMax's throughput-tuned H3, built for faster clips with native synchronized audio and first-to-last frame control. It trades H3's multi-reference conditioning and 2K tier for speed.
A faster, lower-cost tier of MiniMax's H3 Max for text-to-video and first-to-last frame clips of 5 to 15 seconds with native audio, at 480P or 768P.
MiniMax's H3 video model running offline, with synchronised audio. The heaviest local entry by a wide margin: its text encoder alone is a 32B vision-language model.