Alibaba Wan 3.0 Prime
Alibaba's lower-latency Wan 3.0 tier, with the same multimodal workflows, 30-second ceiling, and native audio as the base model.
Wan 3.0 Prime is a cloud video-generation model from Alibaba. It produces short video clips from a text prompt, a starting image, or by interpolating between a first and last frame. Clips at 1080p with a synchronized audio track. It also supports reference images. It runs through Runware, Replicate, and OpenRouter using your own API key, from about $0.14 per second of output.
- Released
- Jul 2026
- Pricing
- $0.14 per second
- Max resolution
- 1080p
- Aspect ratios
- 16:9, 9:16, 1:1 +2 more
- Audio
- Synchronized track
- Image-to-video
- Start & end frame + reference images
Examples
Generated with Wan 3.0 Prime via Runware. The same two prompts run against every video model in the catalog (one landscape, one portrait), so the only thing that changes between two models’ clips is the model.
Camera motion, scene coherence over time
Fluid motion, fine detail, temporal stability
Alibaba
Alibaba's Tongyi research group publishes the Wan video models and the Qwen family of language models.
www.alibabacloud.com ↗Alibaba's flagship Qwen3.8 reasoning model for coding, agentic workflows and document analysis, with image input and a 1M-token context window.
A higher-throughput variant of Alibaba's Qwen3.8 Max, served as a separate tier at a higher price.
Alibaba's fast, low-cost multimodal Qwen3.8 reasoning model for coding assistance, agentic workflows and visual understanding, with a 1M-token context window.
Alibaba's omni-modal Qwen3.8 reasoning model built around agentic work, which reads text, images, audio and video.
High-quality diffusion model for detailed, prompt-faithful image generation.
Alibaba's professional-grade Wan 2.7 image model for higher-fidelity, prompt-faithful generation and editing.