ByteDance Seed Audio 1.0
ByteDance's audio generation model that reads a script word for word or builds a whole scene from one prompt, with a described voice, ambience, music and sound effects, and clones a voice from short reference clips.
Seed Audio 1.0 is a cloud text-to-speech model from ByteDance. It converts written text into natural, spoken audio. Voices can be tuned for speed, and output is available as MP3, WAV, FLAC, and OGG. It runs through Runware and OpenRouter using your own API key, from about $0.0025 per second of output.
- Released
- Jun 2026
- Pricing
- $0.0025 per second
- Type
- Text-to-speech
- Voice controls
- Speed
- Output formats
- MP3, WAV, FLAC, OGG
Examples
Generated with Seed Audio 1.0 via Runware: the same three scripts read by every text-to-speech model in the catalog, so the only thing that changes between two models’ clips is the model. Each model uses its own default voice; there is no shared voice to hold constant.
Baseline naturalness and pacing
The last train had already left, but she decided to walk anyway. The city was quieter than she remembered, and for the first time in weeks, she wasn't in a hurry.
Emotion, emphasis, and pauses
Wait — you're telling me it actually worked? After all that? I can't believe it. Honestly, I thought we'd lost the whole thing.
Decimals, percentages, version numbers
The API returned 3,481 results in 0.42 seconds, a 12 percent improvement over version 2.5. Latency at the 99th percentile dropped from 840 milliseconds to 610.
ByteDance
ByteDance's AI labs ship the Seedream image models and the Seedance video models, among other Doubao-branded research.
www.bytedance.com ↗ByteDance's fast Seed 2.1 reasoning model with a 256K context and output window, native tool calling and image input.
ByteDance's flagship Seed 2.0 reasoning model for complex analysis, coding and agentic tool use, with image input and a 256K context.
Balanced Seed 2.0 reasoning model with tool calling and image input at a lower price than Pro.
The smallest, cheapest Seed 2.0 model for fast everyday tasks, with tool calling and image input.
Preview of ByteDance's coding-tuned Seed 2.0 model for code generation and agentic programming, with tool calling.
Lightweight Seedream variant for fast, high-quality image generation.