ElevenLabs Flash v2.5
ElevenLabs' fastest speech model, with about 75 ms latency across 32 languages at half the per-character price of Multilingual v2.
Flash v2.5 is a cloud text-to-speech model from ElevenLabs. It converts written text into natural, spoken audio. It offers a choice of 26 voices and supports multiple languages. Voices can be tuned for speed, stability, and similarity, and output is available as MP3 and WAV_44100. It runs through Replicate and ElevenLabs using your own API key, from $0.05 per 1,000 characters.
- Released
- 1 year ago (Dec 2024)
- Pricing
- $0.05 / 1k chars
- Type
- Text-to-speech
- Voices
- 26 to choose from
- Languages
- Multilingual
- Voice controls
- Speed, Stability, Similarity
- Output formats
- MP3, WAV_44100
Examples
Generated with Flash v2.5 via Replicate: the same three scripts read by every text-to-speech model in the catalog, so the only thing that changes between two models’ clips is the model. Each model uses its own default voice; there is no shared voice to hold constant.
Baseline naturalness and pacing
The last train had already left, but she decided to walk anyway. The city was quieter than she remembered, and for the first time in weeks, she wasn't in a hurry.
Emotion, emphasis, and pauses
Wait — you're telling me it actually worked? After all that? I can't believe it. Honestly, I thought we'd lost the whole thing.
Decimals, percentages, version numbers
The API returned 3,481 results in 0.42 seconds, a 12 percent improvement over version 2.5. Latency at the 99th percentile dropped from 840 milliseconds to 610.
ElevenLabs
ElevenLabs builds best-in-class voice AI — realistic text-to-speech, voice cloning, music, and sound effects.
elevenlabs.io ↗