ElevenLabs v4 Turbo
The low-latency tier of ElevenLabs v4, with the same audio tags and more than 90 languages at half the per-character price.
From ElevenLabs, v4 Turbo is a cloud text-to-speech model. It converts written text into natural, spoken audio. It supports multiple languages. Voices can be tuned for stability and similarity, and output is available as MP3 and WAV_44100. It runs through ElevenLabs using your own API key, from $0.04 per 1,000 characters.
- Released
- this month (Sep 2026)
- Pricing
- $0.04 / 1k chars
- Type
- Text-to-speech
- Languages
- Multilingual
- Voice controls
- Stability, Similarity
- Output formats
- MP3, WAV_44100
Examples
Generated with v4 Turbo via ElevenLabs: the same three scripts read by every text-to-speech model in the catalog, so the only thing that changes between two models’ clips is the model. Each model uses its own default voice; there is no shared voice to hold constant.
Baseline naturalness and pacing
The last train had already left, but she decided to walk anyway. The city was quieter than she remembered, and for the first time in weeks, she wasn't in a hurry.
Emotion, emphasis, and pauses
Wait — you're telling me it actually worked? After all that? I can't believe it. Honestly, I thought we'd lost the whole thing.
Decimals, percentages, version numbers
The API returned 3,481 results in 0.42 seconds, a 12 percent improvement over version 2.5. Latency at the 99th percentile dropped from 840 milliseconds to 610.
ElevenLabs
ElevenLabs builds best-in-class voice AI — realistic text-to-speech, voice cloning, music, and sound effects.
elevenlabs.io ↗ElevenLabs' most expressive speech model, with closer voice cloning, audio tags for delivery and more than 90 languages.
ElevenLabs' most expressive voice model with rich emotional and tonal range.
ElevenLabs music generation model for creating original AI-composed tracks.
ElevenLabs' fastest speech model, with about 75 ms latency across 32 languages at half the per-character price of Multilingual v2.
ElevenLabs' text-to-sound-effects model: foley, ambiences and impacts up to 30 seconds, with seamless looping for background beds.
ElevenLabs' batch transcription model, which returns word-level timing for subtitles and can label who is speaking.