Skip to content
CSuite
Audio · Text-to-speech · Released Sep 2026

ElevenLabs v4

ElevenLabs' most expressive speech model, with closer voice cloning, audio tags for delivery and more than 90 languages.

v4 is a cloud text-to-speech model built by ElevenLabs. It converts written text into natural, spoken audio. It supports multiple languages. Voices can be tuned for stability and similarity, and output is available as MP3 and WAV_44100. It runs through ElevenLabs using your own API key, from $0.08 per 1,000 characters.

Modality
Audio
Available on
Model ID
elevenlabs/v4
Specs
Released
this month (Sep 2026)
Pricing
$0.08 / 1k chars
Type
Text-to-speech
Languages
Multilingual
Voice controls
Stability, Similarity
Output formats
MP3, WAV_44100
Samples

Examples

Generated with v4 via ElevenLabs: the same three scripts read by every text-to-speech model in the catalog, so the only thing that changes between two models’ clips is the model. Each model uses its own default voice; there is no shared voice to hold constant.

Narration

Baseline naturalness and pacing

The last train had already left, but she decided to walk anyway. The city was quieter than she remembered, and for the first time in weeks, she wasn't in a hurry.

0:00 / 0:10
Expressive

Emotion, emphasis, and pauses

Wait — you're telling me it actually worked? After all that? I can't believe it. Honestly, I thought we'd lost the whole thing.

0:00 / 0:08
Numbers and jargon

Decimals, percentages, version numbers

The API returned 3,481 results in 0.42 seconds, a 12 percent improvement over version 2.5. Latency at the 99th percentile dropped from 840 milliseconds to 610.

0:00 / 0:16
About the creator

ElevenLabs

ElevenLabs builds best-in-class voice AI — realistic text-to-speech, voice cloning, music, and sound effects.

elevenlabs.io ↗
More from ElevenLabs
From the blog

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app