Audio · Speech-to-text · Released Jan 2026
ElevenLabs Scribe v2
ElevenLabs' batch transcription model, which returns word-level timing for subtitles and can label who is speaking.
From ElevenLabs, Scribe v2 is a cloud speech-to-text model. It transcribes spoken audio into written text. It handles multiple spoken languages and can return timestamped segments along with the transcript. It also supports speaker labels. It runs through ElevenLabs using your own API key, from about $0.0037 per minute of audio.
Specs
- Released
- 8 months ago (Jan 2026)
- Pricing
- $0.0037 per minute of audio
- Type
- Speech-to-text
- Languages
- Multilingual
- Timestamps
- Timed segments
- Speaker labels
- Supported
About the creator
ElevenLabs
ElevenLabs builds best-in-class voice AI — realistic text-to-speech, voice cloning, music, and sound effects.
elevenlabs.io ↗More from ElevenLabs
v3
AudioCloud
ElevenLabs' most expressive voice model with rich emotional and tonal range.
Music
AudioCloud
ElevenLabs music generation model for creating original AI-composed tracks.
Flash v2.5
AudioCloud
ElevenLabs' fastest speech model, with about 75 ms latency across 32 languages at half the per-character price of Multilingual v2.
Sound Effects v2
AudioCloud
ElevenLabs' text-to-sound-effects model: foley, ambiences and impacts up to 30 seconds, with seamless looping for background beds.