Skip to content
CSuite
Audio · Speech-to-text · Released Jan 2026

ElevenLabs Scribe v2

ElevenLabs' batch transcription model, which returns word-level timing for subtitles and can label who is speaking.

From ElevenLabs, Scribe v2 is a cloud speech-to-text model. It transcribes spoken audio into written text. It handles multiple spoken languages and can return timestamped segments along with the transcript. It also supports speaker labels. It runs through ElevenLabs using your own API key, from about $0.0037 per minute of audio.

Modality
Audio
Available on
Model ID
elevenlabs/scribe-v2
Specs
Released
8 months ago (Jan 2026)
Pricing
$0.0037 per minute of audio
Type
Speech-to-text
Languages
Multilingual
Timestamps
Timed segments
Speaker labels
Supported
About the creator

ElevenLabs

ElevenLabs builds best-in-class voice AI — realistic text-to-speech, voice cloning, music, and sound effects.

elevenlabs.io ↗
More from ElevenLabs

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app