Skip to content
CSuite
Audio · Speech-to-text · Released Sep 2026

Google Gemini 3.5 Transcribe

Google's speech-to-text model for fast synchronous transcription with automatic language detection; this route returns plain text.

Gemini 3.5 Transcribe is a cloud speech-to-text model built by Google. It transcribes spoken audio into written text. It handles multiple spoken languages. It runs through OpenRouter using your own API key, from about $0.005 per minute of audio.

Modality
Audio
Available on
Model ID
google/gemini-3.5-transcribe
Specs
Released
Sep 2026
Pricing
$0.005 per minute of audio
Type
Speech-to-text
Languages
Multilingual
About the creator

Google

Google and Google DeepMind build the Gemini family of multimodal models, the Imagen and Nano Banana image models, the Lyria music models, and the Veo video models.

deepmind.google ↗
More from Google

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app