Google Gemini 3.5 Transcribe
Google's speech-to-text model for fast synchronous transcription with automatic language detection; this route returns plain text.
Gemini 3.5 Transcribe is a cloud speech-to-text model built by Google. It transcribes spoken audio into written text. It handles multiple spoken languages. It runs through OpenRouter using your own API key, from about $0.005 per minute of audio.
- Released
- Sep 2026
- Pricing
- $0.005 per minute of audio
- Type
- Speech-to-text
- Languages
- Multilingual
Google and Google DeepMind build the Gemini family of multimodal models, the Imagen and Nano Banana image models, the Lyria music models, and the Veo video models.
deepmind.google ↗Google's most capable Flash model, with clear gains over 3.7 Flash in software engineering, agentic tasks, and multi-step reasoning, plus vision and a 1M-token context window.
Google's newest Flash workhorse, a more capable successor to 3.6 Flash with stronger software engineering, better document comprehension and more disciplined tool use, at half the price per token.
Google's most efficient multimodal Flash model: an updated reasoning stack with stronger coding and computer-use quality, at the same speed and scale as the rest of the Flash line.
Google's most capable multimodal model with deep reasoning and 1M token context.
Ultra-fast image generation model optimised for speed and creative output.
Google's fastest, most cost-efficient Nano Banana model — 1K images in ~4 seconds for high-volume workflows.