Skip to content
CSuite
Google · Audio

Gemini 3.1 Flash TTS

Fast Gemini text-to-speech with natural-sounding expressive voices.

Google's Gemini 3.1 Flash TTS is a cloud text-to-speech model. It converts written text into natural, spoken audio. It offers a choice of 30 voices and supports multiple languages. Output is available as MP3, WAV, and FLAC. It runs through Replicate and Runware using your own API key, from $1.00 per million input tokens.

Modality
Audio
Available on
ReplicateRunware
Model ID
google/gemini-3.1-flash-tts
Specs
Pricing
$1.00 / 1M in · $20.00 / 1M out
Type
Text-to-speech
Voices
30 to choose from
Languages
Multilingual
Output formats
MP3, WAV, FLAC
Samples

Examples

Generated with Gemini 3.1 Flash TTS via Runware the same three scripts read by every text-to-speech model in the catalog, so the only thing that changes between two models’ clips is the model. Each model uses its own default voice; there is no shared voice to hold constant.

Narration

Baseline naturalness and pacing

The last train had already left, but she decided to walk anyway. The city was quieter than she remembered, and for the first time in weeks, she wasn't in a hurry.

0:00 / 0:11
Expressive

Emotion, emphasis, and pauses

Wait — you're telling me it actually worked? After all that? I can't believe it. Honestly, I thought we'd lost the whole thing.

0:00 / 0:11
Numbers and jargon

Decimals, percentages, version numbers

The API returned 3,481 results in 0.42 seconds, a 12 percent improvement over version 2.5. Latency at the 99th percentile dropped from 840 milliseconds to 610.

0:00 / 0:16
About the creator

Google

Google and Google DeepMind build the Gemini family of multimodal models, the Imagen and Nano Banana image models, the Lyria music models, and the Veo video models.

deepmind.google
More from Google
From the blog

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app