Microsoft MAI-Voice-2.1 Flash
Microsoft AI's low-latency text-to-speech model, built for real-time responsiveness. It reads in 23 languages through the same preset voices as MAI-Voice-2.1, across 28 locales, and the voice you pick sets the language.
MAI-Voice-2.1 Flash is a cloud text-to-speech model from Microsoft. It converts written text into natural, spoken audio. It offers a choice of 97 voices. Voices can be tuned for speed. It runs through OpenRouter using your own API key, from $0.015 per 1,000 characters.
- Released
- Oct 2026
- Pricing
- $0.015 / 1k chars
- Type
- Text-to-speech
- Voices
- 97 to choose from
- Voice controls
- Speed
Examples
Generated with MAI-Voice-2.1 Flash via OpenRouter: the same three scripts read by every text-to-speech model in the catalog, so the only thing that changes between two models’ clips is the model. Each model uses its own default voice; there is no shared voice to hold constant.
Baseline naturalness and pacing
The last train had already left, but she decided to walk anyway. The city was quieter than she remembered, and for the first time in weeks, she wasn't in a hurry.
Emotion, emphasis, and pauses
Wait — you're telling me it actually worked? After all that? I can't believe it. Honestly, I thought we'd lost the whole thing.
Decimals, percentages, version numbers
The API returned 3,481 results in 0.42 seconds, a 12 percent improvement over version 2.5. Latency at the 99th percentile dropped from 840 milliseconds to 610.
Microsoft
Microsoft builds the open-weight Phi family of small language models for reasoning, math and code. Microsoft AI also makes MAI-Transcribe, a speech-to-text model that covers 60 languages.
azure.microsoft.com/products/phi ↗Microsoft AI's highest-fidelity text-to-speech model, with detailed prosody and a consistent speaker over long passages. It reads in 23 languages through preset voices across 28 locales, and the voice you pick sets the language.
Microsoft AI's multilingual transcription model, which covers 60 languages with automatic language detection and returns plain text.
Microsoft's 14B state-of-the-art open model with strong reasoning, math, and code.
Microsoft's compact 3.8B model with a 128K context and native function calling.
Microsoft's 14B reasoning model fine-tuned for chain-of-thought on math, science, and code.