Fish Audio S2.1 Pro
Fish Audio's recommended production TTS with natural-language cues in brackets, multi-speaker synthesis in one call, and zero-shot voice cloning across 83 languages.
From Fish Audio, Fish Audio S2.1 Pro is a cloud text-to-speech model. It converts written text into natural, spoken audio. It offers a choice of 8 voices. Voices can be tuned for speed, and output is available as MP3, WAV, FLAC, OGG, and OPUS. It runs through Runware, OpenRouter, and Fish Audio using your own API key, from $0.015 per 1,000 characters.
- Released
- 2 months ago (Jul 2026)
- Pricing
- $0.015 / 1k chars
- Type
- Text-to-speech
- Voices
- 8 to choose from
- Voice controls
- Speed
- Output formats
- MP3, WAV, FLAC, OGG, OPUS
Examples
Generated with Fish Audio S2.1 Pro via Runware: the same three scripts read by every text-to-speech model in the catalog, so the only thing that changes between two models’ clips is the model. Each model uses its own default voice; there is no shared voice to hold constant.
Baseline naturalness and pacing
The last train had already left, but she decided to walk anyway. The city was quieter than she remembered, and for the first time in weeks, she wasn't in a hurry.
Emotion, emphasis, and pauses
Wait — you're telling me it actually worked? After all that? I can't believe it. Honestly, I thought we'd lost the whole thing.
Decimals, percentages, version numbers
The API returned 3,481 results in 0.42 seconds, a 12 percent improvement over version 2.5. Latency at the 99th percentile dropped from 840 milliseconds to 610.
Fish Audio
Fish Audio builds expressive speech models with inline emotion cues, multi-speaker synthesis, and zero-shot voice cloning across 80+ languages.
fish.audio ↗