Audio · Text-to-speech · Released May 2026
Fish Audio S2 Pro
Fish Audio's previous-generation open S2 TTS with natural-language cues in brackets, multi-speaker dialogue and zero-shot voice cloning across 80+ languages.
S2 Pro is a cloud text-to-speech model from Fish Audio. It converts written text into natural, spoken audio. Voices can be tuned for speed, and output is available as MP3, WAV, and OPUS. It runs through Fish Audio using your own API key, from $0.015 per 1,000 characters.
Specs
- Released
- 4 months ago (May 2026)
- Pricing
- $0.015 / 1k chars
- Type
- Text-to-speech
- Voice controls
- Speed
- Output formats
- MP3, WAV, OPUS
About the creator
Fish Audio
Fish Audio builds expressive speech models with inline emotion cues, multi-speaker synthesis, and zero-shot voice cloning across 80+ languages.
fish.audio ↗More from Fish Audio
Fish Audio S2.1 Pro
AudioCloud
Fish Audio's recommended production TTS with natural-language cues in brackets, multi-speaker synthesis in one call, and zero-shot voice cloning across 83 languages.
S2.1 Pro Free
AudioCloud
Fish Audio's S2.1 Pro model at $0 for development and testing under fair-use limits, with the same bracket cues, multi-speaker synthesis and voice cloning but no latency or data-processing guarantees.
S1
AudioCloud
Fish Audio's 4B-parameter S1 TTS with 64 emotion, tone and effect cues in parentheses and zero-shot voice cloning across 13 languages.
Fish Audio Transcribe 1
AudioCloud
Fish Audio's speech-to-text model, which detects the spoken language automatically; timestamps come back on the direct route.
Transcribe 1 Pro
AudioCloud
Fish Audio's speech-to-text model for multi-speaker conversations, which labels who is speaking and keeps emotion and vocal-event cues such as laughter in the transcript.