Audio · Text-to-speech · Released Jun 2025
Fish Audio S1
Fish Audio's 4B-parameter S1 TTS with 64 emotion, tone and effect cues in parentheses and zero-shot voice cloning across 13 languages.
From Fish Audio, S1 is a cloud text-to-speech model. It converts written text into natural, spoken audio. Voices can be tuned for speed, and output is available as MP3, WAV, and OPUS. It runs through Fish Audio using your own API key, from $0.015 per 1,000 characters.
Specs
- Released
- 1 year ago (Jun 2025)
- Pricing
- $0.015 / 1k chars
- Type
- Text-to-speech
- Voice controls
- Speed
- Output formats
- MP3, WAV, OPUS
About the creator
Fish Audio
Fish Audio builds expressive speech models with inline emotion cues, multi-speaker synthesis, and zero-shot voice cloning across 80+ languages.
fish.audio ↗More from Fish Audio
Fish Audio S2.1 Pro
AudioCloud
Fish Audio's recommended production TTS with natural-language cues in brackets, multi-speaker synthesis in one call, and zero-shot voice cloning across 83 languages.
S2.1 Pro Free
AudioCloud
Fish Audio's S2.1 Pro model at $0 for development and testing under fair-use limits, with the same bracket cues, multi-speaker synthesis and voice cloning but no latency or data-processing guarantees.
S2 Pro
AudioCloud
Fish Audio's previous-generation open S2 TTS with natural-language cues in brackets, multi-speaker dialogue and zero-shot voice cloning across 80+ languages.
Fish Audio Transcribe 1
AudioCloud
Fish Audio's speech-to-text model, which detects the spoken language automatically; timestamps come back on the direct route.
Transcribe 1 Pro
AudioCloud
Fish Audio's speech-to-text model for multi-speaker conversations, which labels who is speaking and keeps emotion and vocal-event cues such as laughter in the transcript.