Providers · Vendor API
Fish Audio
Fish Audio speech and transcription through your own Fish Audio API key and its REST API, with models such as S2.1 Pro for text-to-speech from the Fish voice library or a cloned sample, a free S2.1 Pro tier for testing, and Transcribe 1 Pro for transcripts with speaker labels. Requests go straight from your machine to Fish Audio and you pay Fish Audio's own rates, billed per byte of text for speech and per second of audio for transcription.
Audio(6 models)
Transcribe 1 Pro
Fish Audio's speech-to-text model for multi-speaker conversations, which labels who is speaking and keeps emotion and vocal-event cues such as laughter in the transcript.
Speech-to-text · this month (Sep 2026)
S2.1 Pro Free
Fish Audio's S2.1 Pro model at $0 for development and testing under fair-use limits, with the same bracket cues, multi-speaker synthesis and voice cloning but no latency or data-processing guarantees.
Text-to-speech · 2 months ago (Jul 2026)
Fish Audio Transcribe 1
Fish Audio's speech-to-text model, which detects the spoken language automatically; timestamps come back on the direct route.
Speech-to-text · 2 months ago (Jul 2026)
Fish Audio S2.1 Pro
Fish Audio's recommended production TTS with natural-language cues in brackets, multi-speaker synthesis in one call, and zero-shot voice cloning across 83 languages.
Text-to-speech · 2 months ago (Jul 2026)
S2 Pro
Fish Audio's previous-generation open S2 TTS with natural-language cues in brackets, multi-speaker dialogue and zero-shot voice cloning across 80+ languages.
Text-to-speech · 4 months ago (May 2026)
S1
Fish Audio's 4B-parameter S1 TTS with 64 emotion, tone and effect cues in parentheses and zero-shot voice cloning across 13 languages.
Text-to-speech · 1 year ago (Jun 2025)