Audio · Speech-to-text · Released Jul 2026
Fish Audio Transcribe 1
Fish Audio's speech-to-text model, which detects the spoken language automatically; timestamps come back on the direct route.
Fish Audio Transcribe 1 is a cloud speech-to-text model built by Fish Audio. It transcribes spoken audio into written text. It handles multiple spoken languages and can return timestamped segments along with the transcript. It runs through OpenRouter and Fish Audio using your own API key, from about $0.006 per minute of audio.
Specs
- Released
- 2 months ago (Jul 2026)
- Pricing
- $0.006 per minute of audio
- Type
- Speech-to-text
- Languages
- Multilingual
- Timestamps
- Timed segments
About the creator
Fish Audio
Fish Audio builds expressive speech models with inline emotion cues, multi-speaker synthesis, and zero-shot voice cloning across 80+ languages.
fish.audio ↗More from Fish Audio
Fish Audio S2.1 Pro
AudioCloud
Fish Audio's recommended production TTS with natural-language cues in brackets, multi-speaker synthesis in one call, and zero-shot voice cloning across 83 languages.
S2.1 Pro Free
AudioCloud
Fish Audio's S2.1 Pro model at $0 for development and testing under fair-use limits, with the same bracket cues, multi-speaker synthesis and voice cloning but no latency or data-processing guarantees.
S2 Pro
AudioCloud
Fish Audio's previous-generation open S2 TTS with natural-language cues in brackets, multi-speaker dialogue and zero-shot voice cloning across 80+ languages.
S1
AudioCloud
Fish Audio's 4B-parameter S1 TTS with 64 emotion, tone and effect cues in parentheses and zero-shot voice cloning across 13 languages.
Transcribe 1 Pro
AudioCloud
Fish Audio's speech-to-text model for multi-speaker conversations, which labels who is speaking and keeps emotion and vocal-event cues such as laughter in the transcript.