Skip to content
CSuite
Audio · Speech-to-text · Released Jul 2026

Fish Audio Transcribe 1

Fish Audio's speech-to-text model, which detects the spoken language automatically; timestamps come back on the direct route.

Fish Audio Transcribe 1 is a cloud speech-to-text model built by Fish Audio. It transcribes spoken audio into written text. It handles multiple spoken languages and can return timestamped segments along with the transcript. It runs through OpenRouter and Fish Audio using your own API key, from about $0.006 per minute of audio.

Modality
Audio
Model ID
fishaudio/transcribe-1
Specs
Released
2 months ago (Jul 2026)
Pricing
$0.006 per minute of audio
Type
Speech-to-text
Languages
Multilingual
Timestamps
Timed segments
About the creator

Fish Audio

Fish Audio builds expressive speech models with inline emotion cues, multi-speaker synthesis, and zero-shot voice cloning across 80+ languages.

fish.audio ↗
More from Fish Audio

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app