Skip to content
CSuite
Audio · Text-to-speech · Released Jun 2025

Fish Audio S1

Fish Audio's 4B-parameter S1 TTS with 64 emotion, tone and effect cues in parentheses and zero-shot voice cloning across 13 languages.

From Fish Audio, S1 is a cloud text-to-speech model. It converts written text into natural, spoken audio. Voices can be tuned for speed, and output is available as MP3, WAV, and OPUS. It runs through Fish Audio using your own API key, from $0.015 per 1,000 characters.

Modality
Audio
Available on
Model ID
fishaudio/s1
Specs
Released
1 year ago (Jun 2025)
Pricing
$0.015 / 1k chars
Type
Text-to-speech
Voice controls
Speed
Output formats
MP3, WAV, OPUS
About the creator

Fish Audio

Fish Audio builds expressive speech models with inline emotion cues, multi-speaker synthesis, and zero-shot voice cloning across 80+ languages.

fish.audio ↗
More from Fish Audio

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app