Audio · Text-to-speech · Released Sep 2025
Kuaishou Kling TTS
Kuaishou's Kling text-to-speech model, reading up to 1,000 characters aloud in English or Chinese system voices.
Kling TTS is a cloud text-to-speech model from Kuaishou. It converts written text into natural, spoken audio. It offers a choice of 27 voices. It runs through Kling using your own API key, from about $0.007 per generation.
Specs
- Released
- 1 year ago (Sep 2025)
- Pricing
- $0.007 per generation
- Type
- Text-to-speech
- Voices
- 27 to choose from
About the creator
Kuaishou
Kuaishou builds the Kling family of cinematic video models — high-resolution clips with native audio, lip sync, and multi-shot storytelling.
klingai.com ↗More from Kuaishou
Kling Image 3.0
ImageCloud
Kuaishou's Kling Image 3.0 model for consistent text-to-image and single-reference image editing at 1K or 2K.
Kling Image 3.0 Omni
ImageCloud
Kuaishou's multi-reference Kling Image 3.0 Omni model with native 2K and 4K output from text and up to ten reference images.
Kling Text to Audio
AudioCloud
Kuaishou's Kling sound effects model, generating 3 to 10 second clips from a short text description.
Kling 3.0
VideoCloud
Kuaishou's cinematic video model generating clips up to 1080p with native audio, lip sync, and sound effects from text or images.
Kling 3.0 Turbo
VideoCloud
Kuaishou's faster Kling 3.0 tier, generating 3 to 15 second clips up to 1080p with native audio from text or a first frame.
Kling 3.0 Omni
VideoCloud
Kuaishou's all-in-one Kling 3.0 model that generates or edits clips up to 4K from text, frames, reference images and a source video.