Audio · Sound effects · Released Aug 2025
Kuaishou Kling Text to Audio
Kuaishou's Kling sound effects model, generating 3 to 10 second clips from a short text description.
Kuaishou's Kling Text to Audio is a cloud sound-effects model. It generates sound effects from a short text description. It can generate clips up to 10 seconds long. Output is available as MP3 and WAV. It runs through Kling using your own API key, from about $0.035 per generation.
Specs
- Released
- 1 year ago (Aug 2025)
- Pricing
- $0.035 per generation
- Type
- Sound effects
- Max length
- Up to 10s
- Output formats
- MP3, WAV
About the creator
Kuaishou
Kuaishou builds the Kling family of cinematic video models — high-resolution clips with native audio, lip sync, and multi-shot storytelling.
klingai.com ↗More from Kuaishou
Kling Image 3.0
ImageCloud
Kuaishou's Kling Image 3.0 model for consistent text-to-image and single-reference image editing at 1K or 2K.
Kling Image 3.0 Omni
ImageCloud
Kuaishou's multi-reference Kling Image 3.0 Omni model with native 2K and 4K output from text and up to ten reference images.
Kling TTS
AudioCloud
Kuaishou's Kling text-to-speech model, reading up to 1,000 characters aloud in English or Chinese system voices.
Kling 3.0
VideoCloud
Kuaishou's cinematic video model generating clips up to 1080p with native audio, lip sync, and sound effects from text or images.
Kling 3.0 Turbo
VideoCloud
Kuaishou's faster Kling 3.0 tier, generating 3 to 15 second clips up to 1080p with native audio from text or a first frame.
Kling 3.0 Omni
VideoCloud
Kuaishou's all-in-one Kling 3.0 model that generates or edits clips up to 4K from text, frames, reference images and a source video.