Skip to content
CSuite
Features · Audio

Voice, music, and sound — one panel.

Narrate a script with a natural voice, score it with generated music, and drop in the sound effects — three kinds of audio generation in one workspace, with a built-in waveform editor for the cleanup. Generation runs on your own API keys; editing runs locally on your machine.

Choose

Three kinds of sound, one picker.

Switch the composer between speech, music, and sound effects — the model picker filters to what supports the active type, and your choice is remembered per category. ElevenLabs for a warm narrator, Lyria for a score, SFX 1.5 for foley: one API key per platform unlocks them all.
Speak

A script in, a voiceover out.

Paste the script, pick a model and a voice, and generate. Voices come from each model's own roster, with per-voice controls — speed, stability, similarity — on models that expose them. Hit the play button beside the picker to audition a voice before committing a whole script to it: the first listen costs a fraction of a cent, and every one after that is cached and free. On models that support cloning, hand it a sample and a transcript instead and it reads your script in that voice.The clip below the mockup is a real, unretouched Eleven v3 generation from CSuite — press play. The prompt in the bar is the exact script it read.
Real output · Eleven v3narration.mp3
0:00 / 0:10
Compose

Describe the mood, get the music.

Instrumentation, tempo, texture — describe it like a brief to a session musician and the model composes it. Royalty questions don't follow you around: it's your generation, made with your key.The 30-second clip below is a real Lyria 3 composition from the exact prompt shown — the same brief every music model in the catalog answers on its detail page.
Real output · Lyria 3ambient.mp3
0:00 / 0:30
Foley

Sound effects on demand.

The door slam, the forest ambience, the projector whir — describe the sound instead of digging through sample libraries, and drop the result straight into a video composition or export it for your editor.Below: a real SFX 1.5 generation of the prompt shown — a heavy wooden door in a stone corridor.
Real output · SFX 1.5impact.mp3
0:00 / 0:10
Enhance

Type the idea, get the whole brief.

For music and effects, press the wand and your own text model turns one line into a full brief: what the track is about first, then genre, instrumentation, tempo, and mood. Lyrics land wherever the model reads them. Speech has no wand, because the text you type is what gets spoken.Music is billed per generation, so the takes you discard are the expensive part.
Edit

Clean it up without leaving the app.

A waveform editor built into the player covers the usual cleanup — no export to another tool. Every edit renders locally on your machine and saves in place or as a new copy, keeping the file's own format so a trimmed MP3 stays an MP3.

Trim to the take

Cut the silence off both ends of a recording, or pull one clean take out of a long session — drag the handles on the waveform and everything outside the selection goes.

Fade in, fade out

Give a generated music bed a smooth entrance and exit before it goes under a voiceover — set the fade lengths and the envelope is applied to the waveform.

Fix the level

A quiet voiceover next to a loud music bed is the most common mix problem there is — boost or cut the clip's volume until the levels sit right.

Change the speed

Turn a 42-minute lecture recording into a 28-minute listen at 1.5×, or slow a fast take down — the duration updates with the rate.

Convert the format

Hand your editor a lossless WAV or FLAC, or squeeze a session down to MP3 or OGG for sharing — pick the target and it transcodes straight from the source file, with no lossy round-trip in between.
Workspace

The rest of the booth.

The small things around the generate button — hearing a voice first, knowing the price, getting back to the prompt that worked.

Audition a voicePlay any voice in the roster straight from the picker before you commit a script to it. First listen costs a fraction of a cent; every one after is cached and free.
Clone a voiceOn models that support it, attach a sample and its transcript and the model reads your script in that voice instead of a preset one.
Write the brief with AIFor music and effects, type the gist and hit the wand: your text model expands it into subject, genre, instrumentation, tempo, and mood, writing the lyrics too on models that sing from the prompt. Speech is left alone, since the text is what gets spoken.
Know what it costs firstThe composer shows the estimated price before you generate — per clip for music, from your live character count for speech — and a status line tracks the run.
Reuse what workedRecent prompts sit as chips above the composer, and every generated clip remembers the model and settings behind it, reloadable in one click.
Files that stay yoursClips are ordinary files in your project folder — drag recordings in, filter a long list by name, and open or back them up with any tool you like.
Catalog

Available models & providers.

Every audio model in the catalog — text-to-speech, music, and sound effects — all cloud-hosted through Runware and Replicate with your own keys, each with real sample clips on its detail page.

RunwareCloud · your API key
ReplicateCloud · your API key
Hugging FaceLocal · your hardware
Gemini 3.1 Flash TTS
Cloud
Fast Gemini text-to-speech with natural-sounding expressive voices.
Runware · Replicate
Grok Text-to-Speech
Cloud
xAI's expressive text-to-speech model with multiple voices and 20-language support.
Runware · Replicate
Lyria 3
Cloud
Enhanced Lyria model with improved musical coherence and audio quality.
Replicate
Lyria 3 Pro
Cloud
Google's professional-grade music model with enhanced fidelity and musical expression.
Replicate
Eleven v3
Cloud
ElevenLabs' most expressive voice model with rich emotional and tonal range.
Replicate
Eleven Turbo v2.5
Cloud
Low-latency ElevenLabs TTS for real-time conversational speech synthesis.
Replicate
Speech 2 Turbo
Cloud
MiniMax fast speech synthesis with natural prosody and multi-language support.
Replicate
Eleven Music
Cloud
ElevenLabs music generation model for creating original AI-composed tracks.
Replicate
Music 2.6
Cloud
MiniMax music model for full-arrangement AI compositions with rich instrumentation.
Runware · Replicate
SFX 1.5
Cloud
Mirelo sound effects model for generating custom audio SFX from text prompts.
Runware
Speech 2.8
Cloud
MiniMax's studio-grade speech model with 332 voices, emotion control, voice cloning, and a 50,000-character input limit.
Runware · Replicate
Fish Audio S2.1 Pro
Cloud
Fish Audio's expressive TTS with inline emotion cues, multi-speaker synthesis in one call, and zero-shot voice cloning across 80+ languages.
Runware
ACE-Step v1.5
Cloud
Open full-track music generation with lyric editing, remix, cover generation, and voice cloning in 50+ languages.
Runware
Supertonic TTS
Local
On-device English text-to-speech that runs entirely in the app, with ten preset voices and no API key.
Hugging Face
FAQ

Audio questions, answered.

How speech, music, and sound-effect generation work in CSuite — voices, costs, editing, and formats.

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app