Narrate a script with a natural voice, score it with generated music, and drop in the sound effects — three kinds of audio generation in one workspace, with a built-in waveform editor for the cleanup. Generation runs on your own API keys; editing runs locally on your machine.
Switch the composer between speech, music, and sound effects — the model picker filters to what supports the active type, and your choice is remembered per category. ElevenLabs for a warm narrator, Lyria for a score, SFX 1.5 for foley: one API key per platform unlocks them all.
Paste the script, pick a model and a voice, and generate. Voices come from each model's own roster, with per-voice controls — speed, stability, similarity — on models that expose them. Hit the play button beside the picker to audition a voice before committing a whole script to it: the first listen costs a fraction of a cent, and every one after that is cached and free. On models that support cloning, hand it a sample and a transcript instead and it reads your script in that voice.The clip below the mockup is a real, unretouched Eleven v3 generation from CSuite — press play. The prompt in the bar is the exact script it read.
voiceover.mp3Eleven v30:000:11ReplicateEleven v3Voice · Rachel ▶MP3The last train had already left, but she decided to walk anyway. The city was quieter than she remembered…Speak
Real output · Eleven v3narration.mp3
0:00 / 0:10
Compose
Describe the mood, get the music.
Instrumentation, tempo, texture — describe it like a brief to a session musician and the model composes it. Royalty questions don't follow you around: it's your generation, made with your key.The 30-second clip below is a real Lyria 3 composition from the exact prompt shown — the same brief every music model in the catalog answers on its detail page.
rainy-morning.mp3Lyria 30:000:30ReplicateLyria 3Music30sSlow ambient piece for a rainy morning: warm analog pads, soft tape hiss, and a simple four-note piano motif repeating. No drums.Generating
Real output · Lyria 3ambient.mp3
0:00 / 0:30
Foley
Sound effects on demand.
The door slam, the forest ambience, the projector whir — describe the sound instead of digging through sample libraries, and drop the result straight into a video composition or export it for your editor.Below: a real SFX 1.5 generation of the prompt shown — a heavy wooden door in a stone corridor.
door-slam.mp3SFX 1.50:000:10RunwareSFX 1.5Sound effectA heavy wooden door slamming shut in a stone corridor, with a short natural reverb tail.Generate
Real output · SFX 1.5impact.mp3
0:00 / 0:10
Enhance
Type the idea, get the whole brief.
For music and effects, press the wand and your own text model turns one line into a full brief: what the track is about first, then genre, instrumentation, tempo, and mood. Lyrics land wherever the model reads them. Speech has no wand, because the text you type is what gets spoken.Music is billed per generation, so the takes you discard are the expensive part.
ComposerEnhance prompt with AIYou typeda sad song about kids with no shoes while others complain about old trainers↓Rewritten by your text model, for the model you pickedSent to the modelA poignant folk ballad about children who have no shoes at all, set against people complaining about last season’s trainers. Fingerpicked acoustic guitar, soft piano, light cello, around 68 BPM.[Verse 1]Bare feet on the gravel roadA little girl walks miles to school[Chorus]And somewhere across the world tonightLyria 3 ProReads lyrics from the prompt itself[Verse] and [Chorus] tags
Edit
Clean it up without leaving the app.
A waveform editor built into the player covers the usual cleanup — no export to another tool. Every edit renders locally on your machine and saves in place or as a new copy, keeping the file's own format so a trimmed MP3 stays an MP3.
Trim to the take
Cut the silence off both ends of a recording, or pull one clean take out of a long session — drag the handles on the waveform and everything outside the selection goes.
podcast-take.mp30:06 → 0:48ResetSave new copy
Fade in, fade out
Give a generated music bed a smooth entrance and exit before it goes under a voiceover — set the fade lengths and the envelope is applied to the waveform.
rainy-morning.wavFade in · 1.5sFade out · 2sSave new copy
Fix the level
A quiet voiceover next to a loud music bed is the most common mix problem there is — boost or cut the clip's volume until the levels sit right.
voiceover.mp3Volume120%
Change the speed
Turn a 42-minute lecture recording into a 28-minute listen at 1.5×, or slow a fast take down — the duration updates with the rate.
lecture-recording.mp30.5×1×1.5×2×42:00 → 28:00
Convert the format
Hand your editor a lossless WAV or FLAC, or squeeze a session down to MP3 or OGG for sharing — pick the target and it transcodes straight from the source file, with no lossy round-trip in between.
field-notes.flacfield-notes.flac→field-notes.wavMP3WAVFLACOGGSave new copy
Workspace
The rest of the booth.
The small things around the generate button — hearing a voice first, knowing the price, getting back to the prompt that worked.
Audition a voicePlay any voice in the roster straight from the picker before you commit a script to it. First listen costs a fraction of a cent; every one after is cached and free.
Clone a voiceOn models that support it, attach a sample and its transcript and the model reads your script in that voice instead of a preset one.
Write the brief with AIFor music and effects, type the gist and hit the wand: your text model expands it into subject, genre, instrumentation, tempo, and mood, writing the lyrics too on models that sing from the prompt. Speech is left alone, since the text is what gets spoken.
Know what it costs firstThe composer shows the estimated price before you generate — per clip for music, from your live character count for speech — and a status line tracks the run.
Reuse what workedRecent prompts sit as chips above the composer, and every generated clip remembers the model and settings behind it, reloadable in one click.
Files that stay yoursClips are ordinary files in your project folder — drag recordings in, filter a long list by name, and open or back them up with any tool you like.
Catalog
Available models & providers.
Every audio model in the catalog — text-to-speech, music, and sound effects — all cloud-hosted through Runware and Replicate with your own keys, each with real sample clips on its detail page.
How speech, music, and sound-effect generation work in CSuite — voices, costs, editing, and formats.
For speech: Eleven v3, Eleven Turbo v2.5, Gemini 3.1 Flash TTS, Speech 02 Turbo, Speech 2.8, Grok Text-to-Speech, and Fish Audio S2.1 Pro. For music: Lyria 3 and Lyria 3 Pro, Eleven Music, Music 2.6, and ACE-Step v1.5. For sound effects: SFX 1.5. All run in the cloud through Runware and Replicate with your own API keys — browse per-model details and real sample clips at csuite.so/models.
A Type selector switches the composer between the three categories, and the model picker filters to models that support the active one. Your model choice is remembered per category, so flipping from a voiceover to a music bed and back doesn't lose your setup.
Each text-to-speech model ships its own voice roster — pick a voice in the generation settings, and on models that expose them, tune per-voice controls like speed, stability, and similarity. Several models also support multiple languages. A play button sits beside the picker so you can hear a voice before writing the whole script: the first play of a voice generates a short sample for a fraction of a cent, and every play after that is cached and free.
Yes, on models that support it — Fish Audio S2.1 Pro today. Choose an audio sample, type what's said in it, and the model reads your script in that voice instead of a preset one. The sample is copied into your project like any other reference file.
For music and sound effects, yes: type the gist and hit the wand, and your text model expands it into a fuller brief — genre, instrumentation, tempo, and mood for music; source, texture, and acoustics for an effect. It leads with what the track is actually about, so a song about something specific stays about that, and it follows the model you picked: one that sings from its prompt gets real verses marked with the section tags it understands, while one with a lyrics field of its own gets a style brief and keeps your lyrics untouched. Text-to-speech deliberately has no wand, because the text you type is what gets spoken and rewriting it would change the words. Recent prompts also stay as chips above the composer, and every generated clip remembers the model and settings behind it.
Generation is cloud-only for now — there are no local audio models in the catalog yet. Editing is fully local, though: trims, fades, volume, speed, and conversion all run on your machine, and your recordings never upload anywhere.
Text-to-speech is typically billed per 1,000 characters, and music or sound effects per generation or per second of output — always by the platform (Runware or Replicate) directly to your account, with no CSuite markup or metering. The app's analytics show spend per provider.
Yes. Drop MP3, WAV, FLAC, OGG, or AAC files into your project folder (or drag them into the sidebar) and they open in the same player: trim a take, fade the ends, adjust volume, change speed, or convert — saving in place or as a new copy alongside the original.
The file's own. An edit renders locally and is re-encoded back to the source format through the bundled ffmpeg, so trimming an MP3 gives you an MP3, not a WAV that's ten times the size. If you do want a different format, the convert tool changes it deliberately — MP3, WAV, FLAC, or OGG — transcoding straight from the source rather than through a lossy round-trip.
In your project folder, as plain files on your own disk — generated clips, imports, and edited copies alike. No proprietary library, no cloud sync.
One-time payment. Yours forever.
No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.