Making audio with CSuite.
Every Type, voice, setting and tool in the audio workspace, with a figure for each, followed by recipes that chain them into finished work. The short version lives on the main guide; this page is for when you want to know exactly what a control does.
Get around
The workspace
Open Audio from the rail. One workspace covers four jobs: speaking a script, composing music, making a sound effect and transcribing a recording. The shape is the same as the other workspaces, a file list, a player in the middle and the prompt box below, and a Type control in the Model popover decides which job the prompt box is doing.
- 1 · Sidebar
- Your project folder pill and every audio file in it, newest first: MP3, WAV, FLAC, OGG, AAC and M4A. A “Generating…” row sits at the top while a clip is being made.
- 2 · Title
- The file name without its extension. Edit it to rename the file.
- 3 · Editing toolbar
- Trim, Fade, Volume, Speed and Format. Each opens an options row under the header and works on the waveform.
- 4 · Info and Model
- Info lists duration, size, format and, for clips CSuite made, the model and prompt. The Model button shows the selected model; its popover holds Type, the picker and the settings.
- 5 · Player
- The waveform with a playhead, then play, a scrub bar and the time. Space plays and pauses.
- 6 · Transcript
- Under the player for a recording that has one: timed lines you can click, Copy and Open in Text.
- 7 · Composer
- The script, music brief, effect description or transcription hint, depending on Type, with the wand where it applies, reference pickers the model takes, the active settings and cost, and Create (or Transcribe).
The ✕ beside the folder pill (“Clear selection”) deselects the file so the next prompt makes a new clip rather than acting on the selected one. Esc does the same outside a text field, as does clicking Audio in the rail again.
Files
- Open a file by clicking its row. The row shows the format and the clip’s length.
- More actions on a row: Show in Finder (Show in File Explorer on Windows, Show in file manager on Linux), Duplicate, which also selects the copy, and Delete, which asks first. Delete or Backspace with a row selected also deletes.
- Add your own recordings by dragging MP3, WAV, FLAC, OGG, AAC or M4A files onto the list (“Drop audio to add it here”). They are copied into the project folder and open in the same player with the same tools. Video goes through the composer’s Upload under Transcribe instead, which extracts its soundtrack.
- Filter and sort appear above 20 files: “Filter audio by name…” and a sort button cycling Newest, Oldest and Name.
- Names: a generated clip takes the first two words of its prompt or script, title-cased; a clash gets a “ 2”, “ 3” suffix. An edit saved as a new copy sits beside the original; a transcript takes the recording’s name with
.txtand.srt.
Generate
The four Types
Open the Model button and the first row is Type: Music, Sound Effects, Text to Speech or Transcribe. The model list filters to match, the composer’s placeholder changes (“Describe the music to generate…”, “Enter text to convert to speech…”), and the wand, reference pickers and cost hint follow the Type.
- A default per Type: each of the four keeps its own model, so switching from a voiceover to a music bed and back loses nothing. A pick made here lasts the session; the saved default is set under Models → Defaults.
- Auto works per Type: under Auto the prompt decides the model among those that do that job, preferring Runware, then Replicate, then OpenRouter, or an installed on-device model with no cloud key.
- Where models run: most through Runware, Replicate and OpenRouter on your platform keys; GPT Audio, GPT Audio Mini and MAI-Voice are OpenRouter-only; Lyria 3.5 needs a Google key; and a key under Models → Vendors runs OpenAI, Google, ElevenLabs, xAI, MiniMax, Kling and Fish Audio models direct. Supertonic TTS and the Whisper models run on the bundled Hugging Face runtime with no key.
- A Type with no models keeps the Type row and says so in the picker’s place, naming the providers to add.
Text to speech
Voice cloning
Models that clone get a Voice cloning control in the Model popover: Fish Audio S2.1 Pro on Runware, OpenRouter or a Fish Audio key, and S2 Pro, S1 and S2.1 Pro Free on a Fish Audio key. Press Choose audio…, pick a clean recording of the voice, then type the exact words spoken in it (“The exact words spoken in the reference sample — the provider needs it to clone the voice”). While a sample is attached the preset Voice field is disabled, since the sample replaces it, and the settings summary reads “Voice: Cloned sample”. Remove the sample to go back to presets.
Seed Audio 1.0 works differently: it takes a Ref. audio clip with no transcript, and can take reference images instead, one kind at a time. The sample is copied into your project like any other reference file.
Music
Describe the track in the composer: what it is about, genre, instruments, tempo, mood. The models differ in what else they take:
- Lyria 3 and Lyria 3 Pro (Replicate, OpenRouter) sing from the prompt and accept up to four Reference images to set a mood; Lyria 3.5 runs on a Google key.
- Music 2.6 (Runware, Replicate) and ACE-Step v1.5 (Runware) add a Lyrics field that reads
[verse]and[chorus]tags; leave it empty for an instrumental. ACE-Step also takes a Duration from 30 seconds to five minutes and a Source audio clip to continue or restyle. - ElevenLabs Music (Replicate or an ElevenLabs key) sets the length in the popover, from a few seconds to five minutes, with an Instrumental switch.
- Seed Audio 1.0 is listed under Music too and takes reference audio or images.
A duration-billed model’s cost hint reads “~$ / clip” from the slider; a flat-priced one shows its price. The status line walks “Starting…”, “Generating…” and “Saving…” and Stop cancels on the provider too.
Sound effects
Describe the sound (“a rusty iron gate creaking open slowly, birds in the distance”) and set a Duration. SFX 1.5 on Runware makes clips of one to ten seconds and takes a Source video, so the sound follows what happens on screen. ElevenLabs Sound Effects v2 on an ElevenLabs key runs from half a second to thirty, with a prompt-influence dial and a Loop switch for seamless beds. Kling Text to Audio runs on a Kling key, and Seed Audio 1.0 is listed here too. Short effects are cheap; the cost hint reads “~$ / clip”.
The wand
The wand appears for Music and Sound Effects only. For music it leads with what the track is about and never drops it, then adds genre, instrumentation, tempo, mood and structure. On a model that sings from its prompt it writes real lines with the section tags that model reads; on a model with its own Lyrics field it writes a style brief and leaves your lyrics alone. For an effect it names the source, texture, acoustics and dynamics. It runs on your text model and puts the rewrite back in the box for you to edit.
Text to Speech has no wand because the text is what gets spoken; Transcribe has none because its box holds a spelling hint, not a prompt. Recent prompts for Music, Sound Effects and Text to Speech stay as chips above the box.
On-device models
Download these under AI model providers and they appear in the picker with no key required. Supertonic TTS speaks English in ten voices (F1 to F5, M1 to M5) with a speed dial and a Steps setting; more steps, cleaner audio, more time. Its voice previews are generated on your machine, so even the first listen is free. Whisper Base, Whisper Small and Whisper Large V3 Turbo transcribe entirely offline; pick the spoken language yourself, since on-device Whisper cannot detect it (English by default). Music and sound effects are cloud-only. Editing is local for every file whatever made it.
Transcribe
Transcribing a recording
The transcript
- Files: a
.txtbeside the recording always, plus an.srtsubtitle file when the model returns timestamps. Whisper, Whisper Large V3 and Turbo, Scribe v2, on-device Whisper and the Fish Audio models on a Fish Audio key do; GPT-4o Transcribe and Mini, MAI-Transcribe 2, Gemini 3.5 Transcribe, Muse Voice Transcribe and the Fish models through OpenRouter return text only. With speaker labels the text is one paragraph per turn and each subtitle cue is prefixed “Speaker N:”. - The view under the player lists timed lines; click one to jump playback there, and the line at the playhead stays highlighted while it plays. An untimed transcript shows as plain text.
- Copy puts the whole text on the clipboard (“Copied”). Open in Text opens the
.txtin the Text workspace for editing or AI rewrites; edits made there show here when you come back.
Player and edit
Player and info
- Waveform and transport: the clip drawn as bars with a playhead, then play, a scrub bar and the time. Space plays and pauses whenever your focus is not in a text field. A file the preview cannot decode still plays but shows “Waveform unavailable”.
- Info (the i in the header) lists Duration, Size, Format and Modified, plus Model and Prompt for clips CSuite made. Use prompt & settings refills the composer: the prompt always, the settings when the same model is still selected.
- Every tool below opens an options row under the header with a Save a new copy switch and Apply. On, the result is a sibling file; off, the file is rewritten in place and the player reloads it. Edits render on your machine and are re-encoded to the file’s own format, so a trimmed MP3 stays an MP3; an edited AAC or M4A is saved as WAV, since there is no bundled encoder for it.
Trim
Press Trim and two handles appear on the waveform (“Drag handles on waveform”). Drag them to keep just the part you want; the options row shows the start and end times. Apply writes the kept range. Cut the silence off both ends of a take, or pull one clean take out of a long session, then trim again for the next one with Save a new copy on.
Fade
Type a Fade in and a Fade out length in seconds, in tenths; either can be zero, and each is capped at half the clip so the two never cross. The ramp is linear and is shown on the waveform before you Apply. A music bed under a voice usually wants a short fade in and a longer fade out.
Volume and speed
- Volume is a slider from 0% (silent) to 200% (double), 100% being unchanged. Push a quiet voiceover up or a loud bed down; pushing past the point where the loudest bars hit the top clips the audio, so prefer cutting the loud clip to boosting the quiet one.
- Speed offers 0.5×, 0.75×, 1×, 1.25×, 1.5× and 2×. The duration updates with the rate, and the pitch shifts with it, unlike the video workspace, which keeps the pitch. A 42-minute lecture at 1.5× is a 28-minute listen.
Format
Format transcodes straight from the source file to MP3, WAV, FLAC or OGG, with no lossy round-trip through the editor. The file’s current format is listed as “Already MP3” and choosing it in place shows a toast; turn on Save a new copy to make a same-format copy. Hand an editor a lossless WAV or FLAC, or squeeze a session down to MP3 or OGG for sharing.
Recipes
A finished voiceover
Show notes from an episode
.srt beside the recording is the subtitle file for the video version.A music bed under a voice
Sound for a clip
Subtitles for a video
.srt is beside the extracted MP3. For wording fixes, Open in Text edits the .txt; for timed cues, open the .srt in Text as a plain file and edit the lines, keeping the timestamps..srt into your video editor or player, or into a composition’s caption track.Narration with no account
Reference
Keyboard shortcuts
⌘ is Ctrl on Windows and Linux. The main guide lists every workspace’s shortcuts together.
Player
Space- Play or pause the selected clip
Esc- Clear the selection (outside a text field)
Files and composer
Enter / Shift+Enter- Create or Transcribe / add a line
Delete / Backspace- Delete the selected file (asks first)
⌘+ / ⌘- / ⌘0- Zoom the interface in, out, or back to 100%
Troubleshooting
- “No API providers yet.” Add Replicate, Runware or OpenRouter under AI model providers, or download an on-device model; Supertonic TTS and Whisper need no key.
- “None of your providers offer Transcribe models.” Transcription runs on OpenAI, OpenRouter, ElevenLabs, Fish Audio or Replicate keys, or on-device Whisper. Add one, or pick another Type.
- “None of your providers offer Music models.” (or Sound Effects) Music and effects are cloud-only; add Runware, Replicate or OpenRouter.
- The wand is missing. It appears for Music and Sound Effects only. Speech is spoken verbatim and Transcribe takes a hint, so neither has one.
- “This model reads the audio only, so it takes no hint.” The hint box is disabled on audio-only transcribers. Use Whisper or GPT-4o Transcribe when spellings matter.
- “Select an audio file to transcribe.” Transcribe works on the selected sidebar file; pick one, or Upload or drop one.
- “Only audio … or video … files can be transcribed.” Convert other formats first. A video with no audio stream is refused.
- No
.srtappeared. The model returns text only (GPT-4o Transcribe and Mini, MAI-Transcribe 2, Gemini 3.5 Transcribe, Muse, Fish via OpenRouter). Use Whisper, Scribe v2 or on-device Whisper for subtitles. - “No speech was recognized in this file.” The recording is silent or the wrong language was forced. On on-device Whisper set the language; on cloud models leave Auto-detect.
- Wrong language on on-device Whisper. It cannot detect it; choose the spoken language in the popover.
- “Couldn’t generate a voice sample.” The provider call failed; check the key and try again. Previews on a cloud model are billed only on the first play.
- The Voice field is disabled. A cloning sample is attached and replaces it; remove the sample to use presets.
- No cost estimate. Per-token speech models, Auto and on-device models show none; character-billed speech shows one once the box has text.
- “Waveform unavailable”. The preview could not decode the file; it still plays and converts. Use Format to make a WAV or MP3 copy for editing.
- “This file is already MP3.” Pick another format, or turn on Save a new copy.
- An edited M4A came back as WAV. There is no bundled AAC encoder; convert the WAV to MP3 with Format if you need it small.
- Speed changed the pitch. That is how the audio workspace works; the video workspace keeps pitch.
- An on-device model fails to load on Windows. Install Microsoft’s Visual C++ Redistributable from the link in the message, then retry.