Skip to content
CSuite
GuideAudio

Making audio with CSuite.

Every Type, voice, setting and tool in the audio workspace, with a figure for each, followed by recipes that chain them into finished work. The short version lives on the main guide; this page is for when you want to know exactly what a control does.

Part 1

Get around

1.1

The workspace

The audio workspace with a generated narration selected and its transcript showing. The numbers match the legend below.

Open Audio from the rail. One workspace covers four jobs: speaking a script, composing music, making a sound effect and transcribing a recording. The shape is the same as the other workspaces, a file list, a player in the middle and the prompt box below, and a Type control in the Model popover decides which job the prompt box is doing.

1 · Sidebar
Your project folder pill and every audio file in it, newest first: MP3, WAV, FLAC, OGG, AAC and M4A. A “Generating…” row sits at the top while a clip is being made.
2 · Title
The file name without its extension. Edit it to rename the file.
3 · Editing toolbar
Trim, Fade, Volume, Speed and Format. Each opens an options row under the header and works on the waveform.
4 · Info and Model
Info lists duration, size, format and, for clips CSuite made, the model and prompt. The Model button shows the selected model; its popover holds Type, the picker and the settings.
5 · Player
The waveform with a playhead, then play, a scrub bar and the time. Space plays and pauses.
6 · Transcript
Under the player for a recording that has one: timed lines you can click, Copy and Open in Text.
7 · Composer
The script, music brief, effect description or transcription hint, depending on Type, with the wand where it applies, reference pickers the model takes, the active settings and cost, and Create (or Transcribe).

The ✕ beside the folder pill (“Clear selection”) deselects the file so the next prompt makes a new clip rather than acting on the selected one. Esc does the same outside a text field, as does clicking Audio in the rail again.

1.2

Files

The sidebar’s row menu and the drop target for files from your desktop.
  • Open a file by clicking its row. The row shows the format and the clip’s length.
  • More actions on a row: Show in Finder (Show in File Explorer on Windows, Show in file manager on Linux), Duplicate, which also selects the copy, and Delete, which asks first. Delete or Backspace with a row selected also deletes.
  • Add your own recordings by dragging MP3, WAV, FLAC, OGG, AAC or M4A files onto the list (“Drop audio to add it here”). They are copied into the project folder and open in the same player with the same tools. Video goes through the composer’s Upload under Transcribe instead, which extracts its soundtrack.
  • Filter and sort appear above 20 files: “Filter audio by name…” and a sort button cycling Newest, Oldest and Name.
  • Names: a generated clip takes the first two words of its prompt or script, title-cased; a clash gets a “ 2”, “ 3” suffix. An edit saved as a new copy sits beside the original; a transcript takes the recording’s name with .txt and .srt.
Part 2

Generate

2.1

The four Types

The Model popover opens with Type on top; the picker below it lists only models that do that job.

Open the Model button and the first row is Type: Music, Sound Effects, Text to Speech or Transcribe. The model list filters to match, the composer’s placeholder changes (“Describe the music to generate…”, “Enter text to convert to speech…”), and the wand, reference pickers and cost hint follow the Type.

  • A default per Type: each of the four keeps its own model, so switching from a voiceover to a music bed and back loses nothing. A pick made here lasts the session; the saved default is set under Models → Defaults.
  • Auto works per Type: under Auto the prompt decides the model among those that do that job, preferring Runware, then Replicate, then OpenRouter, or an installed on-device model with no cloud key.
  • Where models run: most through Runware, Replicate and OpenRouter on your platform keys; GPT Audio, GPT Audio Mini and MAI-Voice are OpenRouter-only; Lyria 3.5 needs a Google key; and a key under Models → Vendors runs OpenAI, Google, ElevenLabs, xAI, MiniMax, Kling and Fish Audio models direct. Supertonic TTS and the Whisper models run on the bundled Hugging Face runtime with no key.
  • A Type with no models keeps the Type row and says so in the picker’s place, naming the providers to add.
2.2

Text to speech

ElevenLabs v3 on an ElevenLabs key: a voice with its preview button, language, and the speed, stability and similarity dials. The script sits in the composer.
1
Pick a voice and hear it
The Voice field lists the model’s roster, with a play button beside it. The first play of a voice generates a short sample (“generates one short clip the first time, then it’s cached”) for a fraction of a cent on a cloud model; after that it is instant and free. On Supertonic TTS the samples are made on your machine.
2
Tune what the model offers
Speed on most; Stability and Similarity on ElevenLabs; Language where the model can be told (otherwise it follows the text); emotion presets on MiniMax Speech; a style instruction on GPT-4o Mini TTS; a Format where the vendor offers several. Only what the model accepts is shown.
3
Write the text and create
The composer is labelled Text and what you type is spoken verbatim; there is deliberately no wand here. A cost hint beside Create grows with the text on models billed per character. The clip lands in your project folder and plays.
TipPunctuation is pacing. A full stop makes a longer pause than a comma, an ellipsis a longer one still, and a line break between paragraphs gives most models a breath. Spell out numbers and abbreviations you want read a particular way.
2.3

Voice cloning

Fish Audio S2.1 Pro with a sample attached: the preset Voice is disabled, the transcript of the sample is required, and the summary reads Cloned sample.

Models that clone get a Voice cloning control in the Model popover: Fish Audio S2.1 Pro on Runware, OpenRouter or a Fish Audio key, and S2 Pro, S1 and S2.1 Pro Free on a Fish Audio key. Press Choose audio…, pick a clean recording of the voice, then type the exact words spoken in it (“The exact words spoken in the reference sample — the provider needs it to clone the voice”). While a sample is attached the preset Voice field is disabled, since the sample replaces it, and the settings summary reads “Voice: Cloned sample”. Remove the sample to go back to presets.

Seed Audio 1.0 works differently: it takes a Ref. audio clip with no transcript, and can take reference images instead, one kind at a time. The sample is copied into your project like any other reference file.

TipOnly clone a voice you have the right to use: your own, or someone who has agreed in writing. The terms rule out cloning without consent.
2.4

Music

ACE-Step v1.5 with a Lyrics field and a Duration slider, a source track attached to restyle, and the run in progress.

Describe the track in the composer: what it is about, genre, instruments, tempo, mood. The models differ in what else they take:

  • Lyria 3 and Lyria 3 Pro (Replicate, OpenRouter) sing from the prompt and accept up to four Reference images to set a mood; Lyria 3.5 runs on a Google key.
  • Music 2.6 (Runware, Replicate) and ACE-Step v1.5 (Runware) add a Lyrics field that reads [verse] and [chorus] tags; leave it empty for an instrumental. ACE-Step also takes a Duration from 30 seconds to five minutes and a Source audio clip to continue or restyle.
  • ElevenLabs Music (Replicate or an ElevenLabs key) sets the length in the popover, from a few seconds to five minutes, with an Instrumental switch.
  • Seed Audio 1.0 is listed under Music too and takes reference audio or images.

A duration-billed model’s cost hint reads “~$ / clip” from the slider; a flat-priced one shows its price. The status line walks “Starting…”, “Generating…” and “Saving…” and Stop cancels on the provider too.

2.5

Sound effects

SFX 1.5 with a Duration slider and a source video attached, so the effect is timed to the picture.

Describe the sound (“a rusty iron gate creaking open slowly, birds in the distance”) and set a Duration. SFX 1.5 on Runware makes clips of one to ten seconds and takes a Source video, so the sound follows what happens on screen. ElevenLabs Sound Effects v2 on an ElevenLabs key runs from half a second to thirty, with a prompt-influence dial and a Loop switch for seamless beds. Kling Text to Audio runs on a Kling key, and Seed Audio 1.0 is listed here too. Short effects are cheap; the cost hint reads “~$ / clip”.

2.6

The wand

The prompt enhancer on a music brief: subject first, then genre, instruments, tempo and mood, with sung lines for a model that reads them.

The wand appears for Music and Sound Effects only. For music it leads with what the track is about and never drops it, then adds genre, instrumentation, tempo, mood and structure. On a model that sings from its prompt it writes real lines with the section tags that model reads; on a model with its own Lyrics field it writes a style brief and leaves your lyrics alone. For an effect it names the source, texture, acoustics and dynamics. It runs on your text model and puts the rewrite back in the box for you to edit.

Text to Speech has no wand because the text is what gets spoken; Transcribe has none because its box holds a spelling hint, not a prompt. Recent prompts for Music, Sound Effects and Text to Speech stay as chips above the box.

2.7

On-device models

What runs on the bundled Hugging Face runtime without a key.

Download these under AI model providers and they appear in the picker with no key required. Supertonic TTS speaks English in ten voices (F1 to F5, M1 to M5) with a speed dial and a Steps setting; more steps, cleaner audio, more time. Its voice previews are generated on your machine, so even the first listen is free. Whisper Base, Whisper Small and Whisper Large V3 Turbo transcribe entirely offline; pick the spoken language yourself, since on-device Whisper cannot detect it (English by default). Music and sound effects are cloud-only. Editing is local for every file whatever made it.

Part 3

Transcribe

3.1

Transcribing a recording

Transcribe on Scribe v2: the box holds an optional hint, Upload leads the composer, the button reads Transcribe, and the status line shows the percentage.
1
Bring the recording in
Select an audio file in the sidebar, or press Upload in the composer, or drop files onto the composer or the empty panel (“Drop audio or video to transcribe it”). Audio is copied into your project; for a video (MP4, WebM, MOV, AVI) only the soundtrack is extracted, as an MP3. The row reads “Adding…” meanwhile, and nothing is sent until you press Transcribe.
2
Choose the model and settings
Set Type to Transcribe and pick a model: Whisper, Whisper Large V3 and Turbo, GPT-4o Transcribe and Mini, MAI-Transcribe 2, Gemini 3.5 Transcribe, Muse Voice Transcribe, Scribe v2, the Fish Audio Transcribe models, or on-device Whisper. Language is Auto-detect on cloud models. Scribe v2 adds Speaker labels and Sound tags; Fish Audio’s Transcribe 1 Pro labels speakers on its own.
3
Add a hint if the model reads one
On Whisper and GPT-4o Transcribe the box takes names, jargon or context so they come out spelled right. Models that only listen (Scribe v2, MAI-Transcribe 2, Gemini 3.5 Transcribe, the Fish models, Muse, on-device Whisper) disable the box with “This model reads the audio only, so it takes no hint.”
4
Transcribe
The cost hint prices the selected file’s length (“~$ / file”). The status line walks “Preparing audio…”, “Transcribing…” with a percentage, and “Saving…”. Long or large recordings are cut into equal pieces for models with a limit and stitched back together; Stop keeps nothing. An empty result fails with “No speech was recognized in this file.” and writes nothing.
3.2

The transcript

A diarized transcript with timed lines: the playhead line is highlighted, and Copy and Open in Text sit in its header.
  • Files: a .txt beside the recording always, plus an .srt subtitle file when the model returns timestamps. Whisper, Whisper Large V3 and Turbo, Scribe v2, on-device Whisper and the Fish Audio models on a Fish Audio key do; GPT-4o Transcribe and Mini, MAI-Transcribe 2, Gemini 3.5 Transcribe, Muse Voice Transcribe and the Fish models through OpenRouter return text only. With speaker labels the text is one paragraph per turn and each subtitle cue is prefixed “Speaker N:”.
  • The view under the player lists timed lines; click one to jump playback there, and the line at the playhead stays highlighted while it plays. An untimed transcript shows as plain text.
  • Copy puts the whole text on the clipboard (“Copied”). Open in Text opens the .txt in the Text workspace for editing or AI rewrites; edits made there show here when you come back.
Part 4

Player and edit

4.1

Player and info

The waveform and transport, with the Info popover open on a generated narration.
  • Waveform and transport: the clip drawn as bars with a playhead, then play, a scrub bar and the time. Space plays and pauses whenever your focus is not in a text field. A file the preview cannot decode still plays but shows “Waveform unavailable”.
  • Info (the i in the header) lists Duration, Size, Format and Modified, plus Model and Prompt for clips CSuite made. Use prompt & settings refills the composer: the prompt always, the settings when the same model is still selected.
  • Every tool below opens an options row under the header with a Save a new copy switch and Apply. On, the result is a sibling file; off, the file is rewritten in place and the player reloads it. Edits render on your machine and are re-encoded to the file’s own format, so a trimmed MP3 stays an MP3; an edited AAC or M4A is saved as WAV, since there is no bundled encoder for it.
4.2

Trim

Trim handles on the waveform; the dimmed bars outside them are dropped on Apply.

Press Trim and two handles appear on the waveform (“Drag handles on waveform”). Drag them to keep just the part you want; the options row shows the start and end times. Apply writes the kept range. Cut the silence off both ends of a take, or pull one clean take out of a long session, then trim again for the next one with Save a new copy on.

4.3

Fade

Fade in and fade out in seconds, drawn as an envelope over the waveform.

Type a Fade in and a Fade out length in seconds, in tenths; either can be zero, and each is capped at half the clip so the two never cross. The ramp is linear and is shown on the waveform before you Apply. A music bed under a voice usually wants a short fade in and a longer fade out.

4.4

Volume and speed

Volume as a percentage, and Speed as six presets.
  • Volume is a slider from 0% (silent) to 200% (double), 100% being unchanged. Push a quiet voiceover up or a loud bed down; pushing past the point where the loudest bars hit the top clips the audio, so prefer cutting the loud clip to boosting the quiet one.
  • Speed offers 0.5×, 0.75×, 1×, 1.25×, 1.5× and 2×. The duration updates with the rate, and the pitch shifts with it, unlike the video workspace, which keeps the pitch. A 42-minute lecture at 1.5× is a 28-minute listen.
4.5

Format

The Format menu lists MP3, WAV, FLAC and OGG; the file’s own format reads Already and is refused in place.

Format transcodes straight from the source file to MP3, WAV, FLAC or OGG, with no lossy round-trip through the editor. The file’s current format is listed as “Already MP3” and choosing it in place shows a toast; turn on Save a new copy to make a same-format copy. Hand an editor a lossless WAV or FLAC, or squeeze a session down to MP3 or OGG for sharing.

Part 5

Recipes

5.1

A finished voiceover

1
Write the script in Text
Draft it in the Text workspace, read it aloud once, and mark pauses with punctuation and paragraph breaks. Copy it.
2
Audition, then speak
Set Type to Text to Speech, play two or three voices from the preview button, pick one, and paste the script in the composer. Check the cost hint, press Create.
3
Clean the ends
Trim any lead-in silence and add a short Fade out. Keep Save a new copy on until you are happy, then delete the drafts from the row menu.
4
Lay it over the picture
Drop the clip into an audio or video composition to mix it with a bed and export.
5.2

Show notes from an episode

1
Transcribe with speakers
Upload the episode, set Type to Transcribe, pick Scribe v2 with Speaker labels on (or Fish Audio’s Transcribe 1 Pro), and press Transcribe. A long episode is split and stitched for you.
2
Check the names
Scrub through the transcript; click a line to hear it. If a guest’s name is misspelled, run it again on Whisper with the names in the Hint box.
3
Turn it into notes
Open in Text, then ask: “Write show notes with a two-sentence summary, five bullet takeaways and a list of topics with timestamps.” The .srt beside the recording is the subtitle file for the video version.
5.3

A music bed under a voice

1
Ask for a bed, not a song
Set Type to Music and describe something that stays out of the way: “warm lo-fi hip hop, no vocals, steady, 60 seconds”. On Music 2.6 or ACE-Step leave the Lyrics empty; on ElevenLabs Music turn Instrumental on. Use the wand if the brief is thin.
2
Shape it
Fade in over a second and out over three, then Volume down to about 40% so the voice sits on top. Save as a new copy.
3
Loop or lengthen
Need more? Attach the clip as Source audio on ACE-Step and ask it to continue in the same style, or use a composition to repeat it under the voice.
5.4

Sound for a clip

1
Give the model the picture
Set Type to Sound Effects, pick SFX 1.5, and attach the video as Source video. Describe what should be heard and set the Duration to the clip’s length (up to ten seconds).
2
Layer, don’t overload
One effect per prompt: the gate, then the birds, then the footsteps. Short clips are a few cents each. Build the mix in the composition editor with each effect on its own row.
5.5

Subtitles for a video

1
Transcribe the video
Under Transcribe, press Upload and pick the MP4. Only the soundtrack is extracted into your project; the video stays where it was. Pick a model that returns timestamps: Whisper, Whisper Large V3 or Turbo, Scribe v2, or on-device Whisper.
2
Fix the text, keep the timing
The .srt is beside the extracted MP3. For wording fixes, Open in Text edits the .txt; for timed cues, open the .srt in Text as a plain file and edit the lines, keeping the timestamps.
3
Use it
Load the .srt into your video editor or player, or into a composition’s caption track.
5.6

Narration with no account

1
Download Supertonic TTS
Under AI model providers, download the model on the Hugging Face runtime. On Windows the runtime needs Microsoft’s Visual C++ Redistributable; the app links to it if it is missing.
2
Pick a voice, free
Set Type to Text to Speech, select Supertonic, and play through the ten voices. Nothing is billed, nothing leaves the machine.
3
Speak and edit offline
Paste the script, raise Steps if the output sounds rough, create, then trim and fade as usual. Add an on-device Whisper model and the round trip, speech and transcription both, never touches the network.
Part 6

Reference

6.1

Keyboard shortcuts

⌘ is Ctrl on Windows and Linux. The main guide lists every workspace’s shortcuts together.

Player

Space
Play or pause the selected clip
Esc
Clear the selection (outside a text field)

Files and composer

Enter / Shift+Enter
Create or Transcribe / add a line
Delete / Backspace
Delete the selected file (asks first)
⌘+ / ⌘- / ⌘0
Zoom the interface in, out, or back to 100%
6.2

Troubleshooting

  • “No API providers yet.” Add Replicate, Runware or OpenRouter under AI model providers, or download an on-device model; Supertonic TTS and Whisper need no key.
  • “None of your providers offer Transcribe models.” Transcription runs on OpenAI, OpenRouter, ElevenLabs, Fish Audio or Replicate keys, or on-device Whisper. Add one, or pick another Type.
  • “None of your providers offer Music models.” (or Sound Effects) Music and effects are cloud-only; add Runware, Replicate or OpenRouter.
  • The wand is missing. It appears for Music and Sound Effects only. Speech is spoken verbatim and Transcribe takes a hint, so neither has one.
  • “This model reads the audio only, so it takes no hint.” The hint box is disabled on audio-only transcribers. Use Whisper or GPT-4o Transcribe when spellings matter.
  • “Select an audio file to transcribe.” Transcribe works on the selected sidebar file; pick one, or Upload or drop one.
  • “Only audio … or video … files can be transcribed.” Convert other formats first. A video with no audio stream is refused.
  • No .srt appeared. The model returns text only (GPT-4o Transcribe and Mini, MAI-Transcribe 2, Gemini 3.5 Transcribe, Muse, Fish via OpenRouter). Use Whisper, Scribe v2 or on-device Whisper for subtitles.
  • “No speech was recognized in this file.” The recording is silent or the wrong language was forced. On on-device Whisper set the language; on cloud models leave Auto-detect.
  • Wrong language on on-device Whisper. It cannot detect it; choose the spoken language in the popover.
  • “Couldn’t generate a voice sample.” The provider call failed; check the key and try again. Previews on a cloud model are billed only on the first play.
  • The Voice field is disabled. A cloning sample is attached and replaces it; remove the sample to use presets.
  • No cost estimate. Per-token speech models, Auto and on-device models show none; character-billed speech shows one once the box has text.
  • “Waveform unavailable”. The preview could not decode the file; it still plays and converts. Use Format to make a WAV or MP3 copy for editing.
  • “This file is already MP3.” Pick another format, or turn on Save a new copy.
  • An edited M4A came back as WAV. There is no bundled AAC encoder; convert the WAV to MP3 with Format if you need it small.
  • Speed changed the pitch. That is how the audio workspace works; the video workspace keeps pitch.
  • An on-device model fails to load on Windows. Install Microsoft’s Visual C++ Redistributable from the link in the message, then retry.

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app