Not just answers — assets. Describe what you need and the assistant generates images, video, voiceovers, music, and sound effects with the models you've chosen, then crops, trims, and converts them in follow-ups. Every result lands in your project folder as a real file.
Pick the model that runs the conversation — a frontier cloud model through Runware or a local one through Ollama — and it writes the replies too. Behind it, each kind of media has its own slot: your image model, your video model, your voices. The assistant calls whichever the request needs.
Models for this conversationTool-capable hostsChatGPT 5.6 SolRunware · CloudGemma 4 12BOllama · LocalThe chat model also writes text — no separate text slot.More modelsused when the assistant generates mediaImageFlux 2 ProVideoVeo 3.1SpeechEleven v3MusicLyria 3Sound effectsSFX 1.5
Generate
Ask in plain language, get the asset.
Describe the shot mid-conversation and the assistant routes it to your image model — no switching workspaces, no re-typing the prompt. The result appears inline and is saved to your project folder at the same moment.The image in the thread is a real, unretouched Flux 2 Pro generation from CSuite — the same sample shown on its catalog page.
Press kit assetsGPT 5.6 Sol · RunwareMake the hero shot for the press kit: a brushed-titanium espresso tamper on wet black slate, hard light, 16:9.image_generateFlux 2 ProSaved · espresso-tamper.pngReal output · Flux 2 ProMessage…Send
Iterate
Follow-ups that actually follow up.
“Crop it to 1:1.” “Make it matte black.” The assistant knows “it” means the last image in the thread: simple edits dispatch straight to the local edit engine with no model round-trip, and creative revisions run image-to-image on the file it just made. Up to six chained tool steps per message — generate, crop, convert, done.
Press kit assetsGPT 5.6 Sol · RunwareCrop it to 1:1 for the product grid.image_cropdirect edit · no model callSaved · espresso-tamper-1x1.pngNow a matte-black variant of the tamper, same lighting.image_updateGenerating from the last image…Message…Generating
Multimodal
Every modality at the table.
One thread can carry a whole deliverable: a voiceover from your speech model, a music bed, a sound effect, a video clip that animates the image two messages up. Speech, music, and sound effects each route to their own model slot, so asking for a voiceover never runs your music model.
Press kit assetsGPT 5.6 Sol · RunwareRead the tagline as a warm voiceover: “Helio. Espresso, dialed in.”speech_generateEleven v30:03Saved · tagline-voiceover.mp3And a 6-second clip of steam rising off the espresso for the site loop.video_generateRendering with Veo 3.1…imagevideospeechmusicsfxedits
Catalog
Available models & providers.
Models you can chat with — hosts that can execute the assistant's tools: cloud models through Runware with your own key, and local models through Ollama, free and offline. (Replicate-only models stay available in the Text workspace.)
How the multimodal chat works in CSuite — capable models, tools and edits, files, offline use, and costs.
Chat needs a host that can execute tools, so the picker offers cloud models through Runware (the GPT 5.6 family, Claude, Gemini, Grok, and more) and local models through Ollama (Gemma, Qwen, Llama, and other open-weight models). Replicate-hosted models can't run chat tools and aren't selectable here — they remain available in the Text workspace.
Beyond conversation, it generates media with the models you've assigned per slot — images (including image-to-image updates of the last image), video clips (using up to two prior images as start and end frames), text-to-speech, music, and sound effects — and performs edits: crop or resize images, trim or convert audio, and crop, resize, or transcode video.
Edits target the most recent file of that kind in the conversation, so "crop it to 1:1" just works. Obvious requests are dispatched directly to the local edit engine without a model round-trip, and the assistant can chain up to six tool steps in one turn — generate, then crop, then convert, from a single message.
Into your project folder, as regular files — the thumbnails and players in the thread point at those files, not at copies trapped in a chat log. Anything the assistant makes is immediately usable in the Image, Audio, and Video workspaces, or in any app on your machine.
Yes — pick a local Ollama model and the conversation, tool-calling, and deterministic edits all run on your machine. Generating media in the same thread uses whatever models you've assigned to those slots: local ones stay offline, cloud ones need their platform key.
The chat model itself — it doubles as the text model, so there's no separate text slot to configure. Image, video, speech, music, and sound-effect generation each have their own model slot under "More models", and if a request needs a slot you haven't configured, the assistant pauses and opens the picker at exactly that slot.
Cloud chat models are billed per token by Runware directly to your account, and any media the assistant generates is billed by the model that made it — with no CSuite markup or metering. Local models cost nothing to run. The app's analytics show spend per provider.
Locally, in the app's database on your machine — conversations, messages, and attachment references never touch our servers. Delete a conversation and it's gone from your disk, not archived in a cloud.
Launch offer · 50% off
One-time payment. Yours forever.
No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.