Not just answers — assets. Describe what you need and the assistant generates images, video, voiceovers, music, and sound effects with the models you've chosen, then crops, trims, and converts them in follow-ups. Every result lands in your project folder as a real file.
Pick the model that runs the conversation — a frontier cloud model through Runware or a local one through Ollama — and it writes the replies too. Behind it, each kind of media has its own slot: your image model, your video model, your voices. The assistant calls whichever the request needs.
Models for this conversationTool-capable hostsChatGPT 5.6 SolRunware · CloudGemma 4 12BOllama · LocalThe chat model also writes text — no separate text slot.More modelsused when the assistant generates mediaImageFlux 2 ProVideoVeo 3.1SpeechEleven v3MusicLyria 3Sound effectsSFX 1.5
Generate
Ask in plain language, get the asset.
Describe the shot mid-conversation and the assistant routes it to your image model — no switching workspaces, no re-typing the prompt. The result appears inline and is saved to your project folder at the same moment.The image in the thread is a real, unretouched Flux 2 Pro generation from CSuite — the same sample shown on its catalog page.
Press kit assetsGPT 5.6 Sol · RunwareMake the hero shot for the press kit: a brushed-titanium espresso tamper on wet black slate, hard light, 16:9.image_generateFlux 2 ProSaved · espresso-tamper.pngReal output · Flux 2 ProMessage…Send
Iterate
Follow-ups that actually follow up.
“Crop it to 1:1.” “Make it matte black.” The assistant knows “it” means the last image in the thread: simple edits dispatch straight to the local edit engine with no model round-trip, and creative revisions run image-to-image on the file it just made. Up to six chained tool steps per message — generate, crop, convert, done.
Press kit assetsGPT 5.6 Sol · RunwareCrop it to 1:1 for the product grid.image_cropdirect edit · no model callSaved · espresso-tamper-1x1.pngNow a matte-black variant of the tamper, same lighting.image_updateGenerating from the last image…Message…Generating
Multimodal
Every modality at the table.
One thread can carry a whole deliverable: a voiceover from your speech model, a music bed, a sound effect, a video clip that animates the image two messages up. Speech, music, and sound effects each route to their own model slot, so asking for a voiceover never runs your music model.
Press kit assetsGPT 5.6 Sol · RunwareRead the tagline as a warm voiceover: “Helio. Espresso, dialed in.”speech_generateEleven v30:03Saved · tagline-voiceover.mp3And a 6-second clip of steam rising off the espresso for the site loop.video_generateRendering with Veo 3.1…imagevideospeechmusicsfxedits
Workspace
Bring your own material.
A thread isn't only what the assistant makes — it's what you hand it, and how easily you can take another run at an answer.
Attach documentsDrop in PDFs, Markdown, HTML, or plain text and ask questions about them. The text is extracted and injected into the prompt, and stays available for follow-ups later in the thread.
Attach an imageOn a vision-capable model, send a picture and ask what's in it — the image goes to the model as real image input, whether it's running locally or in the cloud.
Run a workflowAsk for one of your saved workflows by name and it runs headlessly, posting each saved output back into the thread as an attachment.
RegenerateDidn't like the answer? Regenerate the latest reply and the turn re-runs, reusing whatever image you had attached.
Edit and resendRewrite your last message instead of arguing with a misread one — everything after it is cleared and the turn runs again from the corrected version.
Copy it outCopy a single message, or the entire conversation as Markdown, ready to paste into a doc or an issue.
Catalog
Available models & providers.
Models you can chat with. Because chat runs on tool calls, this list is narrower than the full text catalog: a model needs native function calling and a host that can execute it, which means Runware in the cloud with your own key, or Ollama locally, free and offline. Text models that miss either requirement stay available in the Text workspace.
How the multimodal chat works in CSuite — capable models, tools and edits, files, offline use, and costs.
Chat is tool-calling only, so a model has to clear two bars: its host must be able to execute tool calls, and the model itself must support function calling. That comes to cloud models on Runware whose API does tools (the GPT 5.4 and 5.6 tiers, Claude, Grok, Kimi K2.6, DeepSeek V4) and local Ollama tags flagged for tool calling (Gemma, Qwen, Llama, Granite, and other open-weight models). Everything that misses either bar stays available in the Text workspace instead: Replicate-hosted models, whose prediction API has no tool surface at all; Gemini on Runware, where the tool-continuation round is broken upstream; and open-weight tags without native tool calling, such as Phi-4 and OLMo.
Beyond conversation, it generates media with the models you've assigned per slot — images (including image-to-image updates of the last image), video clips (using up to two prior images as start and end frames), text-to-speech, music, and sound effects — and performs edits: crop or resize images, trim or convert audio, and crop, resize, or transcode video. It can also run one of your saved workflows by name and post the results back into the thread.
Yes. Attach PDFs, Markdown, HTML, or plain text documents and ask questions about them — the text is extracted and carried into the conversation, so follow-ups keep working further down the thread. On a vision-capable model you can attach an image too, and it's sent as real image input rather than a description.
Regenerate the newest reply and the turn re-runs, or edit your own last message and resend — everything after it is cleared and the corrected version runs fresh. Stop genuinely cancels mid-answer, aborting the provider call rather than just hiding the result. Any message, or the whole conversation, copies out as Markdown.
Edits target the most recent file of that kind in the conversation, so "crop it to 1:1" just works. Obvious requests are dispatched directly to the local edit engine without a model round-trip, and the assistant can chain up to six tool steps in one turn — generate, then crop, then convert, from a single message.
Into your project folder, as regular files — the thumbnails and players in the thread point at those files, not at copies trapped in a chat log. Anything the assistant makes is immediately usable in the Image, Audio, and Video workspaces, or in any app on your machine.
Yes — pick a local Ollama model and the conversation, tool-calling, and deterministic edits all run on your machine. Generating media in the same thread uses whatever models you've assigned to those slots: local ones stay offline, cloud ones need their platform key.
The chat model itself — it doubles as the text model, so there's no separate text slot to configure. Image, video, speech, music, and sound-effect generation each have their own model slot under "More models", and if a request needs a slot you haven't configured, the assistant pauses and opens the picker at exactly that slot.
Cloud chat models are billed per token by Runware directly to your account, and any media the assistant generates is billed by the model that made it — with no CSuite markup or metering. Local models cost nothing to run. The app's analytics show spend per provider.
Locally, in the app's database on your machine — conversations, messages, and attachment references never touch our servers. Delete a conversation and it's gone from your disk, not archived in a cloud.
One-time payment. Yours forever.
No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.