Skip to content
CSuite
Features · Chat

A chat that makes things.

Not just answers — assets. Describe what you need and the assistant generates images, video, voiceovers, music, and sound effects with the models you've chosen, then crops, trims, and converts them in follow-ups. Every result lands in your project folder as a real file.

Choose

One chat model, a full crew behind it.

Pick the model that runs the conversation — a frontier cloud model through Runware or a local one through Ollama — and it writes the replies too. Behind it, each kind of media has its own slot: your image model, your video model, your voices. The assistant calls whichever the request needs.
Generate

Ask in plain language, get the asset.

Describe the shot mid-conversation and the assistant routes it to your image model — no switching workspaces, no re-typing the prompt. The result appears inline and is saved to your project folder at the same moment.The image in the thread is a real, unretouched Flux 2 Pro generation from CSuite — the same sample shown on its catalog page.
Iterate

Follow-ups that actually follow up.

“Crop it to 1:1.” “Make it matte black.” The assistant knows “it” means the last image in the thread: simple edits dispatch straight to the local edit engine with no model round-trip, and creative revisions run image-to-image on the file it just made. Up to six chained tool steps per message — generate, crop, convert, done.
Multimodal

Every modality at the table.

One thread can carry a whole deliverable: a voiceover from your speech model, a music bed, a sound effect, a video clip that animates the image two messages up. Speech, music, and sound effects each route to their own model slot, so asking for a voiceover never runs your music model.
Workspace

Bring your own material.

A thread isn't only what the assistant makes — it's what you hand it, and how easily you can take another run at an answer.

Attach documentsDrop in PDFs, Markdown, HTML, or plain text and ask questions about them. The text is extracted and injected into the prompt, and stays available for follow-ups later in the thread.
Attach an imageOn a vision-capable model, send a picture and ask what's in it — the image goes to the model as real image input, whether it's running locally or in the cloud.
Run a workflowAsk for one of your saved workflows by name and it runs headlessly, posting each saved output back into the thread as an attachment.
RegenerateDidn't like the answer? Regenerate the latest reply and the turn re-runs, reusing whatever image you had attached.
Edit and resendRewrite your last message instead of arguing with a misread one — everything after it is cleared and the turn runs again from the corrected version.
Copy it outCopy a single message, or the entire conversation as Markdown, ready to paste into a doc or an issue.
Catalog

Available models & providers.

Models you can chat with. Because chat runs on tool calls, this list is narrower than the full text catalog: a model needs native function calling and a host that can execute it, which means Runware in the cloud with your own key, or Ollama locally, free and offline. Text models that miss either requirement stay available in the Text workspace.

RunwareCloud · your API key
OllamaLocal · your hardware
GPT 5.4
Cloud
OpenAI's frontier model with advanced reasoning and broad multimodal capabilities.
Runware · Replicate
GPT 5.6 Luna
Cloud
OpenAI's fastest, most affordable GPT-5.6 tier — for high-volume, latency-sensitive, and budget-conscious workloads.
Runware · Replicate
GPT 5.6 Terra
Cloud
OpenAI's balanced GPT-5.6 tier — GPT-5.5-level quality at roughly half the cost, for production workloads.
Runware · Replicate
GPT 5.6 Sol
Cloud
OpenAI's flagship GPT-5.6 tier — leads the family on every benchmark, for frontier coding, long-horizon agentic work, and research.
Runware · Replicate
Kimi K2.6
Cloud
Moonshot's frontier open-weight model (1T parameters) built for long-horizon agentic coding, with a 262K context and native vision.
Runware · Replicate
Claude Fable 5
Cloud
Anthropic's Claude 5-family model for the most demanding reasoning, coding, and agentic work — with vision and a 128K output budget.
Runware · Replicate
Claude Opus 4.8
Cloud
Anthropic's most capable model — frontier reasoning, coding, and agentic tool use with vision.
Runware
Claude Sonnet 4.6
Cloud
Anthropic's balanced model for high-quality coding, reasoning, and agentic tool use — with a 1M-token context (beta) and vision.
Runware · Replicate
Claude Haiku 4.5
Cloud
Anthropic's fast, cost-efficient model with strong reasoning, coding, and tool use at low latency.
Runware · Replicate
Grok 4.3
Cloud
xAI's frontier Grok model with strong reasoning, coding, and agentic tool use, plus vision and real-time knowledge from X.
Runware
V4 Flash
Cloud
DeepSeek's fast, cost-efficient V4 mixture-of-experts model (MIT) with long context and native tool calling.
Runware
V4 Pro
Cloud
DeepSeek's high-capability V4 model (MIT) with a 1M-token context, dual thinking modes, and stronger agentic performance than V4 Flash.
Runware
Gemma 4 E2B
Local
Google's efficient 2B multimodal model supporting text, image, and audio input.
Ollama
Gemma 4 E4B
Local
Google's efficient 4B multimodal model — stronger reasoning than E2B with text, image, and audio.
Ollama
Gemma 4 12B
Local
Google's 12B multimodal Gemma 4 model with strong text and image understanding.
Ollama
Gemma 4 26B (MoE)
Local
Google's 26B mixture-of-experts model with long context and strong multimodal reasoning.
Ollama
Gemma 4 31B
Local
Google's flagship 31B dense model with 256K context and top-tier text and image understanding.
Ollama
Mistral Small 3.2 24B
Local
Mistral's 24B instruction-tuned model with vision input, robust function calling, and a 128K context.
Ollama
Ministral 3 14B
Local
Mistral's 14B dense edge model (Apache 2.0) tuned for fast on-device reasoning and tool use.
Ollama
Ministral 3 8B
Local
Mistral's efficient 8B dense edge model (Apache 2.0) for on-device chat, reasoning, and tool use.
Ollama
Ministral 3 3B
Local
Mistral's compact 3B dense edge model (Apache 2.0) for fast on-device text generation and tool use.
Ollama
GPT-OSS 20B
Local
OpenAI's open-weight 20B reasoning model (Apache 2.0) with configurable reasoning effort and native tool calling — runs on 16 GB.
Ollama
Phi-4 Mini 3.8B
Local
Microsoft's compact 3.8B model with a 128K context and native function calling.
Ollama
Laguna XS 2.1 (Q4)
Local
Poolside's 33B (3B active) MoE model for agentic coding and long-horizon local software work, with native reasoning and tool use.
Ollama
Laguna XS 2.1 (Q8)
Local
Poolside's 33B (3B active) MoE agentic-coding model — Q8 quantization for higher fidelity.
Ollama
Laguna XS 2.1 (BF16)
Local
Poolside's 33B (3B active) MoE agentic-coding model — full-precision BF16 weights.
Ollama
Qwen3 0.6B
Local
Alibaba's ultra-compact 0.6B Qwen3 model — runs in-process via HuggingFace or through Ollama, with tool use.
Ollama · Hugging Face
Llama 3.2 1B
Local
Meta's 1B instruction-tuned model — the smallest Llama, runnable in-process via HuggingFace or through Ollama.
Ollama · Hugging Face
Gemma 3 1B
Local
Google's compact 1B Gemma 3 model — runs in-process via HuggingFace or through Ollama.
Ollama · Hugging Face
LFM2 24B
Local
Liquid's 24B foundation model with efficient long-context reasoning.
Ollama
Llama 3.1 8B
Local
Meta's 8B instruction-tuned model with 128K context and strong general reasoning.
Ollama
Llama 3.2 3B
Local
Meta's compact 3B instruction-tuned model — fast on-device text generation.
Ollama · Hugging Face
Granite 4.1 3B
Local
IBM's compact 3B instruction-tuned model with multilingual support, tool use, and structured JSON output.
Ollama
Granite 4.1 8B
Local
IBM's 8B dense decoder-only foundation model for general instruction-following, RAG, and code tasks.
Ollama
Granite 4.1 30B
Local
IBM's flagship 30B dense foundation model — strong reasoning across business, coding, and multilingual tasks.
Ollama
Granite 4.2 3B
Local
IBM's compact 3B instruction-tuned model with multilingual support, tool use, and structured JSON output.
Ollama
Granite 4.2 8B
Local
IBM's 8B dense foundation model for general instruction-following, RAG, and code tasks, with native tool use.
Ollama
Granite 4.2 30B
Local
IBM's flagship 30B foundation model — strong reasoning across business, coding, and multilingual tasks.
Ollama
Muse Glimmer
Local
Meta Superintelligence Labs' 30B multimodal model for always-on local agents — tuned for tool use, long tasks, and failure recovery.
Ollama
Nemotron 3.5 Lightning
Local
NVIDIA's 30B mixture-of-experts model with 3B active parameters, built as the execution layer for always-on agents.
Ollama
Nemotron 3
Local
NVIDIA's multimodal Nemotron mixture-of-experts model unifying text and image understanding for enterprise Q&A, summarization, and document intelligence.
Ollama
Qwen3.8 27B
Local
Alibaba's 27B multimodal Qwen 3.8 model with a 256K context — substantial gains in coding, research, and long-horizon agentic tasks, with native image and video understanding.
Ollama
Qwen3.6 27B
Local
Alibaba's 27B multimodal Qwen 3.6 model with a 256K context and strong vision-language reasoning.
Ollama
Qwen3.6 35B
Local
Alibaba's flagship 35B multimodal Qwen 3.6 model for top-tier vision-language reasoning.
Ollama
Qwen3.5 2B
Local
Alibaba's compact 2B multimodal Qwen model with vision input and a 256K context — fast on-device text+image reasoning.
Ollama
Qwen3.5 4B
Local
Alibaba's efficient 4B multimodal Qwen model with vision input and strong multilingual reasoning.
Ollama
Qwen3.5 9B
Local
Alibaba's 9B multimodal Qwen model balancing vision-language reasoning with on-device performance.
Ollama
LFM2.5 8B
Local
Liquid's efficient 8B Liquid Foundation Model tuned for long-context reasoning on-device.
Ollama
North Mini Code 1 (Q4)
Local
Cohere's compact code-specialized model tuned for fast local code generation, completion, and refactoring.
Ollama
Ornith 9B
Local
Deep Reinforce's 9B self-improving agentic-coding model (MIT) with 256K context and native tool use.
Ollama
Ornith 35B
Local
Deep Reinforce's flagship 35B self-improving agentic-coding model (MIT) with 256K context — state-of-the-art open-source coding for its size.
Ollama
Ornith 1.5 9B
Local
Deep Reinforce's 9B self-improving model (MIT) with 256K context, image input, and native tool use, tuned for reasoning, coding, and agentic work.
Ollama
Ornith 1.5 35B
Local
Deep Reinforce's 35B MoE self-improving model (MIT) with 256K context, image input, and native tool use, strong across reasoning, coding, and agentic tasks.
Ollama
R1 1.5B
Local
DeepSeek's compact 1.5B reasoning model distilled from Qwen2.5 — chain-of-thought on-device.
Ollama · Hugging Face
R1 8B
Local
DeepSeek's 8B reasoning model distilled from Qwen3 (R1-0528) with strong math, code, and logic.
Ollama
R1 14B
Local
DeepSeek's 14B reasoning model distilled from Qwen2.5 with exceptional math and coding benchmarks.
Ollama
R1 32B
Local
DeepSeek's 32B reasoning model distilled from Qwen2.5 for top-tier chain-of-thought reasoning.
Ollama
Qwen3 4B
Local
Alibaba's compact 4B instruction model with strong reasoning and multilingual capabilities.
Ollama · Hugging Face
FAQ

Chat questions, answered.

How the multimodal chat works in CSuite — capable models, tools and edits, files, offline use, and costs.

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app