Skip to content
CSuite
Features · Chat

A chat that makes things.

Not just answers — assets. Describe what you need and the assistant generates images, video, voiceovers, music, and sound effects with the models you've chosen, then crops, trims, and converts them in follow-ups. Every result lands in your project folder as a real file.

Choose

One chat model, a full crew behind it.

Pick the model that runs the conversation — a frontier cloud model through Runware or a local one through Ollama — and it writes the replies too. Behind it, each kind of media has its own slot: your image model, your video model, your voices. The assistant calls whichever the request needs.
Generate

Ask in plain language, get the asset.

Describe the shot mid-conversation and the assistant routes it to your image model — no switching workspaces, no re-typing the prompt. The result appears inline and is saved to your project folder at the same moment.The image in the thread is a real, unretouched Flux 2 Pro generation from CSuite — the same sample shown on its catalog page.
Iterate

Follow-ups that actually follow up.

“Crop it to 1:1.” “Make it matte black.” The assistant knows “it” means the last image in the thread: simple edits dispatch straight to the local edit engine with no model round-trip, and creative revisions run image-to-image on the file it just made. Up to six chained tool steps per message — generate, crop, convert, done.
Multimodal

Every modality at the table.

One thread can carry a whole deliverable: a voiceover from your speech model, a music bed, a sound effect, a video clip that animates the image two messages up. Speech, music, and sound effects each route to their own model slot, so asking for a voiceover never runs your music model.
Catalog

Available models & providers.

Models you can chat with — hosts that can execute the assistant's tools: cloud models through Runware with your own key, and local models through Ollama, free and offline. (Replicate-only models stay available in the Text workspace.)

RunwareCloud · your API key
OllamaLocal · your hardware
GPT 5.6 Luna
Cloud
OpenAI's fastest, most affordable GPT-5.6 tier — for high-volume, latency-sensitive, and budget-conscious workloads.
Replicate · Runware
GPT 5.6 Terra
Cloud
OpenAI's balanced GPT-5.6 tier — GPT-5.5-level quality at roughly half the cost, for production workloads.
Replicate · Runware
GPT 5.6 Sol
Cloud
OpenAI's flagship GPT-5.6 tier — leads the family on every benchmark, for frontier coding, long-horizon agentic work, and research.
Replicate · Runware
Claude Fable 5
Cloud
Anthropic's Claude 5-family model for the most demanding reasoning, coding, and agentic work — with vision and a 128K output budget.
Replicate · Runware
Claude Opus 4.8
Cloud
Anthropic's most capable model — frontier reasoning, coding, and agentic tool use with vision.
Runware
Claude Sonnet 4.6
Cloud
Anthropic's balanced model for high-quality coding, reasoning, and agentic tool use — with a 1M-token context (beta) and vision.
Replicate · Runware
Claude Haiku 4.5
Cloud
Anthropic's fast, cost-efficient model with strong reasoning, coding, and tool use at low latency.
Replicate · Runware
Grok 4.3
Cloud
xAI's frontier Grok model with strong reasoning, coding, and agentic tool use, plus vision and real-time knowledge from X.
Runware
Gemini 3 Flash
Cloud
Fast, cost-efficient Gemini model for high-throughput text and multimodal tasks.
Replicate · Runware
V4 Flash
Cloud
DeepSeek's fast, cost-efficient V4 mixture-of-experts model (MIT) with long context and native tool calling.
Runware
Gemma 4 E2B
Local
Google's efficient 2B multimodal model supporting text, image, and audio input.
Ollama
Gemma 4 E4B
Local
Google's efficient 4B multimodal model — stronger reasoning than E2B with text, image, and audio.
Ollama
Gemma 4 12B
Local
Google's 12B multimodal Gemma 4 model with strong text and image understanding.
Ollama
Gemma 4 26B (MoE)
Local
Google's 26B mixture-of-experts model with long context and strong multimodal reasoning.
Ollama
Gemma 4 31B
Local
Google's flagship 31B dense model with 256K context and top-tier text and image understanding.
Ollama
Gemma 3 Med 4B
Local
Google's medical multimodal model fine-tuned for healthcare and biomedical tasks.
Ollama
Gemma 3 Med 27B
Local
Large medical variant for complex clinical reasoning and biomedical analysis.
Ollama
Mistral Small 4
Local
Mistral AI's unified small model merging reasoning, vision, and agentic coding, with configurable reasoning effort and tool use.
Ollama
Ministral 3 14B
Local
Mistral AI's 14B dense edge model (Apache 2.0) tuned for fast on-device reasoning and tool use.
Ollama
Ministral 3 8B
Local
Mistral AI's efficient 8B dense edge model (Apache 2.0) for on-device chat, reasoning, and tool use.
Ollama
Ministral 3 3B
Local
Mistral AI's compact 3B dense edge model (Apache 2.0) for fast on-device text generation and tool use.
Ollama
GPT-OSS 20B
Local
OpenAI's open-weight 20B reasoning model (Apache 2.0) with configurable reasoning effort and native tool calling — runs on 16 GB.
Ollama
Phi-4 14B
Local
Microsoft's 14B state-of-the-art open model with strong reasoning, math, and code.
Ollama
Phi-4 Mini 3.8B
Local
Microsoft's compact 3.8B model with a 128K context and native function calling.
Ollama
Phi-4 Reasoning 14B
Local
Microsoft's 14B reasoning model fine-tuned for chain-of-thought on math, science, and code.
Ollama
OLMo 3.1 32B
Local
Ai2's fully open 32B model — open weights, data, and training recipe — with long chain-of-thought reasoning.
Ollama
OLMo 3 7B
Local
Ai2's fully open 7B model with open weights, training data, and recipe.
Ollama
OLMo 3 32B
Local
Ai2's fully open 32B model with open weights, training data, and recipe.
Ollama
Laguna XS 2.1 (Q4)
Local
Poolside's 33B (3B active) MoE model for agentic coding and long-horizon local software work, with native reasoning and tool use.
Ollama
Laguna XS 2.1 (Q8)
Local
Poolside's 33B (3B active) MoE agentic-coding model — Q8 quantization for higher fidelity.
Ollama
Laguna XS 2.1 (BF16)
Local
Poolside's 33B (3B active) MoE agentic-coding model — full-precision BF16 weights.
Ollama
Qwen3 0.6B
Local
Alibaba's ultra-compact 0.6B Qwen3 model — runs in-process via HuggingFace or through Ollama, with tool use.
Ollama · Hugging Face
Llama 3.2 1B
Local
Meta's 1B instruction-tuned model — the smallest Llama, runnable in-process via HuggingFace or through Ollama.
Ollama · Hugging Face
Gemma 3 1B
Local
Google's compact 1B Gemma 3 model — runs in-process via HuggingFace or through Ollama.
Ollama · Hugging Face
LFM2 24B
Local
Liquid AI's 24B foundation model with efficient long-context reasoning.
Ollama
Llama 3.1 8B
Local
Meta's 8B instruction-tuned model with 128K context and strong general reasoning.
Ollama
Llama 3.2 3B
Local
Meta's compact 3B instruction-tuned model — fast on-device text generation.
Ollama · Hugging Face
Granite 4.1 3B
Local
IBM's compact 3B instruction-tuned model with multilingual support, tool use, and structured JSON output.
Ollama
Granite 4.1 8B
Local
IBM's 8B dense decoder-only foundation model for general instruction-following, RAG, and code tasks.
Ollama
Granite 4.1 30B
Local
IBM's flagship 30B dense foundation model — strong reasoning across business, coding, and multilingual tasks.
Ollama
Nemotron 3
Local
NVIDIA's multimodal Nemotron mixture-of-experts model unifying text and image understanding for enterprise Q&A, summarization, and document intelligence.
Ollama
Qwen3.6 27B
Local
Alibaba's 27B multimodal Qwen 3.6 model with a 256K context and strong vision-language reasoning.
Ollama
Qwen3.6 35B
Local
Alibaba's flagship 35B multimodal Qwen 3.6 model for top-tier vision-language reasoning.
Ollama
Qwen3.5 2B
Local
Alibaba's compact 2B multimodal Qwen model with vision input and a 256K context — fast on-device text+image reasoning.
Ollama
Qwen3.5 4B
Local
Alibaba's efficient 4B multimodal Qwen model with vision input and strong multilingual reasoning.
Ollama
Qwen3.5 9B
Local
Alibaba's 9B multimodal Qwen model balancing vision-language reasoning with on-device performance.
Ollama
LFM2.5 8B
Local
Liquid AI's efficient 8B Liquid Foundation Model tuned for long-context reasoning on-device.
Ollama
North Mini Code 1 (Q4)
Local
Cohere's compact code-specialized model tuned for fast local code generation, completion, and refactoring.
Ollama
Ornith 9B
Local
Deep Reinforce's 9B self-improving agentic-coding model (MIT) with 256K context and native tool use.
Ollama
Ornith 35B
Local
Deep Reinforce's flagship 35B self-improving agentic-coding model (MIT) with 256K context — state-of-the-art open-source coding for its size.
Ollama
R1 1.5B
Local
DeepSeek's compact 1.5B reasoning model distilled from Qwen2.5 — chain-of-thought on-device.
Ollama · Hugging Face
R1 8B
Local
DeepSeek's 8B reasoning model distilled from Qwen3 (R1-0528) with strong math, code, and logic.
Ollama
R1 14B
Local
DeepSeek's 14B reasoning model distilled from Qwen2.5 with exceptional math and coding benchmarks.
Ollama
R1 32B
Local
DeepSeek's 32B reasoning model distilled from Qwen2.5 for top-tier chain-of-thought reasoning.
Ollama
Qwen3 4B
Local
Alibaba's compact 4B instruction model with strong reasoning and multilingual capabilities.
Ollama · Hugging Face
FAQ

Chat questions, answered.

How the multimodal chat works in CSuite — capable models, tools and edits, files, offline use, and costs.

Launch offer · 50% off

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

$98$49
Pricing

Secure checkout via Stripe. Already have a license? Download the app