Skip to content
CSuite
Features · Text

A real editor, with a model in the loop.

Not a chat window you copy-paste from — a document editor where AI works in place: draft from a prompt, rewrite a selection, accept an autocomplete, drop in a generated image. Frontier cloud models on your own keys, or open-weight models running free and offline on your machine.

Choose

Frontier cloud or free local.

Claude, GPT, Gemini, and Grok through your own API keys when you want the best prose — or Gemma, Qwen, Llama, and Mistral running locally through Ollama and the bundled Hugging Face runtime when you want free, offline, and private. Same editor, same prompt history, switch per document.
Draft

Prompt in, draft streamed onto the page.

Generation streams straight into the document — appending, replacing a selection, or inserting at the caret depending on what you asked. Untitled files name themselves from your first prompt, and models with adjustable reasoning expose an effort control next to the picker.The paragraph in the mockup is a real, unedited Claude Sonnet 4.6 answer to the exact prompt shown — 46 words against a 50-word cap. Every text model in the catalog answers the same prompts on its detail page.
Revise

Fix the sentence, not the document.

Select a clumsy passage and a floating toolbar appears: Shorten, Expand, or Improve — the rewrite streams in over exactly that range while everything around it stays put. The selection stays visibly highlighted while you type a follow-up, so you never lose track of what's being changed.
Editor

An editor first, AI second.

Underneath the AI is a proper writing tool — real formats, real formatting, and files that stay yours. The model helps in place instead of replacing your editor.

Rich text, Markdown, or plain text

Create .html, .md, or .txt documents and format with a full ribbon — headings, bold and italics, lists, case transforms. Everything autosaves as you type, to ordinary files any other editor can open.

Autocomplete that waits for you

Pause for a moment and a greyed continuation appears at the caret — Tab accepts, Shift+Right tries another, Esc dismisses. It never touches the file until you accept, uses whichever model you've picked (local models keep it fully private), and switches off with one toggle.

Generate images inside the doc

Select a phrase, hit Generate image, and a placeholder resolves in place using your image model — the espresso-tamper shot here is a real Flux 2 Pro generation. You can also drag images, audio, and video straight from your project or the OS into the page.
Catalog

Available models & providers.

Every text model in the catalog, in one picker. Cloud models run through Replicate and Runware with your own keys; open-weight models run locally through Ollama and the bundled Hugging Face runtime — free and offline.

ReplicateCloud · your API key
RunwareCloud · your API key
OllamaLocal · your hardware
Hugging FaceLocal · your hardware
GPT 5.4
Cloud
OpenAI's frontier model with advanced reasoning and broad multimodal capabilities.
Replicate
GPT 5.6 Luna
Cloud
OpenAI's fastest, most affordable GPT-5.6 tier — for high-volume, latency-sensitive, and budget-conscious workloads.
Replicate · Runware
GPT 5.6 Terra
Cloud
OpenAI's balanced GPT-5.6 tier — GPT-5.5-level quality at roughly half the cost, for production workloads.
Replicate · Runware
GPT 5.6 Sol
Cloud
OpenAI's flagship GPT-5.6 tier — leads the family on every benchmark, for frontier coding, long-horizon agentic work, and research.
Replicate · Runware
Claude Fable 5
Cloud
Anthropic's Claude 5-family model for the most demanding reasoning, coding, and agentic work — with vision and a 128K output budget.
Replicate · Runware
Claude Opus 4.8
Cloud
Anthropic's most capable model — frontier reasoning, coding, and agentic tool use with vision.
Runware
Claude Sonnet 4.6
Cloud
Anthropic's balanced model for high-quality coding, reasoning, and agentic tool use — with a 1M-token context (beta) and vision.
Replicate · Runware
Claude Haiku 4.5
Cloud
Anthropic's fast, cost-efficient model with strong reasoning, coding, and tool use at low latency.
Replicate · Runware
Grok 4.3
Cloud
xAI's frontier Grok model with strong reasoning, coding, and agentic tool use, plus vision and real-time knowledge from X.
Runware
Gemini 3.5 Flash
Cloud
Google's latest fast Gemini 3 model: frontier-level reasoning at Flash-level latency and cost, tuned for agentic workflows and iterative coding.
Replicate
Gemini 3.1 Pro
Cloud
Google's most capable multimodal model with deep reasoning and 1M token context.
Replicate
Gemini 3 Flash
Cloud
Fast, cost-efficient Gemini model for high-throughput text and multimodal tasks.
Replicate · Runware
Gemini 2.5 Flash
Cloud
Balanced Gemini model with strong reasoning and a 1M token context window.
Replicate
V4 Flash
Cloud
DeepSeek's fast, cost-efficient V4 mixture-of-experts model (MIT) with long context and native tool calling.
Runware
Gemma 4 E2B
Local
Google's efficient 2B multimodal model supporting text, image, and audio input.
Ollama
Gemma 4 E4B
Local
Google's efficient 4B multimodal model — stronger reasoning than E2B with text, image, and audio.
Ollama
Gemma 4 12B
Local
Google's 12B multimodal Gemma 4 model with strong text and image understanding.
Ollama
Gemma 4 26B (MoE)
Local
Google's 26B mixture-of-experts model with long context and strong multimodal reasoning.
Ollama
Gemma 4 31B
Local
Google's flagship 31B dense model with 256K context and top-tier text and image understanding.
Ollama
Gemma 3 Med 4B
Local
Google's medical multimodal model fine-tuned for healthcare and biomedical tasks.
Ollama
Gemma 3 Med 27B
Local
Large medical variant for complex clinical reasoning and biomedical analysis.
Ollama
Mistral Small 4
Local
Mistral AI's unified small model merging reasoning, vision, and agentic coding, with configurable reasoning effort and tool use.
Ollama
Ministral 3 14B
Local
Mistral AI's 14B dense edge model (Apache 2.0) tuned for fast on-device reasoning and tool use.
Ollama
Ministral 3 8B
Local
Mistral AI's efficient 8B dense edge model (Apache 2.0) for on-device chat, reasoning, and tool use.
Ollama
Ministral 3 3B
Local
Mistral AI's compact 3B dense edge model (Apache 2.0) for fast on-device text generation and tool use.
Ollama
GPT-OSS 20B
Local
OpenAI's open-weight 20B reasoning model (Apache 2.0) with configurable reasoning effort and native tool calling — runs on 16 GB.
Ollama
Phi-4 14B
Local
Microsoft's 14B state-of-the-art open model with strong reasoning, math, and code.
Ollama
Phi-4 Mini 3.8B
Local
Microsoft's compact 3.8B model with a 128K context and native function calling.
Ollama
Phi-4 Reasoning 14B
Local
Microsoft's 14B reasoning model fine-tuned for chain-of-thought on math, science, and code.
Ollama
OLMo 3.1 32B
Local
Ai2's fully open 32B model — open weights, data, and training recipe — with long chain-of-thought reasoning.
Ollama
OLMo 3 7B
Local
Ai2's fully open 7B model with open weights, training data, and recipe.
Ollama
OLMo 3 32B
Local
Ai2's fully open 32B model with open weights, training data, and recipe.
Ollama
Laguna XS 2.1 (Q4)
Local
Poolside's 33B (3B active) MoE model for agentic coding and long-horizon local software work, with native reasoning and tool use.
Ollama
Laguna XS 2.1 (Q8)
Local
Poolside's 33B (3B active) MoE agentic-coding model — Q8 quantization for higher fidelity.
Ollama
Laguna XS 2.1 (BF16)
Local
Poolside's 33B (3B active) MoE agentic-coding model — full-precision BF16 weights.
Ollama
Qwen3 0.6B
Local
Alibaba's ultra-compact 0.6B Qwen3 model — runs in-process via HuggingFace or through Ollama, with tool use.
Ollama · Hugging Face
Llama 3.2 1B
Local
Meta's 1B instruction-tuned model — the smallest Llama, runnable in-process via HuggingFace or through Ollama.
Ollama · Hugging Face
Gemma 3 1B
Local
Google's compact 1B Gemma 3 model — runs in-process via HuggingFace or through Ollama.
Ollama · Hugging Face
LFM2 24B
Local
Liquid AI's 24B foundation model with efficient long-context reasoning.
Ollama
Llama 3.1 8B
Local
Meta's 8B instruction-tuned model with 128K context and strong general reasoning.
Ollama
Llama 3.2 3B
Local
Meta's compact 3B instruction-tuned model — fast on-device text generation.
Ollama · Hugging Face
Granite 4.1 3B
Local
IBM's compact 3B instruction-tuned model with multilingual support, tool use, and structured JSON output.
Ollama
Granite 4.1 8B
Local
IBM's 8B dense decoder-only foundation model for general instruction-following, RAG, and code tasks.
Ollama
Granite 4.1 30B
Local
IBM's flagship 30B dense foundation model — strong reasoning across business, coding, and multilingual tasks.
Ollama
Nemotron 3
Local
NVIDIA's multimodal Nemotron mixture-of-experts model unifying text and image understanding for enterprise Q&A, summarization, and document intelligence.
Ollama
Qwen3.6 27B
Local
Alibaba's 27B multimodal Qwen 3.6 model with a 256K context and strong vision-language reasoning.
Ollama
Qwen3.6 35B
Local
Alibaba's flagship 35B multimodal Qwen 3.6 model for top-tier vision-language reasoning.
Ollama
Qwen3.5 2B
Local
Alibaba's compact 2B multimodal Qwen model with vision input and a 256K context — fast on-device text+image reasoning.
Ollama
Qwen3.5 4B
Local
Alibaba's efficient 4B multimodal Qwen model with vision input and strong multilingual reasoning.
Ollama
Qwen3.5 9B
Local
Alibaba's 9B multimodal Qwen model balancing vision-language reasoning with on-device performance.
Ollama
LFM2.5 8B
Local
Liquid AI's efficient 8B Liquid Foundation Model tuned for long-context reasoning on-device.
Ollama
North Mini Code 1 (Q4)
Local
Cohere's compact code-specialized model tuned for fast local code generation, completion, and refactoring.
Ollama
Ornith 9B
Local
Deep Reinforce's 9B self-improving agentic-coding model (MIT) with 256K context and native tool use.
Ollama
Ornith 35B
Local
Deep Reinforce's flagship 35B self-improving agentic-coding model (MIT) with 256K context — state-of-the-art open-source coding for its size.
Ollama
R1 1.5B
Local
DeepSeek's compact 1.5B reasoning model distilled from Qwen2.5 — chain-of-thought on-device.
Ollama · Hugging Face
R1 8B
Local
DeepSeek's 8B reasoning model distilled from Qwen3 (R1-0528) with strong math, code, and logic.
Ollama
R1 14B
Local
DeepSeek's 14B reasoning model distilled from Qwen2.5 with exceptional math and coding benchmarks.
Ollama
R1 32B
Local
DeepSeek's 32B reasoning model distilled from Qwen2.5 for top-tier chain-of-thought reasoning.
Ollama
Qwen3 4B
Local
Alibaba's compact 4B instruction model with strong reasoning and multilingual capabilities.
Ollama · Hugging Face
FAQ

Text questions, answered.

How writing with AI works in CSuite — models, offline use, formats, selection actions, autocomplete, and costs.

Launch offer · 50% off

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

$98$49
Pricing

Secure checkout via Stripe. Already have a license? Download the app