Models.
CSuite supports models across every modality: text, image, audio and video. Run them locally on your machine or in the cloud with your own API keys, private by design. Below are the models bundled with the app out of the box, and you can always add your favorite model inside the app.
Google and Google DeepMind build the Gemini family of multimodal models, the Imagen and Nano Banana image models, the Lyria music models, and the Veo video models.
OpenAI builds GPT, DALL·E, the Sora family, and the open-weight gpt-oss models, and has been a central force behind the modern wave of generative AI.
Alibaba's Tongyi research group publishes the Wan video models and the Qwen family of language models.
ByteDance's AI labs ship the Seedream image models and the Seedance video models, among other Doubao-branded research.
Kuaishou builds the Kling family of cinematic video models — high-resolution clips with native audio, lip sync, and multi-shot storytelling.
Anthropic builds the Claude family of large language models, focused on safety, helpfulness, and honesty.
Meta's FAIR and GenAI labs build the open-weight Llama family of language models — among the most widely used open models for local inference.
IBM Research builds the open-weight Granite family — efficient foundation models tuned for enterprise instruction-following, RAG, code, and tool use.
DeepSeek is a Chinese AI lab whose open-weight models — from the R1 reasoning distills to the V4 mixture-of-experts family — bring frontier-grade reasoning under permissive (MIT) licensing.
Fish Audio builds expressive speech models with inline emotion cues, multi-speaker synthesis, and zero-shot voice cloning across 80+ languages.
ElevenLabs builds best-in-class voice AI — realistic text-to-speech, voice cloning, music, and sound effects.
MiniMax is a Chinese AI lab whose speech and music models are widely used for multilingual voiceover and generative music.
Runway builds the Gen family of generative video models, pioneering high-fidelity text-to-video and image-to-video for filmmakers and creators.
xAI builds the Grok family — frontier reasoning LLMs plus Grok Imagine video and voice models, integrated with real-time knowledge from X.
Mistral is a Paris-based lab that ships fast, efficient open-weight language models with strong reasoning per parameter.
Black Forest Labs makes the FLUX family of image models — state-of-the-art open-weight image generation.
Deep Reinforce builds the Ornith family — self-improving open-weight (MIT) models for agentic coding, tuned with reinforcement learning on coding and terminal benchmarks.
Microsoft Research builds the open-weight Phi family of small language models — efficient, reasoning-focused models that punch above their parameter count.
Recraft builds design-first image models — the V4 family, tuned for brand-consistent typography, layout and colour control, plus vector tiers that generate editable SVG artwork rather than raster output.
Pruna AI builds performance-focused image and video models, including the P-Video line, which generates clips with native audio and offers a draft mode for fast, low-cost iteration.
The Allen Institute for AI (Ai2) builds the OLMo family — fully open language models with open weights, training data, and recipes.
Poolside AI builds foundation models for software engineering — the Laguna family, tuned for agentic coding and long-horizon local development.
Lightricks builds the LTX family of video models — fast, high-resolution generation with native synchronized audio, first-to-last-frame control, and the open-weight LTX-Video research line.
Liquid develops Liquid Foundation Models (LFM) — a new architecture optimized for efficient long-context reasoning.
NVIDIA builds the open-weight Nemotron family of language and multimodal models, tuned for enterprise reasoning, agentic workflows, and document intelligence.
Mirelo focuses on generative audio — sound effects and ambient soundscapes driven by text prompts.
Cohere builds enterprise language models — the Command family and the North agent platform — tuned for retrieval-augmented generation, tool use, and secure private deployment.
Moonshot builds the Kimi family — frontier open-weight models with very long context, native visual understanding, and strong agentic coding.
ACE Studio builds the open ACE-Step music models — full-track generation with lyric editing, remixing, and voice cloning in 50+ languages.
Z.ai develops the open GLM family of language models, built for long-horizon agentic coding and end-to-end engineering work over very large contexts.
Supertone is a Seoul voice-AI company whose Supertonic text-to-speech models are exported to ONNX and built to run entirely on the user's own device.