Single clips from a prompt, a photo, or start- and end-frames — with the frontier model you pick per shot. Then trim, crop, and convert clips, or assemble full videos on a multi-track timeline with AI drafting the scenes. Cloud generation runs on your own API keys, open models render locally through the GGML runtime, and editing and rendering run on your machine.
Veo 3.1 for a hero shot with synced audio, Seedance 2.0 Fast for cheap iterations, Kling 3.0 or Gen-4.5 when a scene needs a different look. One API key per platform unlocks every model it hosts, and your clips, prompt history, and settings stay in the same workspace no matter which model renders.
ProvidersRunwareCloud · key setReplicateCloud · key setGGMLLocal · 5 video modelsVideo modelsKey validVeo 3.1Google · CloudSeedance 2.0ByteDance · CloudKling 3.0Kuaishou · CloudGen-4.5Runway · CloudHappyHorse 1.1Alibaba · CloudWan 2.2 TI2V-5BAlibaba · LocalOne key per platform unlocks every model it hosts.
Generate
From prompt to footage.
Describe the shot — subject, camera move, light — pick an aspect ratio and resolution, and render. Long jobs stream progress into the composer, and finished clips land in your project folder like any other file.The clip playing here is a real, unretouched Veo 3.1 generation from CSuite — the exact prompt is in the bar below it, and the same test renders for every video model on its catalog page.
coastal-road.mp4Real output · Veo 3.1RunwareVeo 3.116:9720pAudio onSlow aerial push in over a coastal cliff road at golden hour, waves breaking on rocks below, a single car travelling along the road…Render
Animate
Start from a still, end in motion.
Attach a photo as the start frame and describe the motion — the model animates from exactly that image, keeping your composition and subject. On models that accept an end frame too, give it the first and last shot and it generates the move that connects them: perfect for product turns and scene transitions.
latte-pour.mp4Real output · Seedance 2.0RunwareSeedance 2.0Start frame9:16Close-up of a barista pouring steamed milk into an espresso, latte art forming in the crema, steam rising from the cup…Generating
Enhance
Type the gist, get the direction.
Press the wand and one line becomes a proper shot: subject and action, camera movement, framing, light, and mood. With a start frame attached it directs the motion instead of re-describing the scene, and it only writes sound when the model actually renders audio.Video is billed per second and takes minutes to come back, so the vague attempt is the costly one.
ComposerEnhance prompt with AIYou typedshe turns around and tells him she is leaving↓Rewritten by your text model, for the model you pickedSent to the modelShe turns from where she stands to face him, the movement slow and deliberate. The camera pushes in to a tight shot on her face, holds a beat, then racks focus to his reaction behind her. She says, quietly, “I’m leaving.” Room tone underneath, the rustle of her coat as she turns, then silence.Seedance 2.5Start frame attached, so it directs the motionAudio on, so dialogue is quoted
Edit
Fix clips without leaving the app.
The everyday cuts that usually mean opening an editor, built in and powered by ffmpeg on your machine — with progress you can cancel, and nothing uploaded anywhere.
Resize for the destination
Drop a full-res render down to 720p for a web embed or 480p for a quick share — pick a preset height from 240p to 1080p and the width follows the clip's aspect ratio.
coastal-road.mp4240p360p480p720p1080p→ 1280×720
Crop to a new aspect ratio
Turn a 16:9 landscape shot into a 9:16 vertical for Shorts and Reels, or a 1:1 square for a feed — crop by aspect ratio and keep the action centered.
coastal-road.mp416:99:161:1Save new copy
Convert the container
Deliver the same clip as MP4, WebM, or MOV without a round-trip through another tool — a WebM for your site, a MOV for the editor who asked for one. Or export a looping GIF at the frame rate and width you want — it lands in your Image workspace, ready to drop into a README or a Slack thread.
latte-pour.mp4latte-pour.mp4→latte-pour.webmMP4WebMMOVSave new copy
TrimSet an in and out point by typing them or parking the playhead and clicking Set. The kept range shows as a band on the scrubber.
SpeedSlow a clip to half speed or push it to 2×, with the pitch of any audio preserved rather than chipmunked.
Audio in or outPull a clip's soundtrack out as its own MP3 — it appears in the Audio workspace — or strip it for a silent cut, which takes seconds since the picture isn't re-encoded.
Grab a frameSave the frame under the playhead as a PNG, or send it straight to the composer as the start frame of the next clip. That's how you extend a shot.
UpscaleTake a clip to 1080p or 4K with Real-ESRGAN — no prompt, just pick it from the AI menu. Runs on your Replicate key.
Remove the backgroundMatte a subject out to a green screen with Robust Video Matting. Built for talking-head and lip-sync footage.
Step frame by frameArrow keys move the playhead a second at a time; hold Shift and it steps a single frame, for finding the exact cut.
Know what you're holdingOne panel with duration, resolution, frame rate, codecs, and bitrate — plus the model and prompt behind anything CSuite generated.
Compositions
Not just clips — whole videos.
A second creation mode for long-form work: assemble titles, images, clips, and audio on a multi-track timeline and render one finished video. Describe it in plain English and the AI drafts the whole thing for you.
A multi-track timeline
Titles, images, video, and audio each get their own track. Drag to move, pull an edge to resize, drag up or down to change what draws on top, and snap cleanly to neighbouring clips, the playhead, or a clip's own true length when you're trimming it. Fades, volume, motion, and exact position live in the layer panel — or just drag and scale a title directly on the preview, with guides that snap it to the centre and the edges.
Tell the assistant what the video is for and it plans the beats and lays them onto the timeline as real clips — titles, voice-over, music, sound effects — then generates the image and audio layers in parallel, so timings and title cards are playable immediately while the rest backfill. It picks one visual theme and holds every image to it, so the shots look like one shoot. Refine with follow-ups: nudge a timing, reword a title, quieten the music. Each request touches only what you asked about and undoes in one step.
AI assistantRoad trip teaser12-second teaser for a coastal road trip film: opening title, drone shot, voice-over, warm music.Drafted 3 tracks · 6 clips. Generating layers in parallel:✓ Titles · ✓ Images · ● Voice-over · ● Music bedRefining timing…Hold the opening title 1s longer, then start the music on the first cut…Generating
Turn a still into a shot
Every layer has a properties panel: timing and fades, exact position and size, motion, volume and looping, and for titles the fonts actually installed on your machine. Image layers get one more — Make video. Describe the motion and your chosen model animates the still from that exact frame; the clip swaps in place, keeping its timing and framing, and one ⌘Z puts the image back. Changed the shape of the whole video? Switch aspect ratio and every clip is rescaled to fit rather than stretched.
Layer · Coastal stillimage clipTimingStart2.4sLength3.0sPositionW72%H64%Generate with AIslow push in, waves breaking belowRegenerateMake videoThe still becomes the first frame — the clip swaps in place, and ⌘Z puts the image back.
Render on your machine
Export at up to 4K and 60 fps as MP4, WebM, or MOV. The render runs locally with ffmpeg — your footage never uploads — and the finished file lands in your project folder like any other clip.
Export · Road trip teaser0:12 · 16:9ResolutionNative720p1080p2160pFramerateNative243060ContainerMP4 · H.264WebM · VP9MOV · H.264Rendering locally…64%Renders on your machine with ffmpeg — nothing uploads.
Catalog
Available models & providers.
Every video model in the catalog, in one picker. Cloud models run through Runware and Replicate with your own keys; open-weight models run locally through the GGML runtime — each with real sample clips on its detail page.
How video generation, clip editing, and compositions work in CSuite — models, sound, costs, and rendering.
The catalog spans the current flagship video models — Veo 3.1 and Veo 3.1 Fast, the Seedance family through 2.5, Kling 3.0, Gen-4.5, LTX-2.5 Pro and Fast, MiniMax H3, Flux 3 Video, HappyHorse 1.1, Wan 2.7, Grok Imagine Video 1.5, and Gemini Omni Flash. Cloud models run through Runware and Replicate with your own API keys, and open-weight models — Wan 2.1 and 2.2, LTX-2.3 Distilled, and MiniMax H3 — run locally through the GGML runtime. Browse the full list with per-model specs and real sample clips at csuite.so/models.
Yes. Type a one-line idea and hit the wand in the composer, and your selected text model rewrites it into a proper shot covering subject and action, camera movement, framing, lighting, and mood. It adapts to the run: with a start frame attached it describes movement and camera instead of re-describing a scene the frame already holds, and it only writes sound and dialogue when the model actually renders audio and the toggle is on. Since video is billed per second and takes minutes to come back, sharpening the prompt for a fraction of a cent first is usually cheaper than another attempt.
Yes. Open-weight video models — Wan 2.1 and 2.2, LTX-2.3 Distilled, and MiniMax H3 — run on your own machine through the bundled GGML runtime, built on stable-diffusion.cpp. Video is the heaviest local workload, so each model lists the hardware it needs; a modern GPU or Apple silicon makes a big difference. Everything after generation is local too: clip edits (trim, resize, crop, speed, audio extraction, GIF export, format conversion) and composition rendering run on your machine with ffmpeg, and your footage never uploads anywhere.
Trim it to an in and out point, resize it, crop it to another aspect ratio, change its speed with the pitch preserved, pull the soundtrack out into its own MP3 or strip it entirely, export a looping GIF, or convert between MP4, WebM, and MOV. You can also grab the frame under the playhead as a PNG — or grab it and drop it straight into the composer as the start frame of the next clip, which is how you extend a shot. Two AI tools need no prompt at all: upscale to 1080p or 4K, and background removal for talking-head footage.
On models that support it, yes — Veo 3.1, for example, can generate a synchronized audio track, and the toggle appears in the generation settings whenever the selected model offers it. For everything else, compositions let you layer voice-over, music, and sound effects onto any clip.
Yes. Models that accept a starting image show a “Start frame” control in the composer — attach a photo and describe the motion. Several models also accept an end frame, generating the motion that connects your first and last shot.
Long-form videos built on a multi-track timeline: titles, images, video clips, and audio arranged over time, with drag-to-move, resize, snapping, fades, and per-clip controls. Describe what you want and the AI assistant lays the clips straight onto the timeline, then generates the image, voice-over, music, and sound-effect layers in parallel. The result renders to a single video file.
A properties panel per clip. Titles get their text plus font — chosen from the fonts actually installed on your machine — size, weight, colour, and alignment. Images and video get their file, how it fills the frame, and exact position, size, and rotation, plus a from/to motion animation. Video and audio get a volume slider and a loop toggle, so a clip that outruns its source either repeats or slows to fill the gap instead of freezing. You can also drag and scale anything directly on the preview.
Yes. Any image clip has a Make video action: the still becomes the first frame, you describe the motion, and your chosen video model animates it. The clip is swapped for the video in place, keeping its timing and framing — one ⌘Z puts the image back, and the rendered file stays in your project folder either way.
Compositions export at native size or 480p / 720p / 1080p / 2160p, at native, 24, 30, or 60 fps, as MP4 (H.264), WebM (VP9), or MOV (H.264). Rendering runs entirely on your machine via ffmpeg — nothing is uploaded, and progress is shown with a cancel option.
Video is the most expensive modality — typically billed per second of output, and a single clip can cost more than a batch of images. CSuite doesn't meter or mark anything up: the platform (Runware or Replicate) bills your account directly, and the app's analytics show spend per provider so there are no surprises.
In your project folder, as plain files on your own disk — generated clips, imported footage, and rendered compositions alike. No proprietary library, no cloud sync.
One-time payment. Yours forever.
No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.