Skip to content
CSuite
GuideVideo AICreatorsAugust 4, 202610 min read

The idea-to-video pipeline: script, stills, motion, sound

One idea, four models, $3.53: we made a 30-second film the slow way and kept every receipt. Watch it, then steal the pipeline.

The pipeline, with receipts
One sentence is a slot machine. A pipeline is a production.
01
Idea
one paragraph · free
02
Script
4 beats + VO · $0.00
03
Stills
4 frames · $0.16
04
Motion
4 clips · $3.20
05
Sound
VO + score · $0.17
Four AI models chained end to end on August 4, 2026: gpt-oss-120b on Replicate for the script, then Seedream 4.5, Veo 3.1 Fast, Gemini TTS, and MiniMax music through Runware’s API. Every visual is the first take, no rerolls, $3.53 billed in total. The 30-second result is embedded below, misses included.

Last August we typed single sentences into four video models and published the contact sheet: beautiful atmosphere, three-handed pianists, a chalkboard reading CNCKNY POKLCK. The lesson was that one prompt buys you a lottery ticket, drawn from everything the model has ever seen. This post is the other way to spend the money. We took one idea, a quiet 30-second film about a lighthouse keeper, and walked it through the full production pipeline: a text model wrote the script, an image model designed each frame, a video model animated the frames, and two audio models recorded the voiceover and the score.

The finished piece is embedded below, along with every intermediate artifact, every price, and the two places the pipeline stumbled. Total model bill: $3.53. The argument is not that the film is great cinema. It is that at every stage a human got to see the work, judge it, and fix it cheaply before the next model made it expensive, and that this loop, not any single model, is what turns generation into production.

Start with a brief, not a prompt

A real production starts before the camera: someone decides what is being made, for whom, and what has to be true of it. Ours fits in a paragraph. Title: “The Keeper.” Thirty seconds, four shots, one character, no dialogue on camera, voiceover and music underneath. Tone: quiet, warm, pre-dawn to sunrise. That paragraph did more work than any prompt in this post, because every later stage inherited its decisions.

The constraints are not aesthetic preferences; they are engineering around known failure modes. Our text-to-video test found that models paint gorgeous atmosphere and fumble on-screen text, object counts, and fine hand work. So the brief bans all three: no signs, no counting, no close-up manipulation. One character, because two doubles the identity problem. Sunrise, because warm directional light flatters every image model ever trained. You can design a brief that fights the medium or one that plays to it, and the price of the film mostly follows that choice.

A slate exists so everyone agrees on what shot this is before film rolls. The brief does that job for a model pipeline. Photo by Avel Chuklanov on Unsplash.

A script costs a tenth of a cent and saves ten dollars

Stage two hands the brief to a text model and asks for structure: four shots, each with a one-sentence visual, a camera move, and a voiceover line under 14 words. We ran gpt-oss-120b on Replicate, an open model whose median hosted price is $0.15 in and $0.60 out per million tokens. Our whole script stage used 2,302 tokens across two calls: about a tenth of a cent. It was also the only stage that needed a retry for a dumb reason; the first reply spent our 800-token budget on its own internal reasoning and cut off mid-sentence.

The script, verbatim from the model
SHOT 1
The keeper, silver-braided, in a navy sweater, walks the rocky path toward the lighthouse under a dim, pre-dawn sky.
CAMERA: steady dolly forward along the path
VO: “Before the world awakens, she begins her quiet vigil.
SHOT 2
Inside the lantern room, she lifts a brass lantern, the blue-gray light spilling onto her face.
CAMERA: slow push in from the doorway
VO: “She tends the flame that will guide sailors through night.
SHOT 3
On the cliff's edge, she watches the horizon blush as sunrise lifts.
CAMERA: slow pan right across the sea
VO: “The day arrives, painting the sea with hope.
SHOT 4
She closes her weathered journal, the sunrise flooding the room with gold.
CAMERA: track backward as she walks toward the open door
VO: “And with the light, she keeps the lighthouse's promise.
Raw output from gpt-oss-120b on Replicate, August 4, 2026, unedited. It also returned a one-line music direction: “Gentle ambient piano with distant low strings, swelling softly as the sunrise brightens.” Before recording, we fixed the grammar in shot 2’s line and rewrote shot 3’s (“painting the sea with hope” is a greeting card); shots 1 and 4 went to the voice model as written.

Why bother, when you could write four beats yourself? You should edit them yourself, and we did: one grammar fix, one line rewritten. But the model’s draft did something valuable that a blank page does not. It converted a vibe into named, numbered assets. From here on, nothing in the pipeline is “the film”; everything is shot 2 or line 3, and a problem in one asset never forces a do-over of the whole. The same discipline that makes a good prompt specific, one subject, one action, one camera note, is being imposed here on the entire production at once.

Stills are the cheapest place to change your mind

Stage three is where most one-prompt users skip ahead and pay for it. Before any video is generated, each beat becomes a designed still: the exact frame the clip will open on. We sent each shot’s visual to Seedream 4.5 through Runware at $0.04 per 2560×1440 frame, prefixed with a character sheet: “a woman lighthouse keeper in her early sixties, silver hair in a single braid, weathered kind face, dark navy wool sweater.” The same sentence, word for word, on all four frames, plus a shared style suffix for the 35mm look and the slate-and-amber palette.

Shot 1the path, pre-dawn
Shot 2the lantern room
Shot 3the cliff at sunrise
Shot 4the journal, golden hour
All four boards on the first take: Seedream 4.5 via Runware, August 4, 2026, $0.04 per frame, no rerolls, no fixed seeds. The character sheet held her face, braid, and sweater across all four. Look closer, though: the lighthouse in shot 3 is not the lighthouse from shot 1.

This is the stage where iteration is nearly free. Wrong outfit, wrong framing, wrong light: four cents and ten seconds to try again. We accepted all four boards on the first pass, but honesty requires the catch in the caption: our character sheet pinned the woman and said nothing about the building, so shot 3 grew a different, older lighthouse with a red door. We noticed at this stage and let it go, gambling that no viewer tracks architecture across a 30-second film the way they track a face. For a brand product, that gamble is a reshoot: pin a reference image or a written spec for every recurring thing in the frame, not just the people.

Animate a frame you approved, not a sentence you hoped

Stage four is the expensive one, and the pipeline earns its keep here. Instead of asking a video model to imagine each scene from text, we handed it the approved still as a locked first frame plus the script’s camera move: Veo 3.1 Fast’s image-to-video mode through Runware, 8 seconds at 720p, native audio off because sound is its own stage. Each clip cost $0.80 and took about a minute, in line with Google’s own $0.10-per-second list price. Every pixel of drift now measures against a frame we chose, not a scene we described. It shows: her face survives motion because the model starts from her face, not from “a woman in her sixties.”

Shot 2, first take, straight from the API: the approved still, the script’s “slow push in,” and 8 seconds of a steady lantern flame. Veo 3.1 Fast, image-to-video, $0.80, 65 seconds of waiting.
Shot 3, first take, our one miss: the script said “slow pan right across the sea,” and the model obeyed so literally that the keeper exits her own film. We kept it because the empty sunrise happens to fit the voiceover line; on a client job this is a $0.80 retake with a rewritten camera note.

Be clear about what the first-frame lock does and does not buy. It guarantees the opening frame is exactly your approved still, so identity, wardrobe, palette, and composition all start correct, and drift has to work against a fixed anchor instead of a vague sentence. It does not guarantee the middle of the move: fast action can still smear a face or a label at second five, which is why our brief kept every motion slow and continuous. The cost asymmetry is the whole strategy. A still is $0.04 and ten seconds; a clip is $0.80 and a minute. Twenty to one, so every decision that can possibly be made at the still stage should be, and the video model should be handed something to execute rather than something to invent.

That miss is the most instructive artifact in the post. The camera note was written by a text model that has never seen a still, executed by a video model that never read the brief. Neither is wrong; the handoff is. A pan away from your subject is a fine instruction for a landscape and a bad one for a character film, and only the human holding both contexts can catch that. We should have; we read the camera note, approved it, and learned the lesson for $0.80 instead of a reshoot day.

Sound is cents, and it carries half the feeling

Stage five costs almost nothing and changes almost everything. Watch any clip above muted, then with sound: the silent version is a tech demo, the scored one is a film. Our four voiceover lines went to Gemini 3.1 Flash TTS through Runware, voice Kore, at well under a cent per line, 16 seconds of finished narration for $0.02. The score came from MiniMax Music 2.6 for $0.15, prompted with the script model’s own one-line music direction. We asked for 40 seconds and got 52; length obedience is still loose in music models, so generate long and trim. The full method, including the consent rule for cloning real voices, is in our voiceover post.

The soundtrack, before the mix
Voiceover, line two
Gemini 3.1 Flash TTS via Runware, voice Kore: 4.2 seconds, $0.004, back in 3.8 seconds.
The score
MiniMax Music 2.6 via Runware, instrumental: we asked for 40 seconds, it returned 52. $0.15, back in 96 seconds.
Both first takes, generated August 4, 2026. The music prompt was the script model’s own direction, pasted verbatim with a length and a no-drums note added.

The finished film, and what it cost

Assembly is the one stage with no AI in it: ffmpeg concatenated the four clips with half-second crossfades, laid the score under the whole 30 seconds at about a third of its volume with a fade at each end, and dropped each voiceover line at the start of its shot, first line at 1.5 seconds, last ending two and a half seconds before the fade to black. Then one verification pass a human cannot skip by watching: we ran the finished soundtrack through a speech-to-text model, which confirmed all four lines land inside their intended shots. It also transcribed a phantom “The End” during the instrumental fade-out, a small reminder that verification models hallucinate too. Here is “The Keeper,” 30 seconds, four models, one afternoon.

“The Keeper,” complete: script by gpt-oss-120b, stills by Seedream 4.5, motion by Veo 3.1 Fast, voice by Gemini TTS, score by MiniMax, assembled with ffmpeg on August 4, 2026. Every visual is a first take.
The receipt, August 4, 2026
Every generation behind the film, with per-task cost reporting on. Two retries all run: the script model’s first reply hit our token cap mid-sentence, and the music request timed out once before succeeding. Every image and clip is the first result.
Script: gpt-oss-120b, 2 calls$0.00
Stills: 4 × Seedream 4.5$0.16
Motion: 4 × Veo 3.1 Fast, 8 s i2v$3.20
Voiceover: 4 × Gemini 3.1 Flash TTS$0.02
Score: MiniMax Music 2.6$0.15
Total, one 30-second film$3.53

Two honest asterisks on the $3.53. First, it excludes the retries a normal project should budget: we got lucky on visuals, and our 30-second ad build with its planned rerolls is the more typical bill. Second, it excludes the most expensive component: the model wait time summed to eight minutes, but the human in the director’s chair, reading beats, judging boards, catching the lighthouse swap, timing the mix, spent about two hours. The pipeline does not remove the filmmaker. It gives the filmmaker cheaper, faster departments.

Fix every problem at the earliest stage it appears

Everything that went wrong in this production was visible before the stage where it would have become expensive. The truncated script cost a fraction of a cent to rerun. The lighthouse swap was visible in a $0.04 still; we chose to accept it. The pan that loses the keeper was latent in a camera note we could have rewritten for free. Quality compounds through a pipeline, and so does negligence: a sloppy beat becomes a wrong still becomes a wasted clip becomes a re-recorded voiceover. The craft is not prompting; it is inspection at handoffs.

The stage nobody automates: the edit is where every earlier decision gets judged together. Photo by Jakob Owens on Unsplash.

The reusable checklist, in production order. Write a brief that bans what video models fumble: on-screen text, counts, fine hand work. Script with a text model into numbered shots, then edit the lines yourself; you are the writers’ room, it is the typist. Pin a character sheet for every recurring person and thing, and reuse it word for word. Approve a still per shot before paying for motion; four cents beats eighty. Animate from approved frames, and read every camera note while looking at the frame it will move. Lay sound as its own stage, generated long and trimmed. And keep the receipts, because the numbers argue better than the showreel: our slot-machine post spent $9.02 learning what one sentence buys, and this one spent $3.53 proving that the same money, walked through a pipeline, comes back as a film.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app