The idea-to-video pipeline: script, stills, motion, sound
One idea, four models, $3.53: we made a 30-second film the slow way and kept every receipt. Watch it, then steal the pipeline.
Last August we typed single sentences into four video models and published the contact sheet: beautiful atmosphere, three-handed pianists, a chalkboard reading CNCKNY POKLCK. The lesson was that one prompt buys you a lottery ticket, drawn from everything the model has ever seen. This post is the other way to spend the money. We took one idea, a quiet 30-second film about a lighthouse keeper, and walked it through the full production pipeline: a text model wrote the script, an image model designed each frame, a video model animated the frames, and two audio models recorded the voiceover and the score.
The finished piece is embedded below, along with every intermediate artifact, every price, and the two places the pipeline stumbled. Total model bill: $3.53. The argument is not that the film is great cinema. It is that at every stage a human got to see the work, judge it, and fix it cheaply before the next model made it expensive, and that this loop, not any single model, is what turns generation into production.
Start with a brief, not a prompt
A real production starts before the camera: someone decides what is being made, for whom, and what has to be true of it. Ours fits in a paragraph. Title: “The Keeper.” Thirty seconds, four shots, one character, no dialogue on camera, voiceover and music underneath. Tone: quiet, warm, pre-dawn to sunrise. That paragraph did more work than any prompt in this post, because every later stage inherited its decisions.
The constraints are not aesthetic preferences; they are engineering around known failure modes. Our text-to-video test found that models paint gorgeous atmosphere and fumble on-screen text, object counts, and fine hand work. So the brief bans all three: no signs, no counting, no close-up manipulation. One character, because two doubles the identity problem. Sunrise, because warm directional light flatters every image model ever trained. You can design a brief that fights the medium or one that plays to it, and the price of the film mostly follows that choice.
A script costs a tenth of a cent and saves ten dollars
Stage two hands the brief to a text model and asks for structure: four shots, each with a one-sentence visual, a camera move, and a voiceover line under 14 words. We ran gpt-oss-120b on Replicate, an open model whose median hosted price is $0.15 in and $0.60 out per million tokens. Our whole script stage used 2,302 tokens across two calls: about a tenth of a cent. It was also the only stage that needed a retry for a dumb reason; the first reply spent our 800-token budget on its own internal reasoning and cut off mid-sentence.
Why bother, when you could write four beats yourself? You should edit them yourself, and we did: one grammar fix, one line rewritten. But the model’s draft did something valuable that a blank page does not. It converted a vibe into named, numbered assets. From here on, nothing in the pipeline is “the film”; everything is shot 2 or line 3, and a problem in one asset never forces a do-over of the whole. The same discipline that makes a good prompt specific, one subject, one action, one camera note, is being imposed here on the entire production at once.
Stills are the cheapest place to change your mind
Stage three is where most one-prompt users skip ahead and pay for it. Before any video is generated, each beat becomes a designed still: the exact frame the clip will open on. We sent each shot’s visual to Seedream 4.5 through Runware at $0.04 per 2560×1440 frame, prefixed with a character sheet: “a woman lighthouse keeper in her early sixties, silver hair in a single braid, weathered kind face, dark navy wool sweater.” The same sentence, word for word, on all four frames, plus a shared style suffix for the 35mm look and the slate-and-amber palette.
This is the stage where iteration is nearly free. Wrong outfit, wrong framing, wrong light: four cents and ten seconds to try again. We accepted all four boards on the first pass, but honesty requires the catch in the caption: our character sheet pinned the woman and said nothing about the building, so shot 3 grew a different, older lighthouse with a red door. We noticed at this stage and let it go, gambling that no viewer tracks architecture across a 30-second film the way they track a face. For a brand product, that gamble is a reshoot: pin a reference image or a written spec for every recurring thing in the frame, not just the people.
Animate a frame you approved, not a sentence you hoped
Stage four is the expensive one, and the pipeline earns its keep here. Instead of asking a video model to imagine each scene from text, we handed it the approved still as a locked first frame plus the script’s camera move: Veo 3.1 Fast’s image-to-video mode through Runware, 8 seconds at 720p, native audio off because sound is its own stage. Each clip cost $0.80 and took about a minute, in line with Google’s own $0.10-per-second list price. Every pixel of drift now measures against a frame we chose, not a scene we described. It shows: her face survives motion because the model starts from her face, not from “a woman in her sixties.”
Be clear about what the first-frame lock does and does not buy. It guarantees the opening frame is exactly your approved still, so identity, wardrobe, palette, and composition all start correct, and drift has to work against a fixed anchor instead of a vague sentence. It does not guarantee the middle of the move: fast action can still smear a face or a label at second five, which is why our brief kept every motion slow and continuous. The cost asymmetry is the whole strategy. A still is $0.04 and ten seconds; a clip is $0.80 and a minute. Twenty to one, so every decision that can possibly be made at the still stage should be, and the video model should be handed something to execute rather than something to invent.
That miss is the most instructive artifact in the post. The camera note was written by a text model that has never seen a still, executed by a video model that never read the brief. Neither is wrong; the handoff is. A pan away from your subject is a fine instruction for a landscape and a bad one for a character film, and only the human holding both contexts can catch that. We should have; we read the camera note, approved it, and learned the lesson for $0.80 instead of a reshoot day.
Sound is cents, and it carries half the feeling
Stage five costs almost nothing and changes almost everything. Watch any clip above muted, then with sound: the silent version is a tech demo, the scored one is a film. Our four voiceover lines went to Gemini 3.1 Flash TTS through Runware, voice Kore, at well under a cent per line, 16 seconds of finished narration for $0.02. The score came from MiniMax Music 2.6 for $0.15, prompted with the script model’s own one-line music direction. We asked for 40 seconds and got 52; length obedience is still loose in music models, so generate long and trim. The full method, including the consent rule for cloning real voices, is in our voiceover post.
The finished film, and what it cost
Assembly is the one stage with no AI in it: ffmpeg concatenated the four clips with half-second crossfades, laid the score under the whole 30 seconds at about a third of its volume with a fade at each end, and dropped each voiceover line at the start of its shot, first line at 1.5 seconds, last ending two and a half seconds before the fade to black. Then one verification pass a human cannot skip by watching: we ran the finished soundtrack through a speech-to-text model, which confirmed all four lines land inside their intended shots. It also transcribed a phantom “The End” during the instrumental fade-out, a small reminder that verification models hallucinate too. Here is “The Keeper,” 30 seconds, four models, one afternoon.
Two honest asterisks on the $3.53. First, it excludes the retries a normal project should budget: we got lucky on visuals, and our 30-second ad build with its planned rerolls is the more typical bill. Second, it excludes the most expensive component: the model wait time summed to eight minutes, but the human in the director’s chair, reading beats, judging boards, catching the lighthouse swap, timing the mix, spent about two hours. The pipeline does not remove the filmmaker. It gives the filmmaker cheaper, faster departments.
Fix every problem at the earliest stage it appears
Everything that went wrong in this production was visible before the stage where it would have become expensive. The truncated script cost a fraction of a cent to rerun. The lighthouse swap was visible in a $0.04 still; we chose to accept it. The pan that loses the keeper was latent in a camera note we could have rewritten for free. Quality compounds through a pipeline, and so does negligence: a sloppy beat becomes a wrong still becomes a wasted clip becomes a re-recorded voiceover. The craft is not prompting; it is inspection at handoffs.
The reusable checklist, in production order. Write a brief that bans what video models fumble: on-screen text, counts, fine hand work. Script with a text model into numbered shots, then edit the lines yourself; you are the writers’ room, it is the typist. Pin a character sheet for every recurring person and thing, and reuse it word for word. Approve a still per shot before paying for motion; four cents beats eighty. Animate from approved frames, and read every camera note while looking at the frame it will move. Lay sound as its own stage, generated long and trimmed. And keep the receipts, because the numbers argue better than the showreel: our slot-machine post spent $9.02 learning what one sentence buys, and this one spent $3.53 proving that the same money, walked through a pipeline, comes back as a film.


