Skip to content
CSuite
TutorialVideo AIMarketingMay 26, 202613 min read

How to make a 30-second AI video ad, end-to-end

Forty-seven minutes. Three dollars and seventy-three cents. Every prompt, every cost, one API — and the finished ad, embedded.

Six beats · thirty seconds · one afternoon
$3.73 · 47 minutes
A polished 30-second ad, built on a laptop, on a Sunday afternoon, for less than the price of a coffee.
Beat 01
0–3s
Hook: hand opens mailbox, finds the bag
Veo 3.1 Fast (i2v)
Beat 02
3–9s
Bag close-up, label rotates into focus
Veo 3.1 Fast (i2v)
Beat 03
9–15s
Beans into grinder, slow-mo pour
Veo 3.1 Fast (i2v)
Beat 04
15–21s
Cup on a kitchen counter, morning light
Veo 3.1 Fast (i2v)
Beat 05
21–27s
Subscriber holds bag, smiles at camera
Veo 3.1 Fast (i2v)
Beat 06
27–30s
Logo + CTA card, captioned
Nano Banana still

Forty-seven minutes. Three dollars and seventy-three cents. One thirty-second vertical ad for a single-origin coffee subscription, ready to upload to Meta, TikTok, and YouTube Shorts. No camera. No actor. No studio. No agency. Just a laptop, a kitchen window, and a stack of AI models that did not exist as a usable workflow eighteen months ago.

This is the post that walks the whole thing, and this time the receipts are literal: we ran the workflow for real on July 27, 2026, every generation billed through one API, and the finished ad is embedded right below. Every prompt that matters is in here, verbatim. Every minute is counted, every dollar is itemized. By the end you have a recipe you can run against your own product on a Sunday afternoon. If you have ever sat through a five-figure video shoot to produce something a 30-second AI workflow could match, this is the read.

The deliverable, unretouched: 720×1280, 24fps, 30.0 seconds, 9.3 MB. Five Veo 3.1 Fast clips, one Nano Banana end card, six TTS lines, one music bed, cut with ffmpeg. Total generation spend: $3.73.

One framing note before we start. AI ads are not magic and they are not free. They are cheap, fast, and surprisingly controllable, if you treat them as a craft. Most of the time savings show up because you skip the parts of a traditional shoot that don’t need a human: scouting, model release forms, scheduling, the second-camera angle that never gets used. The parts you keep are editorial: the brief, the cut, and the judgement about what to keep and what to throw away. AI is a cheap session crew. You are still the director.

The brief is the part nobody wants to write

A 30-second ad is a six-beat structure. Naming the beats before you open a single tool is the only way to keep the AI from ad-libbing for thirty seconds about nothing. The brief for this ad fits in one paragraph. The product: Bayfield Coffee, a fictional single-origin bag-per-month subscription, $22 a month, ships on the first of each month, beans roasted three days before they leave the warehouse. The audience: home brewers who already own a grinder, who are mildly bored with their current grocery-store bag. The platform: Meta and TikTok first, 9:16 vertical, ≤30 seconds, captioned. The call to action: a promo code BAYFIELD10 for $10 off the first bag.

Three things in that paragraph save the next 40 minutes. The first is the platform: deciding 9:16 vertical now keeps every later step from accidentally drifting into 16:9. The second is the audience-already-owns-a-grinder constraint: it removes a beat (no need to explain what a grinder is) and unlocks the slow-motion-pour beat that actually sells the product. The third is the $10-off CTA: a specific number for the closing card is more useful than “learn more”. Write a brief like this for any product you are about to feed to an AI workflow. It is the highest-leverage twenty minutes in the whole process.

A 30-second ad is six beats. The LLM is good at the beats.

Open Claude, ChatGPT, or Gemini. The LLM doesn’t matter much for this step, all three are competent. Paste in the brief from above, then ask for a beat sheet. Here is the prompt, verbatim:

You are writing a 30-second vertical video ad for Bayfield Coffee,
a $22/month single-origin coffee subscription. Audience: home brewers
who already own a grinder. Platform: Meta and TikTok, 9:16, captioned.
CTA: promo code BAYFIELD10 for $10 off the first bag.

Output: a 6-beat script with these columns for each beat:

1. timestamp range (e.g. 0-3s)
2. one-line visual description (must be filmable as image-to-video)
3. voiceover line (max 8 words, conversational)
4. caption text (max 5 words, key word emphasized)

Constraints:
- Open with a pattern interrupt in the first 1.5 seconds.
- Show the bag in the first 9 seconds.
- End on the promo code, not the brand name.
- No actor faces in close-up (we want product-led, not testimonial).

Two minutes of generation, four minutes of trimming. The first draft had eight beats and a script that was six words too long for a calm voiceover at 28 seconds. The second pass dropped the brand-history beat (nobody cares in three seconds) and tightened the VO. The cuts are still yours: an LLM will give you a competent draft, but it will always give you slightly too much. The right move is to cut beats, not cram them in.

On hook structure: TikTok’s own creative team is explicit that 71% of whether a user keeps watching is decided in the first three seconds, and Sprout Social’s 2026 algorithm breakdown confirms the same binary pass-fail at the three-second mark. Beat one earns the rest of the ad. Spend more LLM time on it than on any other beat. The hand-opens-a-mailbox visual is a pattern interrupt (ordinary frame, unexpected object), and it is the cheapest part of the whole workflow to iterate on.

Storyboard frames: one per beat, fidelity over creativity

Skip text-to-video on the first pass. Generate a still image for each beat, then animate the still in step three. There are two reasons. First, image generation is roughly thirty times cheaper than video generation per attempt, so you iterate without watching a meter. Second, an image-to-video pass gives you something the text-to-video path does not: a locked-in starting frame the model has to honor.

We generated all six frames with Google’s Nano Banana 2 Lite through Runware’s image API: $0.10 per frame at 2K (768×1376 costs a third of that if 1K is enough), $0.62 for the set, about twenty seconds per frame. The reason it gets the job over prettier texture models like FLUX.2 or Seedream is the same reason it costs what it does: it is a Gemini-family model, and it renders typography. A product ad lives and dies on the label. All six frames came back with “BAYFIELD” in the same dark-green serif on the same cream label, legible at phone size, on the first try.

The consistency was not luck; it was the prompt. Every one of the six prompts ends with the same pasted brand block, and the block describes the bag the way a props department would:

The coffee bag is always the same product: a matte kraft-paper
coffee bag with a cream-colored rectangular label, the word "BAYFIELD"
printed large in dark forest-green serif capitals, a small minimal
line drawing of two mountain peaks above the word, and "SINGLE ORIGIN"
in small letterspaced capitals below it. Natural warm morning light,
shallow depth of field, photorealistic, vertical 9:16 composition.

One take did not fully behave, and honesty demands it be on the record: the grinder-pour frame invented a dark-green side panel on the bag that exists in no other frame. In a real client job that is a re-roll (ten more cents); for this ad it is on screen for six seconds at an angle and nobody has ever noticed until this sentence. The other five frames, including the flat-design CTA card with the promo code set in a button shape, were first-take keepers. When a $0.10 generation misbehaves you re-roll it; that is the entire luxury of doing your storyboard in images instead of video.

The six frames, as generated · Nano Banana 2 Lite · $0.62 total
Beat 01 · Mailbox hook
Beat 02 · Bag close-up
Beat 03 · Grinder pour
Beat 04 · Cup, morning light
Beat 05 · Subscriber, no face
Beat 06 · CTA end card

Video shots: image-to-video, one model, every time

The first version of this post mixed three video models across five shots — and then spent a paragraph apologizing for the style mismatches and the three sets of credentials. The re-shoot does what we said we’d do next time: one model, Veo 3.1 Fast, for all five clips, with each beat’s storyboard frame locked in as the first frame. Through Runware it bills at $0.10 per second at 720p with native audio off, and native audio off is exactly what you want when the voiceover and music are coming from dedicated models anyway.

Five shots: the 3-second mailbox hook (generated at 4 seconds — Veo’s floor — and trimmed), then four 6-second beats. 28 seconds of footage, $2.80, about 50 seconds of wall-clock per clip. Because every clip starts from a frame the image model already got right, the video model’s only job is motion: the hand lifts the bag, the label rotates into focus, the beans cascade, the steam curls. Zero re-rolls. That is not a brag about prompting skill; it is the payoff of the storyboard-first workflow. When the start frame is locked, the failure surface of the video step collapses.

The honest wrinkle: mid-motion, while the hand grabs the bag in the hook shot, the label smears toward “RAYFIELD” for a few frames before snapping back. Frame-locking guarantees the first frame, not every frame. It reads as motion blur at full speed, and the raw clip below is exactly what the model returned, so you can judge for yourself. If a few smeared frames are a dealbreaker for your brand, the fix is the same as ever: re-roll, or cut around it.

The raw 4-second mailbox hook, straight from Veo 3.1 Fast — first take, $0.40, no audio. Watch the label during the grab: locked first frame, briefly smeared mid-motion, recovered by the lift.

Model choice still matters at the margins. Kling 3.0 is cheaper per second and fine for locked-off B-roll; Seedance 2.0 is the motion specialist; our head-to-head from earlier this month runs the same prompts across the current field. But for a product ad where the same physical object must survive five shots, the deciding factor is reference fidelity, and Veo with a locked first frame is the strongest at it. The single-model workflow also means one API, one key, one billing meter — which is, not coincidentally, the argument of the curation post.

Voiceover and music: TTS is ready. Music is controllable-ish.

Voiceover went to Gemini 3.1 Flash TTS — same Runware key, six lines, generated as six separate files so each one could be placed at its beat’s timestamp in the edit rather than praying one long take lands on the cuts. Total cost: just under two cents. The craft moves from the booth to the text: write the promo code as words (“Code Bayfield ten. Ten dollars off.”) or the model will gamble on how to pronounce BAYFIELD10; keep every line under eight words; and put the pauses in as punctuation — “One bag. First of the month. Every month.” reads three beats where a comma would read one.

Pick a voice that matches the brand, not the demo. The default voice (Kore) is excellent and radio-polished, which is exactly wrong for a small-roastery brand; we shipped Zephyr, which sits closer to “friend who is slightly too into coffee”. The most-played voice on any TTS pricing page and the right voice for your brand are almost always different voices.

Music was the one genuinely bumpy step. MiniMax Music 2.6 produces an ad-quality instrumental bed for $0.15 a generation, but it exposes no duration parameter: our first take came back at 21 seconds, nine short. The workaround is blunt — ask for “a 60 second” bed in the prompt, get 45, trim to 30 with a fade — and it cost one extra generation ($0.15) to learn. On licensing, the short version stands: AI music vendors grant commercial use on paid tiers but none of them indemnify you against third-party claims while the training-data lawsuits work through the courts. For a small ad the risk is nominal; for a seven-figure media budget, buy a stock license instead. The full landscape is in the audio catalog.

The cut is where AI ads still live or die

We cut this one with ffmpeg, because the whole run was scripted and ffmpeg is free, but nothing about the edit requires a terminal: CapCut Desktop is the free GUI equivalent, has a vertical preset, and exports 1080×1920 H.264 without a paywall. Use whichever you’ll actually open. The edit itself is the same five decisions in either tool.

Twelve minutes, and here is where they went. Concatenate five clips and hold the CTA card for the last three seconds with a slow push-in (a still end card with a 16% zoom over three seconds reads as “designed”; a static one reads as “ran out of video”). Place each voiceover file at its beat’s timestamp. Duck the music bed to roughly a fifth of its volume under the VO, fade it in over the first half-second and out over the last two and a half. Burn the captions. Normalize the whole mix to −14 LUFS, which is the loudness the big platforms normalize to anyway. One honest hiccup for the ledger: Homebrew’s ffmpeg ships without the text filter, so the captions became six transparent PNGs rendered with Python’s Pillow and overlaid with timed enable windows — five minutes of the twelve, and the kind of yak-shave a GUI editor simply doesn’t have.

Captions are not optional. The TikTok-and-Reels default is autoplay-muted on a vertical phone, and a vertical ad without captions is a vertical ad with the sound off, which means it is half-dead. Ours are the voiceover’s key phrases, five words or fewer, one emphasized word per line, parked above the platform UI zone. And because the whole pipeline was scripted, we closed the loop with a transcription check: the finished ad’s audio went through Whisper, which returned all six lines at their intended timestamps, promo code intact. Costs a fraction of a cent, catches the one TTS line that silently failed to render before your audience does.

One disclosure detail worth flagging here, because it is the kind of thing that gets a paid ad rejected on Tuesday morning. Meta and TikTok have both rolled out AI-disclosure requirements in 2025 and 2026. The threshold across both platforms is now “label any significantly-AI-generated content”, and TikTok provides a one-click toggle in the upload flow that adds the standardized label. TikTok’s policy on synthetic and manipulated media is the strictest of the major platforms: it is mandatory for any ad whose imagery could be mistaken for real footage. This ad qualifies. Toggle the label. The reach hit is small. The policy violation hit is large.

The receipt

The whole point of writing this kind of post is to leave a receipt the reader can mentally hold. Forty-seven minutes is faster than booking a studio. $3.73 is less than the price of a single stock-photo download. These are the numbers from this run, one run, errors and re-rolls included — not a median and not the cherry-picked floor. Your own numbers will move; the shape of the spend will not. Video is the line item that dominates. Audio and the LLM are rounding-error. The cut is the part you cannot offload.

Time & cost ledger · one 30-second vertical ad · July 27, 2026
Step
Tool
Min
USD
01 · Script & storyboard outline
Claude (free tier)
6
$0.00
02 · Storyboard frames (6 × 2K)
Nano Banana 2 Lite, via Runware
8
$0.62
03 · Video shots (5, 28s of footage)
Veo 3.1 Fast, via Runware
12
$2.80
04 · Voiceover (6 lines)
Gemini 3.1 Flash TTS, via Runware
5
$0.02
05 · Music bed (2 takes)
MiniMax Music 2.6, via Runware
4
$0.30
06 · Cut & export
ffmpeg (free)
12
$0.00
Totals
47
$3.73

Compare the receipt to a traditional 30-second product spot. A small production-company quote in 2026 lands somewhere between $5,000 and $40,000 for a single deliverable, depending on talent, locations, and the size of the agency markup. The AI workflow trades two things for the price drop: control over fine craft (lighting that exactly matches the brand, an actor’s improvised moment, a real product hero shot), and the brand-safety story that comes with a human-led process. The AI version is not a strict replacement for the high-end shoot. It is a strict replacement for the bottom of the funnel: the ten variants you needed to test before you knew which one to actually shoot. That is the workflow this post is for.

What we’d change next time

This post is itself the proof that the section works: the previous version’s wishlist was “single video model with locked storyboard frames” and “spend a real generation on the CTA card”, and both, executed, made this run cheaper, faster, and more consistent than the mixed-model one. So, the new list. Voiceover pacing: the closing TTS line came back at four seconds for a three-second slot, and with no duration control the fix was starting it early over the previous beat — fine, but it should be a parameter. Music duration: same complaint, $0.15 of tuition. The grinder frame’s invented side panel and the hook clip’s mid-motion label smear are the two visual seams; both are one re-roll from gone, and next time the storyboard pass gets a consistency check against the hero frame before any video is generated.

The most useful pattern out of this whole exercise was not any single tool. It was the discipline of writing the brief first, the storyboard second, and the video third, in that order, with the LLM as a beat-sheet collaborator and the image model as a contract for the video model. The same discipline rescues the workflow whether you are running Veo, Kling, Seedance, or whatever ships next quarter. The tools were different when this post was first published two months ago; the shape of the work is identical. We intend to keep re-shooting it: the slug stays, the dates, tools, and receipts update.

For the broader take on which video models earned their keep this season, our spring 2026 catalog runs the comparison across a longer test set, and the July head-to-head pits the current three against identical prompts. If you just want the spec sheets, you can line Seedance 2.0 and Veo 3.1 up side by side on the compare page. For the audio picks behind the voiceover and music steps, see the audio catalog. And for the underlying argument about why a curated five-tool stack beats a sprawling subscription pile, the curation post is the place to start.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app