Skip to content
CSuite
ExplainerVideo AIAI ModelsAugust 3, 20269 min read

Text-to-video in 2026: what a sentence gets you now

Twelve first takes across four models, $9.02 billed, no rerolls: the shots that stunned us and the pianist who grew a third hand.

The contact sheet, not the showreel
Every AI video you’ve seen was the best of many takes. These are our first twelve.
take 01
wok flame, posted
take 02
sign reads CNCKNY POKLCK
take 03
pianist grows a third hand
take 04
noodles levitate
take 05
stove in the bedroom
take 06
wine glass floats
take 07
FRESH TODAY, written right
take 08
night market, at noon
Twelve clips, four current models, two shared prompts and four hard ones, generated through Runware’s API on August 3, 2026. First result kept every time, no rerolls, every cost and wait time in the post. The labels above are real notes from our takes.

Type one sentence, wait about a minute, and a video file lands in your folder: a rainy night market, neon in the puddles, a wok throwing flame, ambient street noise in sync. It costs less than a dollar. That is not a demo reel claim; we generated that exact clip while writing this, and it is embedded below. The part the demo reels leave out is what the other takes looked like, because every viral AI clip you have seen was the best of many attempts, picked by a human with taste and a budget.

So this post publishes the contact sheet instead: twelve clips across the four text-to-video models that matter in August 2026, first result kept every time, no rerolls, no seeds, every cost and wait time printed. The wins are real and so are the three-handed pianist, the floating wine glass, and the chalkboard that reads CNCKNY POKLCK. By the end you will know what a sentence actually buys, and, more useful, which sentences to spend it on.

Every clip you’ve seen was the best of many takes

A film set solves this honestly: shoot forty takes, print one, and nobody calls the poster a lie. The showreels that sell AI video do the same thing while implying they didn’t. The gap between the take you see and the takes you don’t is the single most useful thing to understand about this technology in 2026, because it decides both your budget (you will pay for the misses) and your workflow (someone has to watch them all).

Film photographers printed every frame on one sheet and circled the keeper. AI video needs the same habit. Photo by Markus Spiske on Unsplash.

Our method for this post is a contact sheet with receipts. We picked four models you can buy through one API today, each from a lab inside the top ten of the Artificial Analysis text-to-video arena: Kling 3.0 Standard from Kuaishou, Seedance 2.0 Fast from ByteDance, HappyHorse 1.1 from Alibaba (the model that topped the arena anonymously in April before Alibaba claimed it), and Veo 3.1 Fast from Google. Two prompts went to all four models unchanged. Four more prompts, each designed to press on a known weakness, went to one model each. First result kept, every time. If a clip is good, it earned it; if it is broken, you get to see that too.

Atmosphere is a solved problem

The first shared prompt is the kind of sentence the demo reels are made of, stated in full: “A street food vendor tosses noodles in a flaming wok at a rainy night market, neon signs reflected in the puddles, steam and sparks rising, handheld camera slowly pushing in, ambient market sound.” Mood, weather, light, motion, camera direction, sound. No counts, no text, no rules. This is the genre these models were born for.

Veo 3.1 Fast, first take: $0.90 for six seconds with audio, back in 50 seconds. Coherent street, rain that reads as rain, a scooter passing behind, and the push-in we asked for. The wok flame blooms into a small fireball mid-clip, which a human operator would call a safety incident and a marketer would call cinematic.
Kling 3.0 Standard, first take: $0.63 for five seconds with audio. The most photoreal vendor of the four, wet-street texture to match, and a flame that briefly detaches from the wok and hovers. The neon signs are confident gibberish, pseudo-Chinese glyphs no dictionary contains.
Seedance 2.0 Fast, first take: $0.65 for five seconds with its always-on audio, and the slowest wait of the day at 137 seconds. A beautifully dressed stall, symmetric framing, and one physics-free moment where the entire noodle mass rises out of the wok and hangs in the air like a planet.
HappyHorse 1.1, first take: $0.72 for five seconds. The most documentary rain and skin of the four, and a neon sign in convincing fake Thai. One big deviation: we asked for a night market and got noon in a downpour. Atmosphere delivered, brief ignored.

Squint past the details and all four clips are usable footage. Light, weather, texture, and camera behavior, the things that took a location scout and a rain rig two years ago, are now table stakes at under fifteen cents a second. If your need is atmosphere, b-roll, a mood, a world, text-to-video in 2026 simply works, and the cheapest model in this test delivers it as convincingly as the priciest. That is the honest good news, and it is genuinely news; none of these clips would have been possible at any price from a text prompt in early 2024.

Counts, hands, and chalkboards are not

The second shared prompt looks easier and is much harder: “A baker in a white apron places exactly four croissants on a wooden counter, then writes the words FRESH TODAY on a small chalkboard sign, static camera, soft morning light.” A number that must stay true, hands doing fine work, and five letters of on-screen text. Three of the oldest failure modes in generative video, in one domestic sentence.

Veo 3.1 Fast: gorgeous flour-dusted baker, warm light, and a chalkboard that comes up already written, reading CNCKNY POKLCK. The croissant count drifts from three to five as the clip plays. Every individual frame is lovely; the sentence’s two hard constraints both fail.
Kling 3.0 Standard: frames the shot from the neck down, places croissants one at a time, and lands on exactly four. The board arrives pre-printed with FRESH and the baker only mimes writing, so TODAY never appears. Half right, and the most literal read of the instruction in the test.
Seedance 2.0 Fast: window light a cinematographer would sign, four croissants at the end, and a chalkboard that starts at FRESH and gains TODAY letter by letter while the hand hovers nearby, chalk never touching slate. The words morph between frames like wet ink.
HappyHorse 1.1: the only model to pass the whole sentence. Four croissants land on the counter and stay four, then the hand writes FRESH, then TODAY, stroke by stroke, letterforms wobbly but correct. The arena’s dark horse earned its ranking on the first take.

Score the sentence strictly and three of four models fail it, each in a different way: painted-over text, a dodged instruction, morphing letters. The underlying pattern is that these models paint outcomes rather than execute procedures; writing is a procedure with a visible intermediate state at every frame, and counts drift because nothing inside the model keeps a tally. Then there is the surprise: the newest model in the test passed the whole thing, first take. One take is one data point, not a benchmark, and we would still not put a brand launch on it. But it says the hard column in this table is improving faster than the folk wisdom about AI hands and text has updated.

The plain numbers: about a dollar for six seconds

Here is what the sentence actually costs, from our own bills rather than pricing pages. For orientation: Google’s list prices run from $0.05 per second for Veo 3.1 Lite at 720p to $0.40 for full Veo 3.1, and the Runware model pages quote per-run examples that match what we were charged to the cent.

What we were billed · 720p with audio where offered · August 3, 2026
Model
Per sec
Our clip
Kling 3.0 Standard
$0.126
$0.63 / 5 s
Seedance 2.0 Fast
$0.131
$0.65 / 5 s
HappyHorse 1.1
$0.145
$0.72 / 5 s
Veo 3.1 Fast
$0.15
$0.90 / 6 s
Per-second rates are our billed clip cost divided by its length, through Runware’s API. Waits are wall-clock from request to file. Rates change monthly in this market; treat the spread, not the digits, as the durable fact.

Three numbers deserve a highlight. First, the money: a five-to-six second 720p shot with sound costs $0.63 to $0.90 across this whole field. Call it a dollar a take. A hundred-dollar budget buys you a hundred and some takes, which, at demo-reel hit rates, is a real afternoon of production. Second, the wait: 50 to 141 seconds per clip means iteration speed varies almost 3x between vendors, and across an afternoon that gap decides how many ideas you get to try. Third, the ceiling: every model here caps a single generation between 8 and 15 seconds. Anything longer than a shot is an editing job, not a prompt.

The receipt, August 3, 2026
Every video generation billed for this post, with per-task cost reporting on. No hidden rerolls; three failed HappyHorse requests (a parameter our script sent that the model rejects) cost $0.00.
2 × Veo 3.1 Fast clips$1.80
3 × Kling 3.0 Standard clips$1.89
3 × Seedance 2.0 Fast clips$1.96
3 × HappyHorse 1.1 clips$2.17
1 × Veo 3.1 Fast long take$1.20
Total$9.02

For the wider context of what a mixed hour of AI costs, video included, our itemized-hour post found one video clip dominating the whole bill. These per-second rates are why: video remains the most expensive sentence you can type, even after this year’s price drops.

The failure tour: four ways a good clip goes wrong

The bake-off showed the everyday misses. For the tour, we wrote four prompts that each press on one documented weakness, and gave one to each model. No trickery in the wording; these are sentences a normal user would type on week one. This section is the part of the contact sheet the showreels burn.

Hands, Seedance 2.0 Fast: “Overhead close-up of a pianist’s hands playing a fast run on a grand piano, all fingers clearly visible, studio lighting.” The keys, the Steinway lettering, and the lighting are immaculate. There are three hands. A fourth flickers at the frame edge mid-run.
Physics, Kling 3.0 Standard: “A full wine glass tips off the edge of a kitchen counter, falls, and shatters on the tile floor, slow motion, side view.” The glass leaves the counter, then floats horizontally in midair while the wine stays obediently in the bowl, before teleporting to the floor as shards. Slow motion became no gravity.
Object permanence, HappyHorse 1.1: “A basketball player spins the ball on one finger, then passes it behind her back and catches it, full body, static camera, gym interior.” The gym and the player hold rock steady. The ball spins a visible inch above her fingertip, never touching it, and the behind-the-back pass resolves with the ball in her hands a beat before it should arrive.
Continuity, Veo 3.1 Fast, 8 seconds, $1.20: “One continuous shot follows a golden retriever running from the front door of a house, through the kitchen, and up the stairs into a bedroom.” The strongest result of the tour: rooms connect, the camera flows. But watch the finale: the bedroom has a stainless kitchen stove against the wall, and the retriever arrives noticeably shaggier than it left, halfway to a collie.

None of these are cherry-picked disasters; they are first takes on ordinary sentences, and each failure is a category, not a fluke. Object counts drift, procedures render as outcomes, physics is painted rather than simulated, and long shots slowly forget the world they started in. Video models are improving fast on all four axes, and every one of these clips would have been much worse a year ago. But if your shot depends on one of these axes holding, budget takes, or restructure the shot so it doesn’t.

What already ships today

Knowing the failure map, the practical question flips: which real work clears it? Three genres already do, daily, in production.

The product loop: one hero object, controlled motion, no hands, no text, no counting. The genre AI video was born ready for. Illustration generated with Seedream 4.5 via Runware.

First, mood and b-roll: establishing shots, weather, texture, the rainy market above. Editors cut around specifics anyway, so the failure modes never show. Second, product loops: one object, slow camera moves, seamless backgrounds; there is nothing to miscount and nobody’s hands are on screen. We built a full 30-second ad this way for $3.73, and the method still holds. Third, stylized shorts: animation, surreal, dream-logic content where physics slips read as style. What does not ship from a single sentence: dialogue scenes that must match a script across cuts, brand text on screen, product demos where the object must stay exactly your object, and anything longer than one shot. Those need pipelines, reference images, and edits, not better adjectives.

Prompt like a director, and budget for take two

The single highest-leverage skill in the clips above is not model choice; it is shot language. Our market prompt worked across all four models because it reads like a shot list: subject, action, setting, light, camera move, sound. Vague prompts produce the mush in the middle of the model’s training data; specific camera direction (“handheld, slowly pushing in”) was executed by three of the four models almost literally. The same grammar that fixes text prompts fixes video prompts, with one addition: describe outcomes, not procedures. A sign that already says the words beats an actor writing them, in every model we have ever tested.

The jobs on this set that survive into the AI version: deciding the shot, judging the take, and saying “again.” Photo by Jakob Owens on Unsplash.

Then budget like a producer. Our twelve takes cost $9.02 and about eighteen minutes of waiting; a realistic project multiplies that by a hit rate. On easy genres, expect most takes to be usable. On hard shots, one keeper in three to five takes matched our experience here, which turns a dollar a take into three to five dollars a shot, still absurdly cheap by any production standard and still worth planning for. When the model matters more than the prompt, our head-to-head method post shows how to run your own bake-off for under ten dollars, and the four models here each have a live spec page in the CSuite catalog.

Date-stamp everything in this post, including the praise. These numbers were billed on August 3, 2026; the roster came from the arena the same day. The last time we wrote about this market, one of its three famous names was already a shutdown notice. Six months from now the prices will be lower, the failure tour will be shorter, and some model in this post will have been replaced by its own maker. The contact-sheet habit is the part that will still be true: first takes, receipts, and your own eyes over anyone’s showreel, ours included.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app