Text-to-video in 2026: what a sentence gets you now
Twelve first takes across four models, $9.02 billed, no rerolls: the shots that stunned us and the pianist who grew a third hand.
Type one sentence, wait about a minute, and a video file lands in your folder: a rainy night market, neon in the puddles, a wok throwing flame, ambient street noise in sync. It costs less than a dollar. That is not a demo reel claim; we generated that exact clip while writing this, and it is embedded below. The part the demo reels leave out is what the other takes looked like, because every viral AI clip you have seen was the best of many attempts, picked by a human with taste and a budget.
So this post publishes the contact sheet instead: twelve clips across the four text-to-video models that matter in August 2026, first result kept every time, no rerolls, no seeds, every cost and wait time printed. The wins are real and so are the three-handed pianist, the floating wine glass, and the chalkboard that reads CNCKNY POKLCK. By the end you will know what a sentence actually buys, and, more useful, which sentences to spend it on.
Every clip you’ve seen was the best of many takes
A film set solves this honestly: shoot forty takes, print one, and nobody calls the poster a lie. The showreels that sell AI video do the same thing while implying they didn’t. The gap between the take you see and the takes you don’t is the single most useful thing to understand about this technology in 2026, because it decides both your budget (you will pay for the misses) and your workflow (someone has to watch them all).
Our method for this post is a contact sheet with receipts. We picked four models you can buy through one API today, each from a lab inside the top ten of the Artificial Analysis text-to-video arena: Kling 3.0 Standard from Kuaishou, Seedance 2.0 Fast from ByteDance, HappyHorse 1.1 from Alibaba (the model that topped the arena anonymously in April before Alibaba claimed it), and Veo 3.1 Fast from Google. Two prompts went to all four models unchanged. Four more prompts, each designed to press on a known weakness, went to one model each. First result kept, every time. If a clip is good, it earned it; if it is broken, you get to see that too.
Atmosphere is a solved problem
The first shared prompt is the kind of sentence the demo reels are made of, stated in full: “A street food vendor tosses noodles in a flaming wok at a rainy night market, neon signs reflected in the puddles, steam and sparks rising, handheld camera slowly pushing in, ambient market sound.” Mood, weather, light, motion, camera direction, sound. No counts, no text, no rules. This is the genre these models were born for.
Squint past the details and all four clips are usable footage. Light, weather, texture, and camera behavior, the things that took a location scout and a rain rig two years ago, are now table stakes at under fifteen cents a second. If your need is atmosphere, b-roll, a mood, a world, text-to-video in 2026 simply works, and the cheapest model in this test delivers it as convincingly as the priciest. That is the honest good news, and it is genuinely news; none of these clips would have been possible at any price from a text prompt in early 2024.
Counts, hands, and chalkboards are not
The second shared prompt looks easier and is much harder: “A baker in a white apron places exactly four croissants on a wooden counter, then writes the words FRESH TODAY on a small chalkboard sign, static camera, soft morning light.” A number that must stay true, hands doing fine work, and five letters of on-screen text. Three of the oldest failure modes in generative video, in one domestic sentence.
Score the sentence strictly and three of four models fail it, each in a different way: painted-over text, a dodged instruction, morphing letters. The underlying pattern is that these models paint outcomes rather than execute procedures; writing is a procedure with a visible intermediate state at every frame, and counts drift because nothing inside the model keeps a tally. Then there is the surprise: the newest model in the test passed the whole thing, first take. One take is one data point, not a benchmark, and we would still not put a brand launch on it. But it says the hard column in this table is improving faster than the folk wisdom about AI hands and text has updated.
The plain numbers: about a dollar for six seconds
Here is what the sentence actually costs, from our own bills rather than pricing pages. For orientation: Google’s list prices run from $0.05 per second for Veo 3.1 Lite at 720p to $0.40 for full Veo 3.1, and the Runware model pages quote per-run examples that match what we were charged to the cent.
Three numbers deserve a highlight. First, the money: a five-to-six second 720p shot with sound costs $0.63 to $0.90 across this whole field. Call it a dollar a take. A hundred-dollar budget buys you a hundred and some takes, which, at demo-reel hit rates, is a real afternoon of production. Second, the wait: 50 to 141 seconds per clip means iteration speed varies almost 3x between vendors, and across an afternoon that gap decides how many ideas you get to try. Third, the ceiling: every model here caps a single generation between 8 and 15 seconds. Anything longer than a shot is an editing job, not a prompt.
For the wider context of what a mixed hour of AI costs, video included, our itemized-hour post found one video clip dominating the whole bill. These per-second rates are why: video remains the most expensive sentence you can type, even after this year’s price drops.
The failure tour: four ways a good clip goes wrong
The bake-off showed the everyday misses. For the tour, we wrote four prompts that each press on one documented weakness, and gave one to each model. No trickery in the wording; these are sentences a normal user would type on week one. This section is the part of the contact sheet the showreels burn.
None of these are cherry-picked disasters; they are first takes on ordinary sentences, and each failure is a category, not a fluke. Object counts drift, procedures render as outcomes, physics is painted rather than simulated, and long shots slowly forget the world they started in. Video models are improving fast on all four axes, and every one of these clips would have been much worse a year ago. But if your shot depends on one of these axes holding, budget takes, or restructure the shot so it doesn’t.
What already ships today
Knowing the failure map, the practical question flips: which real work clears it? Three genres already do, daily, in production.
First, mood and b-roll: establishing shots, weather, texture, the rainy market above. Editors cut around specifics anyway, so the failure modes never show. Second, product loops: one object, slow camera moves, seamless backgrounds; there is nothing to miscount and nobody’s hands are on screen. We built a full 30-second ad this way for $3.73, and the method still holds. Third, stylized shorts: animation, surreal, dream-logic content where physics slips read as style. What does not ship from a single sentence: dialogue scenes that must match a script across cuts, brand text on screen, product demos where the object must stay exactly your object, and anything longer than one shot. Those need pipelines, reference images, and edits, not better adjectives.
Prompt like a director, and budget for take two
The single highest-leverage skill in the clips above is not model choice; it is shot language. Our market prompt worked across all four models because it reads like a shot list: subject, action, setting, light, camera move, sound. Vague prompts produce the mush in the middle of the model’s training data; specific camera direction (“handheld, slowly pushing in”) was executed by three of the four models almost literally. The same grammar that fixes text prompts fixes video prompts, with one addition: describe outcomes, not procedures. A sign that already says the words beats an actor writing them, in every model we have ever tested.
Then budget like a producer. Our twelve takes cost $9.02 and about eighteen minutes of waiting; a realistic project multiplies that by a hit rate. On easy genres, expect most takes to be usable. On hard shots, one keeper in three to five takes matched our experience here, which turns a dollar a take into three to five dollars a shot, still absurdly cheap by any production standard and still worth planning for. When the model matters more than the prompt, our head-to-head method post shows how to run your own bake-off for under ten dollars, and the four models here each have a live spec page in the CSuite catalog.
Date-stamp everything in this post, including the praise. These numbers were billed on August 3, 2026; the roster came from the arena the same day. The last time we wrote about this market, one of its three famous names was already a shutdown notice. Six months from now the prices will be lower, the failure tour will be shorter, and some model in this post will have been replaced by its own maker. The contact-sheet habit is the part that will still be true: first takes, receipts, and your own eyes over anyone’s showreel, ours included.


