What one AI image actually costs: cloud pricing decoded
We bought the same product shot from six AI models. The spread was 167x, and the best image wasn't the priciest. The decoder ring inside.
We bought the same photograph six times on the same afternoon. Same prompt, same fixed seed where the model would take one: a bottle of cold brew coffee on wet slate, moody lighting, the kind of shot a small brand pays a studio for. The cheapest version cost $0.0006. The most expensive cost $0.10, which is 167 times more. Neither was the best one.
If that spread surprises you, it’s because generation pricing is quoted in units designed to resist comparison: per image, per megapixel, per token, per GPU-second, per credit. This post translates the zoo into plain cents, prices video and speech the same way, adds the multiplier nobody prints (retries), and then does the math everyone eventually asks for: at what volume does buying your own GPU beat renting the cloud? Every number is dated August 5, 2026, and most of them we paid, not read.
The confusing units are the real obstacle
Nobody sells you an image for cents, even though that’s what you are buying. Replicate prices most image models per output image, which is the honest unit, but drops you to per-GPU-second billing for anything custom. Black Forest Labs prices FLUX 2 Max per megapixel with a surcharge for extra resolution. OpenAI prices images in output tokens, at $32 per million for gpt-image-1.5, and points you to a calculator to learn what a picture costs. ElevenLabs sells credit packs where one credit is roughly one character of speech. Same product, five currencies.
Credits deserve special suspicion, because they add a second exchange rate on top of the first. A credit pack converts dollars to credits, then credits to output, and both rates are set by the seller. They also expire: on most subscription plans, unused monthly credits vanish or roll over with limits, so the effective per-minute price of a lightly used plan is far worse than the advertised one. ElevenLabs’ free tier sounds generous until you convert it: 10,000 credits is roughly ten minutes of speech a month, one podcast intro and change.
The fix is one habit: before comparing anything, divide the dollars by the finished outputs and write down cents per image, cents per second of video, or cents per minute of speech. That single conversion is most of this post’s work, and it’s the part you can reuse long after these specific numbers expire.
One image costs between a twentieth of a cent and a quarter
Translate everything into per-image terms and the market sorts itself into three tiers. Draft tier, under a cent: FLUX Schnell billed us $0.0006 through Runware and lists at $3 per thousand on Replicate; GPT Image 2’s low tier is $0.006. Work tier, two to seven cents: FLUX Dev, Qwen Image 2, Seedream 4.5, Imagen 4, Nano Banana 2. Premium tier, nine cents to a quarter: Ideogram V3 Quality, Gemini 3 Pro Image, and GPT Image 2 at its high setting, $0.211.
Two footnotes worth reading twice. First, Google’s pricing page lists Imagen 4 as shutting down on August 17, 2026, twelve days after this post: cloud price lists are not just volatile, the products underneath them are. Second, resolution quietly multiplies cost on several models. The same Nano Banana 2 render that costs $0.034 at 1K billed us $0.10 at 2K in July. When a price looks surprisingly good, check which resolution it’s quoting.
How to shop the tiers: match the tier to the job, not to your self-image. Drafts, moodboards, and anything you will regenerate twenty times belong in the under-a-cent tier, where a hundred variations cost less than a dime. Client-facing work sits comfortably in the four-cent tier; that’s where the workhorse models live. Reserve the twenty-cent tier for jobs the cheap tiers demonstrably fail: legible text inside the image, precise instruction-following, brand-critical finals. Paying premium rates for iteration is the single most common way people overspend on this technology.
Same prompt, six models: a 167x price spread
Tables tell you what providers charge. They don’t tell you what the money buys, so we ran the experiment: one prompt, “studio product photograph of a glass bottle of cold brew coffee on a wet slate slab, dramatic amber side lighting, dark moody background,” sent unchanged to six models through one API, first result kept, no rerolls.
The result that matters: price does not track quality. The $0.10 FLUX 2 Max shot is gorgeous, but the second-cheapest image, GPT Image 2 at $0.018, invented a coherent brand, set the bottle beside a bowl of beans, and produced arguably the most usable shot of the six. The $0.0006 FLUX Schnell frame has no label and flatter light, and would still pass as a draft or a moodboard frame. Meanwhile the two mid-priced models made choices you’d reject for a product brief: Qwen left the bottle open and capless, Seedream cropped to a close-up that hides the product.
What actually moves the number is not visible quality but three dials: resolution (FLUX 2 Max charges $0.07 for the first megapixel and $0.03 per extra, which is exactly why our 2 MP frame billed $0.10), quality tiers (GPT Image 2 spans $0.006 to $0.211, a 35x range inside one model), and brand premium. Cheap models are not a compromise for drafts, iterations, and anything you’ll re-generate twenty times. Spend the quarter on the final.
A second of video costs ten images; a minute of speech costs a third of one
Apply the same translation to the other modalities and the hierarchy is stark. One second of Veo 3.1 at Google’s list price is $0.40, the cost of ten Seedream images or six hundred FLUX Schnell drafts. An eight-second clip at that rate is $3.20. The budget lane is real but still expensive relative to stills: Veo 3.1 Fast billed us $0.10 a second in July, Kling 3.0 Standard $0.126, Seedance 2.0 Fast $0.131, and Replicate lists Wan 2.1 at $0.09 for 480p.
Speech runs the other way: it’s nearly free at API rates. OpenAI’s tts-1 at $15 per million characters works out to just over a cent per narrated minute. The four voiceover lines in our 30-second film billed about $0.012 total through Gemini’s Flash TTS. The expensive speech is subscription speech: ElevenLabs’ Creator plan works out to roughly $0.18 a minute, an order of magnitude over per-character API pricing, and its free tier’s ten monthly minutes evaporate on one podcast intro.
The practical consequence: for a mixed project, video is the budget and everything else is a rounding error. A 30-second promo needs about $3 to $12 of raw video at these rates, a few cents of stills to plan it, and about a cent of narration. Which is also the correct mental model for choosing where to be cheap: saving money on stills or speech changes nothing, while one avoided video re-render pays for a hundred images.
Real projects cost two to three times the sticker
Sticker prices assume every generation ships. None of ours do, and none of yours will. You reroll a hand, tighten a crop, fix a label the model misspelled, try a warmer light. When we built a 30-second film through the full pipeline yesterday, the billed total was $3.53: the script and all four stills together cost $0.16, while the four motion clips cost $3.20, so a changed mind cost four cents at the stills stage and $0.80 a take at the motion stage. The budget rule that has survived every project we’ve receipted: multiply the sticker by two to three.
Priced that way, real work is still cheap. A 40-shot product batch at four cents is $1.60 on paper, call it $5 with retries and upscales. A wedding’s worth of invitation art is a dollar or two. The 30-second promo is $10, not $3.20. The place the multiplier genuinely hurts is video, where a single changed mind costs more than a hundred stills, which is why the working discipline is to iterate in images and only animate frames you’ve already approved.
Your own GPU starts winning near 1,000 images a month
Every heavy user eventually does this math, so here it is with the assumptions visible. The hardware: an RTX 5060 Ti 16GB, a $429-list card selling from about $580 on August 5, 2026, after a year that saw it as low as $379, so check before you buy. Running an open-weights model like FLUX Dev in 8-bit form, a card in this class produces an image in roughly 40 to 55 seconds. Electricity at the August 2026 US average of 18.8¢/kWh adds about eight hundredths of a cent per image. Marginal cost: effectively zero. The card is the whole cost.
So the crossover sits near a thousand images a month: below it, rent; above it, own. Three honest asymmetries complicate the clean line. Cloud’s frontier models (GPT Image 2, Nano Banana 2, Seedream) are not downloadable at any price, so local means open-weights quality, which is excellent and still a step behind the best hosted models for text and instruction-following. Local seconds are slower than cloud seconds, 50 versus 10, and your time is worth something. But if a capable GPU already sits in your gaming PC or a recent Mac, the purchase price is sunk, the crossover collapses to zero, and iteration-heavy work becomes free in a way no credit pack can match. Our RAM and VRAM guide covers what your existing hardware can actually run.
Mac owners get a version of the same shortcut without a graphics card purchase. Apple Silicon’s unified memory runs the same open-weights image models, slower than a discrete GPU but at the same near-zero marginal cost, and a machine you bought for other reasons needs no payback math at all. The question stops being “should I buy hardware” and becomes “is the hardware I own already good enough for the drafts,” with the cloud reserved for finals. For most people producing at moderate volume, that hybrid is the cheapest honest answer: iterate locally for free, spend the four cents only on frames that survive.
The prices will rot; the arithmetic will not
Several numbers in this post will be wrong within months, and one product in the price table dies in twelve days. That’s not a caveat, it’s the finding: generation prices have been falling for two years and providers reshuffle tiers constantly, which is why locked-in subscriptions age so badly. What keeps working is the method. Convert every quote to cents per finished output. Multiply by two to three for retries. Price the project, not the image. And only run the buy-a-GPU math on your real monthly volume, not the volume you imagine having.
The number to remember: a good AI image costs about four cents today. Everything else in this post is that fact, translated.


