How much energy does AI actually use? Your prompt vs. the grid
One prompt costs nine seconds of TV. The fleet serving them will out-consume Japan by 2030. Both are true; here's how they fit.
In August 2025, Google did something AI companies had spent years avoiding: it published a meter reading. The median Gemini text prompt, measured across its production fleet, consumes 0.24 watt-hours of electricity: about nine seconds of watching TV, with five drops of water for cooling. Eleven months later, households across thirteen eastern US states are watching electric bills climb roughly 15%, with data centers taking most of the public blame.
Both facts are real, and most arguments about AI and energy come from people holding only one of them. Prompt-level guilt is innumerate: your chatbot habit is a rounding error inside your own house. Grid-level complacency is innumerate too: the fleet serving everyone’s prompts is on track to out-consume Japan by 2030. This post walks through the honest numbers at both scales, what a video prompt costs compared to a text one (the gap is about 4,000x), and which of the two scales your decisions actually touch.
Three rival labs landed on the same number
For years the standard citation was that one ChatGPT query used about 3 Wh, ten times a Google search. That figure came from early-2023 assumptions about GPT-3.5-era hardware, and it survived in headlines long after the hardware died. In February 2025 the research group Epoch AI rebuilt the estimate from scratch using current GPUs and realistic token counts and got roughly 0.3 Wh, a tenth of the legacy number. In June 2025 Sam Altman put OpenAI’s own average at 0.34 Wh. Two months later came Google’s 0.24 Wh, the first figure measured rather than modeled.
The convergence is what makes the number trustworthy. A vendor with every incentive to flatter itself, a rival’s CEO, and an independent estimator with no stake all landed within a tenth of a watt-hour. Google’s methodology is also the strictest of the three: the 0.24 Wh includes the host CPU and RAM, the idle machines held in reserve for failover, and the data center’s cooling overhead. Count only the AI chip doing the math and the figure drops to 0.10 Wh. The rest is overhead, and Google’s overhead is unusually lean: a fleet-wide power usage effectiveness of 1.09, meaning 9% extra for cooling and power distribution on top of the computing itself. Older or smaller facilities commonly run 50% or more.
What killed the 3 Wh figure was not a correction; it was time. The estimate described 2023 hardware serving 2023 models with 2023 software. Two hardware generations, better serving techniques, and smaller specialized models later, the measured number is a tenth of it, and it will be stale too. For anything that moves as fast as AI hardware, recency matters more than pedigree: an energy claim without a date on it is wrong.
Text is cheap. Video is a different animal
One number does not cover AI, though, because “AI” is not one workload. The useful mental model is a ladder with a huge jump at the top rung. MIT Technology Review, working with university and Hugging Face researchers, measured open models across modalities in May 2025. A small 8-billion-parameter text model answered for around 0.03 Wh. A large one needed about twenty times more, and complex prompts cost up to nine times as much as simple ones, because energy scales with the tokens processed (the same reason reasoning modes that think in long chains cost more). A high-quality image from Stable Diffusion 3 landed near 0.6 to 1.2 Wh, in text-prompt territory.
Then comes video. Generating a single five-second, 16-frames-per-second clip with the open CogVideoX model consumed about 3.4 megajoules: roughly 950 Wh, or four thousand text prompts, or a microwave running for about an hour. Commercial video models publish no numbers and are likely more efficient per clip, but the shape of the ladder is not in dispute. When we priced a real mixed hour of AI use, one video clip was 88% of the bill; energy follows the same skew.
The practical reading: nobody needs to ration chat messages, and nobody should treat “generate 40 video variations and pick one” as free. Text is the nine-seconds-of-TV rung. Video is the run-the-microwave-for-an-hour rung. They deserve different reflexes.
Your chatbot habit is a rounding error at home
Put the prompt number against the only meter you control. The average US household runs through about 28,000 Wh per day. A heavy user firing 100 text prompts a day adds 24 Wh: less than a tenth of one percent of the house, and a full year of it totals under 9 kWh, which your home consumes in about seven and a half hours of just being a home. Skipping one dryer cycle offsets months of prompting. Even MIT Technology Review’s deliberately heavy scenario (15 text queries, 10 image generations, and 3 video attempts in a day) came to 2.9 kWh, a tenth of a typical household’s daily use, and the videos were nearly all of it.
This is why “think before you prompt” advice, offered as climate virtue, misleads more than it helps. The personal arithmetic only turns real at the video rung, and even there it takes volume. If individual restraint is your instrument, it is pointed at the wrong scale. The scale where the numbers turn real is the next section.
The fleet is the story: a second Japan by 2030
Multiply a tiny number by a planet of users, then add the training runs, the video workloads, and the reserve capacity, and the sum stops being tiny. The International Energy Agency put global data center consumption at 415 TWh in 2024, about 1.5% of world electricity, and projects roughly 945 TWh by 2030: a doubling, to approximately Japan’s entire annual consumption, with AI as the most important driver. Data center demand has grown about 12% per year since 2017, four times faster than electricity use overall. By 2035 the IEA’s base case reaches about 1,200 TWh, with a scenario band of 700 to 1,700 TWh wide enough to hold two different futures.
One framing note: everything above and below measures inference, the serving of answers. Training a frontier model is a separate, enormous, one-time cost. But a trained model serves billions of queries over its life, so the ongoing draw of the fleet, the thing the grid has to plan for, is dominated by inference plus the continuous churn of new training runs, not by any single model’s creation.
The US carries the largest share: 45% of global data center consumption. Lawrence Berkeley National Laboratory measured US data centers at 176 TWh in 2023, 4.4% of national electricity, and projects between 6.7% and 12% by 2028. China holds another 25% of the global total and Europe 15%, so this is a three-region story with one loud protagonist. Note the width of the Berkeley range: nearly a factor of two, from a national lab, over a five-year horizon. A May 2026 comparison of forecasters shows the same spread for 2030: between 325 and 793 TWh for the US depending on whose model you believe. Everyone agrees on the direction; nobody agrees on the slope. Any confident single number about AI’s future energy use is someone’s scenario wearing a fact’s clothing.
Power bills are rising, and the blame is messy
The place most people now meet AI’s energy story is their own electric bill. In PJM, the grid serving 67 million people across thirteen eastern US states, household bills are running about 15% higher in 2026 than before the data center boom, roughly $25 to $30 a month. The mechanism is the capacity auction, the market that pays generators to guarantee supply during peaks: its clearing price jumped from $29 to $270 per megawatt-day in two years as forecasted data center demand collided with retiring power plants.
Before filing this entirely under “AI did it,” read the dissent in the same SemiAnalysis piece. Texas absorbed a comparable AI buildout on its own grid, ERCOT, and wholesale prices moved only a few percent. PJM’s own demand forecasts have been revised downward two years running, forward energy markets price nothing like a ninefold squeeze, and some new data centers generate most of their own power. On that reading, the bill spike is demand growth colliding with slow interconnection queues and a capacity market that panics on forecasts, not a simple AI surcharge. Both readings agree on what matters here: the costs are real, they arrive through market plumbing rather than your prompt count, and the fix is infrastructure and market design, not chat abstinence.
Efficiency is winning per prompt, losing in total
Here is the twist that breaks most intuitions: the per-prompt number is collapsing while the total climbs. Google reports the energy of a median Gemini prompt fell 33x in the twelve months to May 2025, with carbon per prompt down 44x, from better models, better serving, and cleaner power. Epoch’s revision cut the old ChatGPT estimate tenfold. Chips keep improving. Every curve that measures energy per task points down and to the right.
Economists have a name for why the total rises anyway: the Jevons paradox. When something useful gets cheaper, we use enough more of it to swamp the savings. A 33x cheaper prompt does not mean 33x less energy; it means prompts get embedded in every search box, every spreadsheet, every camera roll, and the fleet grows. That is exactly what the IEA’s doubling projection assumes, and it is why per-prompt efficiency, genuinely excellent news, is not a reason to stop watching the aggregate. The two curves are not in contradiction. One is a ratio, the other is a sum.
The two numbers worth remembering
If you keep only two numbers from this post, keep these: a text prompt is about 0.3 Wh, and the world’s data centers are headed for about 945 TWh a year by 2030. The first number tells you how to feel about your own usage: calm. Prompt freely, and save your scrutiny for high-volume video generation, the one rung where a person can spend meaningful energy. The second tells you where the real decisions live: in siting, interconnection queues, capacity markets, and who pays for new generation. Those are policy and engineering questions, and treating them as personal-virtue questions lets the actual decision-makers off the hook.
Two habits fall out of the math for practitioners. Workloads that nobody is waiting on can run when the grid is slack; most AI work can wait, and vendors already discount it. And inference that runs on your own machine puts the meter where you can see it: a local model’s draw shows up on your own power bill in watts you can measure at the wall, part of the broader shift back to personal compute. The energy story of AI is not a secret and no longer needs to be argued from vibes. The meters exist. The readings are published. The only mistake left is quoting a number without its date, or its scale.


