Google measured a median text prompt to its Gemini apps at 0.24 watt-hours of energy, 0.03 grams of CO2 equivalent and 0.26 millilitres of water. Mistral’s lifecycle assessment put a 400-token response at 1.14 grams of CO2 and 45 millilitres of water. Both figures are honest. They are also measuring different things.
How much energy does one AI prompt use?
About a quarter of a watt-hour, on the only large-scale measurement a provider has published. Google’s engineering team released a methodology paper in August 2025 putting the median Gemini text prompt at 0.24 Wh, which it describes as “equivalent to watching TV for less than nine seconds”.
That figure is far lower than most estimates circulating before it. It is also far higher than the number Google would have reported using the common shortcut, and the gap is the most useful part of the paper.
Counting only the active chip, the same prompt measures 0.10 Wh, 0.02 gCO2e and 0.12 mL. The full figure adds idle capacity held in reserve, host machine overhead, and data-centre cooling and power conversion. Google notes that calculations covering “only active machine consumption” therefore represent “theoretical efficiency instead of true operating efficiency at scale”.
| Measurement | Energy | Water |
|---|---|---|
| Google, active chip only | 0.10 Wh | 0.12 mL |
| Google, full stack | 0.24 Wh | 0.26 mL |
| Mistral, full lifecycle, 400 tokens | Not stated separately | 45 mL |
Why is Mistral’s water figure 170 times higher?
Because it is a lifecycle assessment rather than an operational measurement. Mistral’s study, published July 2025 with Carbone 4 and the French environment agency ADEME, follows ISO 14040/44 and counts the whole chain, including hardware manufacture, model training, and the electricity and water embodied upstream of the data centre.
Google’s number is narrower on purpose. It reports what the serving infrastructure consumed while answering, which is the question an operator can actually measure per request.
Neither is wrong and neither can be substituted for the other. Comparing a 0.26 mL figure with a 45 mL figure and concluding one company is 170 times more efficient is a category error, and it is the mistake almost every article about AI water use makes.
Two credible numbers for the same question can differ by two orders of magnitude when nobody agrees where the system stops.
What does training cost compared with use?
A great deal in one payment, then very little per query. Mistral reported that training Mistral Large 2 produced 20.4 kilotonnes of CO2 equivalent and consumed 281,000 cubic metres of water, measured as of January 2025.
That cost is fixed and then amortised across every request the model ever serves. A model answering billions of prompts spreads its training footprint thinly; one answering thousands does not. It is one reason the International Energy Agency’s work on energy and AI models data-centre demand at sector level rather than per query.

This is why per-prompt figures move so fast in the provider’s favour. Google reported that over the twelve months to May 2025 the energy footprint of the median prompt fell 33-fold and its total carbon footprint 44-fold, driven by model efficiency work, better serving software and a cleaner electricity supply. The grid emissions factor for a given region does much of that work on its own.
Does a falling per-prompt cost mean falling total impact?
No, and this is where the published numbers are most often misused. Per-unit efficiency and total consumption are separate quantities, and a 33-fold improvement in the first is entirely compatible with growth in the second if usage rises faster.
Usage has risen faster. This is the Jevons paradox in its ordinary form: assistant features have been added to products that previously made no model calls at all, and reasoning modes spend far more compute per answer than a plain reply. The cheap prompt is what made both possible.
The same dynamic drove the assistant button appearing in every application. Falling unit costs do not reduce spending; they remove the reason to ration.
Which numbers are worth quoting?
Only the ones that name their boundary. A figure for AI energy or water use is meaningless without knowing whether it counts idle capacity, cooling, hardware manufacture and training, and whether the prompt was text, image or video generation.
- Check the boundary. Active chip, full data centre, or full lifecycle, since the three differ by orders of magnitude.
- Check the workload. A median text prompt is not a long reasoning trace, and neither is a video clip.
- Check the date. Efficiency figures from more than a year ago describe different hardware and different serving software.
- Check who measured. Provider self-reporting and independent estimates rarely use the same method.
Very few widely quoted statistics survive those four checks. The often-repeated claim that a single chat consumes a bottle of water traces back to a 2023 preprint on AI water footprints built on data-centre averages rather than per-request measurement, and predates both papers described here.
What actually reduces the footprint?
Model size and workload choice, far more than usage habits. Running a small model where a small model will do changes consumption by more than any amount of prompt discipline, because energy scales with the compute a request triggers rather than with how politely it is worded.
Media generation is the outlier worth knowing about. Image and especially video generation cost orders of magnitude more per output than text, which is why free video tiers are so thin compared with free text and image ones. The pricing reflects the compute directly.
Running a model on your own hardware changes the arithmetic rather than removing it, since a workstation drawing several hundred watts for a slow answer is not automatically cheaper than a purpose-built data centre; the efficiency of the surrounding facility is doing most of the work. Where the electricity comes from matters as much as how much is used. Google’s carbon figure fell faster than its energy figure over the same period, which is a grid and procurement result rather than a software one.
The bottom line
On the best available measurement, a single text prompt costs about 0.24 watt-hours and a quarter of a millilitre of water. That is small. The aggregate is not small, because the number of prompts is growing faster than the cost of each one is falling.
Any figure quoted without its accounting boundary should be treated as unusable. The difference between Google’s 0.26 mL and Mistral’s 45 mL is not a disagreement about facts; it is a disagreement about where the system ends.
Frequently asked questions
How much energy does a single AI prompt use?
Google measured its median Gemini text prompt at 0.24 watt-hours, including idle capacity, host overhead and cooling. It compares that to watching television for under nine seconds. Counting the active chip alone gives 0.10 watt-hours.
How much water does an AI query use?
Google reported 0.26 millilitres per median text prompt, roughly five drops, measured at the data centre. Mistral’s full lifecycle assessment reported 45 millilitres for a 400-token response, because it also counts water embodied in hardware and training.
Why do published AI energy figures differ so much?
Because they draw the system boundary in different places. Active-chip measurements, full data-centre measurements and full lifecycle assessments answer different questions, and the results can differ by two orders of magnitude without either being wrong.
Does using AI less actually help?
Marginally, at individual scale. Choosing a smaller model and avoiding unnecessary image or video generation changes consumption far more than reducing the number of text prompts, because energy tracks the compute a request triggers.
Is AI getting more efficient?
Per prompt, sharply. Google reported a 33-fold drop in per-prompt energy and a 44-fold drop in carbon over twelve months to May 2025. Total consumption is a separate question, since cheaper prompts have driven far more of them.


