How to Measure AI’s Environmental Cost Without Slogans
For a useful estimate of AI’s environmental cost, first ask what activity is being measured, where the measurement begins and ends, and what unit is reported. Training a model and answering a prompt are different workloads; electricity use is not the same as emissions; and a figure for one prompt cannot stand in for a data centre’s total demand. A practical comparison records these distinctions before drawing conclusions.
Start by separating training from inference
Training is the computation used to create or update a model. Inference, also called serving, is the computation used to produce an output after deployment. A single training run may consume substantial energy over a defined period, while inference is repeated across users and requests. Over time, the total from repeated serving can become relevant alongside training, but the relative shares depend on model, usage, and the period counted. The two should be measured and reported separately before any combined total is presented. The International Energy Agency describes both training and deployment as data-centre activities, and a Google technical study specifically measures inference while leaving model training outside its boundary. IEA, “Energy demand from AI” and Elsworth et al., “Measuring the environmental impact of delivering AI at Google Scale”
For training, useful reporting includes the run’s duration, energy consumed, hardware involved, and whether the figure covers only the main accelerators or the wider computing system. For inference, define the service and task: for example, text generation for a named assistant during a stated measurement period. Include whether the reported unit is per request, per token, per generated output, or total service energy. These denominators answer different questions. A per-request value can vary with prompt and response length, model routing, and workload conditions; a per-token value can hide overheads that do not scale directly with tokens.
Check the boundary before comparing figures
The measurement boundary identifies which equipment and activities are counted. A narrow estimate might count electricity drawn by active GPUs or other AI accelerators. A broader serving estimate may also include host CPUs and memory, idle machines reserved for availability, power conversion, cooling, and other data-centre overhead. Some boundaries exclude the network outside the facility, end-user devices, model training, or data storage. An estimate can be internally valid while still being unsuitable for comparison with one that counts more—or different—parts of the system.
Google’s published Gemini Apps estimate illustrates why the boundary matters. Its May 2025 estimate for the median text prompt was 0.24 watt-hours (Wh) using its comprehensive serving method; a narrower method in the same study, which excluded host CPU and memory, idle machines, and facility overhead, produced 0.10 Wh. Those are results for one product, date, metric, and accounting boundary, not universal values for an AI prompt. The study also excludes training and end-user devices. Google Cloud, “Measuring the environmental impact of AI inference” and the accompanying technical paper
When reading any per-task number, look for a short boundary statement. It should say whether the total is accelerator-only or full-stack, whether data-centre overhead is included, and whether training and non-data-centre activities are included. If that information is missing, treat the number as incomplete for comparison. Do not silently add an overhead multiplier from another source: facility efficiency and workload conditions differ.
Keep the units attached to the claim
Electricity is commonly reported in watt-hours (Wh) for a small task, kilowatt-hours (kWh) for larger totals, and megawatt-hours or terawatt-hours for facility or regional totals. One kWh equals 1,000 Wh. Power, measured in watts (W) or kilowatts (kW), is a rate at a moment or over an interval; energy in Wh or kWh accumulates across time. A statement about a data centre’s power capacity therefore does not, by itself, state how much electricity it used during a year.
Emissions are usually expressed as mass of carbon dioxide equivalent, such as grams of CO₂e per task or tonnes of CO₂e per year. To estimate electricity-related emissions, energy use is combined with an emissions factor for electricity, whose value depends on the grid, time period, and accounting method. Embodied emissions from manufacturing equipment are a separate component and may be allocated across hardware lifetime or usage under a chosen method. Keep energy and emissions as distinct results; a lower emissions figure may reflect a less carbon-intensive electricity supply, not lower electricity use. Water use is another distinct metric and needs its own boundary and unit.
Treat hardware efficiency as one input, not the whole result
Hardware efficiency can be expressed as useful computation per unit of energy, or as energy per unit of useful work. But a chip’s peak performance per watt is not the same as the energy used to complete a real service task. Real-world results also depend on utilization, software and model design, workload size, batching, system support equipment, and the facility. The IEA notes that servers account for a large share of data-centre electricity on average, while the shares for cooling and other components vary substantially by facility type. IEA, “Energy demand from AI”
For a personal comparison, prefer the same defined task on the same service and period, with the provider’s measured energy per task if available. If comparing two models or services, check that task, output length, measurement boundary, and treatment of idle capacity are comparable. A hardware specification or benchmark can help explain efficiency potential, but it does not independently establish production energy use.
Use data-centre totals for scale, not per-prompt precision
A facility or national total answers a different question from a per-request estimate. Total electricity consumption reflects all workloads served, facility equipment, and changes in demand over a chosen interval. Dividing such a total by an estimated number of requests may create a rough average only if the numerator and denominator cover the same services, locations, and period. Shared infrastructure also serves non-AI workloads, so assigning all facility electricity to AI would overstate AI’s share unless the allocation method justifies it.
The U.S. Department of Energy’s 2024 Lawrence Berkeley National Laboratory report estimates historical U.S. data-centre electricity use and presents a range of future demand scenarios through 2028. That is a national data-centre analysis, not a direct measurement of the energy for an individual AI request. Its scale and scenario approach are useful for understanding uncertainty in aggregate demand, while product-level operational telemetry is more relevant to a narrowly defined serving task. Lawrence Berkeley National Laboratory, “2024 United States Data Center Energy Usage Report”
Describe uncertainty rather than hiding it
Uncertainty can enter through incomplete access to operational data, estimated rather than metered workloads, assumptions about idle capacity, changing hardware utilization, and choices about how shared equipment is allocated. Emissions add further variation through electricity location and timing, emissions factors, and the accounting treatment of purchased electricity. Data-centre totals also rely on estimates and scenarios when complete metered data are unavailable. The IEA explicitly notes substantial uncertainty in current and future data-centre electricity consumption. IEA, “Energy demand from AI”
A useful report therefore names its observation period, describes the boundary, identifies whether figures are directly metered or modelled, and states key assumptions. Where estimates depend on scenarios or allocation choices, give a range or explain how those choices affect the result. Avoid presenting many decimal places when the underlying measurement does not support them. A median, mean, or representative task should also be labelled accurately: each summarizes a workload distribution differently, and none describes every request.
A practical checklist for evaluating a claim
Before repeating or comparing an environmental figure for AI, ask:
Is it training, inference, or a combined total?
What service, workload, period, and functional unit does it describe?
Does it count only accelerators, or also hosts, idle capacity, and facility overhead?
Is the value energy, power, emissions, or water—and what units are used?
For emissions, what electricity factor, location, time basis, and embodied-emissions treatment were applied?
Is the result measured, estimated, or scenario-based, and what important exclusions or uncertainties remain?
For ordinary personal choices, this checklist is more actionable than a decontextualized “energy per chat” number. It helps distinguish a measured result for one defined service from a broad claim about AI as a whole, and it shows which missing details would make a comparison unreliable.
