When we talk about the digital revolution, we often picture sleek data centers humming quietly in remote locations, invisible streams of data flowing through fiber-optic cables, and a world made more efficient by software. What we rarely picture is the plume of smoke from a coal-fired power plant, the water rushing through a hydroelectric dam’s turbines, or the rare minerals ripped from the earth to build specialized hardware. Yet these physical realities are the bedrock of every major computational achievement, and nowhere is that more apparent than in the training of large-scale machine learning systems.
Rui Mendes has spent years thinking about systems—how they interact, where their boundaries lie, and what we miss when we focus only on the shiny outputs. When a new model breaks records for language understanding or image generation, the headlines celebrate the benchmark scores. But Rui’s curiosity pulls him toward the less glamorous question: what was the material cost of that breakthrough? How much water, how much electricity, how much rare metal had to be extracted and transformed to make a single training run possible?
This article is not a condemnation of technology. It is an attempt to map the invisible supply chain that powers modern machine learning, to trace the environmental toll from the mine to the motherboard, and to ask whether the way we measure progress needs a deeper recalibration.
The Energy Appetite of a Single Training Run
Let’s start with the most direct metric: electricity consumption. Training a state-of-the-art model is not like running a laptop for a few hours. It involves thousands of specialized processors—GPUs or TPUs—operating in parallel for weeks or even months. Each of these chips draws hundreds of watts, and a typical cluster might contain tens of thousands of them.
A widely cited 2019 study from the University of Massachusetts Amherst estimated that training a single large natural language processing model can emit over 284 tonnes of carbon dioxide equivalent—roughly the same as five average American cars over their entire lifetimes, including manufacturing. Since then, models have grown by orders of magnitude. The compute used in the largest training runs has been doubling every 3.4 months, a pace that far outstrips Moore’s Law. What was once a 284-tonne problem is now, for the largest experiments, likely in the thousands of tonnes.
But carbon emissions are only part of the story. The electricity that feeds these clusters has to come from somewhere. In regions where the grid is still dominated by fossil fuels, the carbon intensity is high. Even in areas with significant renewable penetration, the sheer scale of demand can strain local infrastructure, sometimes leading utilities to fire up peaker plants—often natural gas—to meet the load. The location of a data center matters enormously. A training run powered by hydroelectricity in Quebec has a vastly different footprint than one plugged into a coal-heavy grid in parts of the Midwest or Asia.

The Water That Cools the Cloud
Electricity is the most visible resource, but water is the silent partner. Data centers generate immense heat, and cooling systems are essential to prevent equipment failure. Many facilities use evaporative cooling, which consumes water directly. Others rely on electricity generated by thermoelectric power plants, which themselves withdraw vast quantities of water for cooling.
A 2021 study estimated that training a single large model can consume up to 700,000 liters of freshwater—enough to fill an Olympic-sized swimming pool more than a quarter of the way. This figure includes both on-site cooling and off-site water use at power plants. In regions already facing water stress, such as parts of the southwestern United States, the siting of large data centers has become a point of tension. Residents and local governments are beginning to ask whether the economic benefits of hosting these facilities outweigh the strain on aquifers and reservoirs.
Rui Mendes finds this systems-level view compelling because it reveals interdependencies that are easy to ignore. A model trained in a data center cooled by a closed-loop system still draws water indirectly if its electricity comes from a coal or nuclear plant that uses once-through cooling. The water footprint extends far beyond the server room.
The Hardware Lifecycle: From Mine to E-Waste
Beyond operational energy and water, there is the embodied cost of the hardware itself. The GPUs and TPUs that power training runs are marvels of engineering, but they are also products of extractive industries. Rare earth elements, cobalt, tantalum, and gold are mined, refined, and shipped across the globe to fabrication plants. Mining operations are water-intensive and often leave behind toxic tailings. The refining process requires high temperatures and chemical baths, adding further energy and water demands.
Once fabricated, these chips have a limited lifespan. The relentless pace of hardware improvement means that accelerators are often retired after three to five years, even if they remain functional. E-waste is the fastest-growing waste stream on the planet, and data centers contribute a significant share. While some components are recycled, the complex mix of materials makes full recovery difficult. Many end up in informal recycling operations in developing countries, where workers are exposed to hazardous substances without proper protection.
Rui Mendes sees this as a classic systems trap: we optimize for performance per dollar or per watt at the chip level, but we fail to account for the full lifecycle costs. A more efficient chip might reduce operational energy but increase embodied energy if it requires rarer materials or more complex manufacturing. Without a full view, we risk shifting the burden rather than reducing it.

The Scaling Race and Its Environmental Multiplier
The environmental cost of training a single model is significant, but the real concern is the scaling trend. The field has been driven by a simple empirical observation: bigger models, trained on more data with more compute, tend to perform better. This has led to an arms race where each new generation of models requires an order of magnitude more resources than the last.
Consider the compute used in notable projects over the past decade. Early breakthroughs required compute budgets measured in petaflop/s-days. By 2020, the largest training runs were consuming exaflop/s-days—a thousandfold increase. The trend shows no sign of slowing. If anything, the competitive dynamics of the industry encourage overprovisioning: teams will often train multiple versions of a model, discarding all but the best-performing one. The discarded runs still consumed energy, water, and hardware lifespan.
This scaling has a multiplier effect on environmental impact. A tenfold increase in compute does not necessarily mean a tenfold increase in performance, but it does mean a roughly tenfold increase in resource consumption. The question Rui Mendes finds most pressing is whether the marginal benefit justifies the marginal cost. At what point does the pursuit of a slightly better benchmark score become an irresponsible use of shared resources?
Carbon Accounting: What Gets Measured Gets Managed
One of the challenges in addressing this issue is the lack of standardized reporting. Unlike the airline or automotive industries, where emissions are regulated and publicly reported, the computational sector has no mandatory carbon accounting. Some research labs voluntarily disclose the energy and carbon cost of their experiments, but many do not. Even when they do, the methodologies vary widely, making comparisons difficult.
A 2022 paper proposed a standardized “model card” that would include not only performance metrics but also energy consumption, carbon emissions, and hardware details. The idea is to make environmental cost a first-class consideration in model development, alongside accuracy and speed. Some conferences have begun encouraging such disclosures, but adoption remains patchy.
Rui Mendes notes that this is a classic case of externalities: the benefits of a large model are captured by the organization that trains it, while the environmental costs are distributed globally. Without mechanisms to internalize those costs—whether through regulation, market pricing, or cultural norms—the incentive to minimize them remains weak.
Geographic Arbitrage and Carbon Accounting Games
One subtle but important factor is the geographic arbitrage of carbon intensity. A training run powered by a grid with high renewable penetration will have a lower operational carbon footprint than one powered by coal. Organizations can, and do, choose locations partly on this basis. But this raises questions about additionality: does siting a data center in a region with clean energy actually reduce global emissions, or does it simply displace other consumers onto dirtier sources?
The answer depends on the specifics of the grid. In some cases, new demand is met by bringing additional renewable capacity online, which can have a genuinely positive effect. In others, the clean energy is already fully utilized, and new demand is met by fossil fuels. Without careful accounting, location choice can become a form of greenwashing—claiming low emissions based on average grid intensity while ignoring the marginal impact of new load.
There is also a temporal dimension. Training runs often last weeks, and the carbon intensity of a grid varies hour by hour. Some researchers have proposed “carbon-aware” scheduling, shifting workloads to times when renewable penetration is high. This is technically feasible but requires coordination that is not yet common practice.

Water Use in Water-Stressed Regions
The geographic dimension also applies to water. Data centers are often sited in arid regions because of low land costs and abundant solar energy potential. But these same regions are frequently water-stressed. In places like Arizona, New Mexico, and parts of Chile, data center water consumption has become a contentious issue. Local communities question whether the economic benefits of hosting these facilities outweigh the strain on already scarce water resources.
Some facilities use air-cooled or closed-loop systems that minimize direct water use, but these systems typically require more energy, creating a trade-off between water and carbon. In a world where both carbon budgets and freshwater supplies are under pressure, optimizing for one at the expense of the other is a delicate balancing act.
Rui Mendes sees this as a classic systems optimization problem with no easy answer. The variables are interconnected, the constraints are local, and the objectives are global. A solution that works in Scandinavia may fail in the American Southwest. A one-size-fits-all approach is unlikely to succeed.
Hardware Embodied Carbon: The Overlooked Giant
While operational energy and water use are beginning to receive attention, the embodied carbon and resource extraction associated with hardware manufacturing remain largely invisible. Producing a single GPU involves hundreds of steps across multiple countries, each with its own energy mix and environmental regulations. The supply chain is opaque, and manufacturers are not required to disclose lifecycle emissions.
Some estimates suggest that embodied carbon can equal or exceed operational carbon over the lifetime of a server, especially when the server is replaced every few years. If the industry continues to shorten hardware refresh cycles in pursuit of performance gains, the embodied carbon problem will only worsen. Extending the useful life of accelerators, designing for recyclability, and improving manufacturing transparency are all necessary steps—but they require coordination across an industry that currently has little incentive to act.
Rethinking Efficiency Metrics
Current efficiency metrics in machine learning focus almost exclusively on performance per compute unit: how many operations per second, how many parameters per watt. These metrics drive hardware design, model architecture, and even research agendas. But they ignore the full environmental picture.
Rui Mendes argues that we need new metrics that account for the entire lifecycle. What is the total carbon emitted per unit of useful work over a model’s lifetime, including training, deployment, and inference? What is the water footprint per query served? How much rare material was extracted to build the hardware, and what was the social cost in the communities where mining took place?
These questions are not easy to answer, but they are essential if the field is to take responsibility for its environmental impact. Without them, efficiency improvements can become a shell game: reducing operational energy while increasing embodied energy, or shifting water consumption from one watershed to another.
Inference: The Long Tail of Environmental Cost
Much of the discussion focuses on training, but inference—the process of using a trained model to make predictions—also carries a significant environmental burden. A model may be trained once but used millions or billions of times. For large-scale deployed models, the cumulative energy and water cost of inference can far exceed that of training.
Consider a model that answers millions of queries per day. Each query requires a forward pass through the network, consuming energy and generating heat. Over the model’s operational lifetime, the inference cost can dwarf the training cost. Yet inference efficiency receives far less attention than training efficiency in both research and public discourse.
Rui Mendes points out that this is a classic systems blind spot: we focus on the one-time cost of creation and ignore the ongoing cost of operation. A full lifecycle assessment must include both, and for many widely deployed models, the operational phase dominates.
What Can Be Done? A Systems Approach
Addressing the environmental cost of large-scale computing requires action on multiple fronts. No single intervention will solve the problem, but a combination of technical, organizational, and policy changes could bend the curve.
1. Transparent Reporting
Mandatory disclosure of energy use, carbon emissions, water consumption, and hardware lifecycle data would create accountability and enable better decision-making. Standardized reporting frameworks, like those proposed in recent research, could make environmental costs visible and comparable across projects.
2. Carbon-Aware Scheduling
Shifting training workloads to times and places with lower carbon intensity can reduce emissions without sacrificing performance. This requires coordination between cloud providers, grid operators, and research teams, but the technical barriers are surmountable.
3. Hardware Longevity and Circularity
Extending the useful life of accelerators, designing for repairability and recyclability, and creating markets for refurbished hardware can reduce embodied carbon and e-waste. This requires changes in procurement practices and manufacturer incentives.
4. Efficiency Research Beyond FLOPs
Research into model compression, sparse architectures, and efficient inference can reduce the environmental cost per query. But these efforts must be guided by full-scope metrics that capture the whole picture, not just computational efficiency.
5. Siting and Infrastructure Policy
Local governments can play a role by requiring environmental impact assessments for large data centers, including water use and grid effects. Siting decisions should consider not only economic factors but also the strain on local resources and communities.
Frequently Asked Questions
How much energy does training a single large model actually consume?
Estimates vary widely depending on model size, hardware efficiency, and grid carbon intensity. A 2019 study found that training a large natural language processing model emitted around 284 tonnes of CO₂ equivalent. Since then, models have grown significantly, and the largest training runs today likely emit several thousand tonnes. To put that in perspective, the average American car emits about 4.6 tonnes of CO₂ per year.
Why is water use a concern for data centers?
Data centers consume water both directly, through cooling systems that use evaporative methods, and indirectly, through the water used to generate the electricity they consume. In water-stressed regions, this can compete with agricultural, residential, and ecological needs. A single large training run can consume hundreds of thousands of liters of freshwater when both direct and indirect use are accounted for.
Does the environmental cost continue after training is complete?
Yes. Once a model is deployed, every query or prediction it makes consumes energy and water. For models that serve millions of users, the cumulative cost of inference can far exceed the cost of training. This ongoing operational impact is often overlooked in discussions that focus only on the training phase.
Can renewable energy solve the problem?
Renewable energy can significantly reduce the carbon footprint of training and inference, but it is not a complete solution. Water use, hardware lifecycle impacts, and land use for renewable infrastructure all remain concerns. Additionally, the availability of renewable energy varies by location and time, and new demand can sometimes displace other users onto fossil fuel sources if not carefully managed.
Conclusion: Seeing the Whole System
The environmental cost of training large models is not a simple problem with a single villain. It is an emergent property of a system that prioritizes performance above all else, that externalizes environmental costs, and that lacks the feedback loops necessary to self-correct. Rui Mendes believes that the first step toward a solution is simply to see the system clearly—to trace the connections from the mine to the data center to the cloud, and to recognize that every digital achievement rests on a physical foundation.
This is not an argument against progress. It is an argument for a more honest accounting of what progress costs, and for a more thoughtful approach to how we pursue it. The goal is not to stop building models, but to build them in a way that respects the finite resources of the planet we all share.
The next time a new record is set on a benchmark, Rui Mendes will still be curious. But his first question will not be about the score. It will be about the smoke, the water, and the earth that made it possible.