When we picture the internet’s environmental footprint, we usually imagine vast server farms humming under fluorescent lights, the scramble for rare minerals, or the electricity siphoned by billions of screens. But there’s a quieter, more concentrated cost that rarely makes headlines: the brute energy needed to train a single, massive neural network from scratch. Rui Mendes, a systems-minded observer drawn to the hidden inputs behind modern computation, has been following this thread. What he’s uncovered is a story of runaway scale, uncomfortable trade-offs, and a surprising tether to the physical infrastructure we often take for granted.

The Scale of a Single Training Run
To grasp the numbers, you have to look beyond the familiar image of a desktop computer. Training a state-of-the-art model isn’t a quick job you kick off before lunch. It’s a marathon that can stretch across weeks or months, running nonstop on clusters of thousands of specialized chips—GPUs or TPUs—each gulping 300 to 400 watts under full load. Multiply that by a thousand chips, then by 24 hours a day for a month, and the figures stop being abstract. They start resembling the energy footprint of a small town.
Rui notes that the chips themselves are only part of the story. The supporting cast—memory banks, storage arrays, networking gear, and the cooling systems that keep everything from melting—adds a hefty surcharge. In a typical data center, the overhead can be 20% to 60% on top of the computing energy, a ratio captured by the PUE metric. When a single training run pulls megawatts continuously for weeks, the total electricity consumed can match what dozens of households use in a year.
Carbon Emissions That Rival Physical Industries
Researchers have started putting numbers to the carbon bill. One well-cited study found that training a large transformer model can belch out over 284 tonnes of CO2 equivalent—roughly five times the lifetime emissions of an average car, including its manufacture. Rui finds this comparison jarring because it connects the ethereal world of code to the smokestacks and tailpipes we usually associate with pollution. Software feels weightless, but its birth can leave a heavy atmospheric mark.
Where the data center sits on the map matters enormously. A training run hosted in hydro-rich Quebec or wind-swept Denmark will have a fraction of the carbon cost of the same run executed in a region still hooked on coal. Yet the decision often comes down to electricity price, not carbon intensity. That mismatch—between economic logic and environmental sense—is a tension Rui keeps circling back to.

Water in the Machine: The Overlooked Cooling Cost
Electricity grabs the headlines, but Rui’s systems lens pulls another resource into focus: water. Data centers often rely on evaporative cooling towers that literally vaporize water to dump heat, while others draw from power plants that themselves consume huge volumes for cooling. In drought-prone regions, this creates a quiet competition between server racks and local farms or households.
A single large training run can be responsible for millions of liters of water. In a warm climate, a facility might evaporate more than a liter for every kilowatt-hour of energy used. Stretch that over a megawatt-scale run lasting weeks, and the cumulative water footprint becomes hard to ignore. Rui sees this as a classic systems blind spot: optimize aggressively for compute speed, and you end up externalizing costs onto a local aquifer that nobody thought to measure.
The Hardware Lifecycle: From Mines to E-Waste
The environmental story doesn’t begin when the servers power on. It starts in mines where cobalt, lithium, tantalum, and rare earth elements are pulled from the earth—often leaving toxic tailings and strained communities in their wake. Fabricating GPUs and TPUs involves energy-hungry plants, harsh chemicals, and globe-spanning supply chains. Rui insists that any honest accounting has to include this embodied energy and material toll, not just the electricity meter during training.
Then there’s the tail end. Hardware cycles are brutal; specialized accelerators can slide into obsolescence within a few years. A cluster built for a single training run might be decommissioned or shuffled to less demanding tasks, but the churn generates mountains of electronic waste. Some materials get recycled, but the process is leaky, and plenty ends up in landfills or informal scrapyards where the health and environmental costs are borne by people far from the data center’s air-conditioned halls.
Why Efficiency Gains Might Not Save Us
It’s comforting to think that smarter engineering will bail us out. Each new chip generation does more calculations per watt. But Rui points to an old pattern that systems thinkers know well: Jevons paradox. When something gets more efficient, the cost per unit of work drops, and total demand often swells instead of shrinking. We don’t use the same energy to do more; we do more and end up using the same—or even more—total energy.
You can see it playing out in real time. As hardware gets beefier, models get hungrier, datasets balloon, and training runs multiply. What was once a week-long job on a handful of GPUs has morphed into a months-long slog on thousands of accelerators. The energy per run hasn’t fallen; it’s shot upward. Efficiency gains are being devoured by scale, not converted into a lighter footprint.

Geographic and Temporal Shifting: A Flexibility Worth Exploring
One of the more hopeful threads Rui has tugged on is the idea of moving training workloads through space and time to soften their impact. Unlike a factory bolted to the ground, a training run can, in principle, be pointed at a data center with a cleaner grid. Some teams are already dabbling in this—scheduling heavy compute for when the wind is blowing hard or the sun is high, even shifting workloads between regions to chase renewable peaks.
Temporal shifting is another lever. Training doesn’t have to be a nonstop sprint; it can be paused and resumed. By syncing the heavy lifting with periods of excess renewable supply—when the grid’s carbon intensity dips—the emissions footprint can drop sharply. This takes clever orchestration and a willingness to accept a somewhat longer wall-clock time, but the environmental payoff could be large. Rui sees it as a tidy example of systems tradecraft: swap a little speed for a lot less collateral damage.
The Transparency Gap
One of the biggest roadblocks is simply not knowing. Unlike cars or refrigerators, where energy labels are mandatory, there’s no requirement to disclose the energy or carbon cost of training a specific model. Some research papers now voluntarily include estimates, but they’re spotty and use inconsistent methods. Rui notes that without standardized reporting, it’s nearly impossible for the public, policymakers, or even fellow practitioners to compare options or push for better practices.
The opacity stretches to water use and hardware lifecycle impacts, which are almost never mentioned. A systems thinker would argue that you can’t manage what you don’t measure. The first move toward shrinking the environmental cost of training is simply making that cost visible—turning an externality into a tracked metric that sits alongside accuracy and speed on the dashboard.
Rethinking the Value Proposition
At bottom, it’s a question of worth. Training a gargantuan model burns through enormous resources, but what does society get back? Some models crack open problems in medicine, climate science, or fundamental physics—areas where the environmental price tag might be justified by the potential for outsized good. Others are trained for tasks that could likely be handled with simpler, less thirsty methods. Rui suggests we need a more discerning eye: not every problem demands a model with hundreds of billions of parameters, and “bigger is better” shouldn’t be the default reflex.
This calls for a mindset shift from maximizing scale to optimizing for a given outcome with minimal resource draw. It’s a classic engineering trade-off, but one that gets steamrolled when incentives reward headline-grabbing performance numbers over resource-aware efficiency. Rui’s curiosity leads him to wonder: what if we celebrated the most resource-efficient solutions instead of just the largest or most accurate? That reframing could nudge the whole field toward saner habits.
Frequently Asked Questions
How much energy does training a large model actually use?
Estimates vary widely depending on model size, hardware, and duration, but some of the largest training runs have consumed over 1,000 megawatt-hours of electricity—enough to power an average U.S. household for more than 100 years. The associated carbon emissions can reach hundreds of tonnes of CO2 equivalent, though this depends heavily on the energy mix of the local grid.
Why is water usage a concern for data centers?
Data centers often use water for evaporative cooling or to dissipate heat through cooling towers. In water-stressed regions, this consumption can compete with local agricultural, industrial, or residential needs. Additionally, the electricity powering the data center may come from thermoelectric plants that themselves require large amounts of water for cooling, creating an indirect water footprint.
Can training be made more sustainable without sacrificing performance?
Yes, several strategies can help. Scheduling training to coincide with times of high renewable energy availability, using data centers in regions with cleaner grids, and optimizing model architectures to require less computation for the same task are all viable approaches. However, these require conscious effort and often a willingness to accept slightly longer training times or invest in more efficient hardware.
What can individuals or smaller organizations do to reduce impact?
For those not training massive models from scratch, the biggest difference comes from thoughtful use of existing pre-trained models. Fine-tuning a large model on a specific task uses orders of magnitude less energy than training from scratch. Additionally, choosing to run inference on efficient hardware and being mindful of unnecessary computation can cumulatively make a difference.
Looking Ahead: A Call for Systems Awareness
Rui Mendes doesn’t frame this as a doom scroll. He sees it as an invitation to widen the lens. The environmental cost of training large models emerges from a tangle of interconnected systems: energy grids, hardware supply chains, water infrastructure, economic incentives, and research culture. Tweaking any one piece in isolation won’t fix the whole, but a coordinated push toward transparency, efficiency, and value-based choices could bend the curve.
As training runs keep swelling, the conversation needs to escape niche academic circles and enter the broader public square. The benefits of these models are widely shared, but so are the environmental costs—whether they show up on our electricity bills or not. Rui’s hope is that by dragging these hidden costs into the light, we can start making more intentional choices about what we build, how we build it, and whether the trade-off is genuinely worth it.