Why We Should Look at the Power Cord, Not Just the Output
When a new machine learning system nails a breakthrough—say, it suddenly gets uncannily good at recognizing faces or translating obscure languages—the headlines gush about the cleverness. But behind every one of those wins is a physical reality that rarely gets a headline: racks of servers humming for weeks or months, cooling systems gulping water and pushing heat into the air, and a steady, invisible drain on the electrical grid. Rui Mendes has spent years thinking about the systems that make digital life tick, and lately he’s been drawn to the material footprint of these computational feats. The question isn’t whether the results are impressive. It’s whether we’ve even started to tally what they actually cost.
Training a big model isn’t a one-click affair. It’s an industrial process, often spread across thousands of specialized processors running in parallel. The energy draw can rival that of a small town, and the water used to cool the data centers can strain local supplies. Yet when a new model is announced, these numbers are almost never part of the story. Rui thinks that if we’re serious about building responsible technology, we need to start measuring what matters—and that includes the environmental load that comes with every training run.
The Scale of a Single Training Run
To get a feel for the problem, let’s look at a concrete example. Take a large natural language model trained on a web-scraped corpus. The process might involve hundreds of graphics processors running flat-out for several months. Each of those chips draws power, and the facility housing them needs constant cooling to keep things from melting down. The electricity consumption for a single training run can easily reach hundreds of megawatt-hours—roughly the annual energy use of dozens of households.
But electricity is only part of the story. The hardware itself carries a carbon and resource cost. Specialized chips require rare minerals, and their manufacturing is energy-intensive. If we’re going to talk about environmental impact, we should include the embedded emissions of the equipment, not just the operational energy. Rui often points out that this is where systems thinking becomes non-negotiable: a narrow focus on one metric, like kilowatt-hours during training, can blind us to the bigger picture.
Water Use in the Data Center
Another factor that’s easy to overlook is water. Many data centers rely on evaporative cooling, which consumes huge amounts of fresh water. In regions already dealing with water stress, this can pit the needs of a data center against those of the surrounding community. A single large training run can evaporate millions of liters of water—water that might otherwise go to crops or drinking supplies. Rui finds this especially jarring because the digital world is so often imagined as weightless and clean, when in fact it’s thirsty and tethered to very material resources.
Location matters enormously. A data center plugged into a coal-heavy grid will have a much higher carbon footprint than one running on renewables. Similarly, water use in an arid region hits differently than in a water-rich area. But transparency about these factors is rare, which makes it tough for researchers—or anyone else—to assess the true environmental load of a given model.
Why Efficiency Gains Can Fool Us
There’s a common counterargument: as hardware gets more efficient, the energy cost per computation drops. That’s true, but it doesn’t necessarily shrink overall energy consumption. In practice, efficiency gains often lead to bigger models and more ambitious training runs—a phenomenon sometimes called the rebound effect. When it becomes cheaper to train a model, the incentive is to train an even larger one, chasing marginal performance improvements that demand disproportionately more resources.
Rui sees this as a classic systems trap. Without a deliberate effort to cap resource use or account for environmental costs, efficiency improvements can actually accelerate consumption. The real question isn’t how efficient a single operation is, but how much total energy and water a project consumes from start to finish—and whether that total is justified by the value it creates.
Comparing Training to Other Energy-Intensive Activities
To put the numbers in perspective, training a single large model can emit as much carbon dioxide as several cars do over their entire lifetimes. Some estimates peg the carbon footprint of a state-of-the-art training run at over 280 tonnes of CO₂ equivalent. That’s roughly the lifetime emissions of five average cars, including their manufacturing. And since many models are trained multiple times during experimentation, the cumulative impact grows fast.
Rui finds it useful to compare these figures to activities that get far more environmental scrutiny. Air travel, for instance, is constantly discussed in terms of its carbon footprint, and many people actively try to fly less. Yet the computational equivalent—training a massive model—rarely enters the same conversation, even though its emissions can be comparable to a transatlantic flight for each researcher involved. The gap in awareness is something Rui thinks we need to close.
The Lifecycle Perspective: Beyond Training
Training is only one phase of a model’s life. Once deployed, a model keeps drawing energy every time it’s used. For a popular service handling millions of requests a day, the inference phase can dwarf the training phase in total energy consumption. This is especially true for models that need powerful hardware to run, like those used in real-time video analysis or large-scale recommendation systems.
Rui emphasizes that a full lifecycle assessment is necessary to understand the true environmental cost. That means accounting for the energy and materials to manufacture the hardware, the energy for training, the energy for inference, and the energy for eventual decommissioning and recycling. Without this wider view, we risk optimizing one phase while ignoring bigger impacts elsewhere.
What Could Transparency Look Like?
One of the biggest roadblocks to tackling this issue is the lack of transparency. Most organizations don’t disclose the energy consumption, carbon emissions, or water use tied to their models. Rui argues this has to change. Just as some companies now publish environmental impact reports for their physical products, developers of large models should provide similar disclosures. That could include the total electricity used during training, the carbon intensity of the grid at the data center location, and the water consumed for cooling.
Such transparency would let researchers, policymakers, and the public make informed comparisons. It would also create an incentive for developers to minimize environmental impact, rather than simply chasing maximum model performance. Rui believes that measurement is the first step toward management, and without open reporting, the environmental costs will stay invisible.
Rethinking the Metrics of Success
The current culture of machine learning research rewards ever-larger models and incremental improvements on benchmark datasets. But those benchmarks rarely include environmental cost as a factor. Rui suggests the community should consider new metrics that balance performance against resource consumption. For example, a model that hits 95% accuracy with half the energy use of a model that reaches 96% might be the smarter choice from a systems perspective.
This shift would require changes in how research is funded, published, and celebrated. Conferences and journals could require submissions to include estimates of computational resources used. Funding agencies could prioritize projects that demonstrate efficiency alongside effectiveness. Rui sees this as a chance to align the values of the research community with broader societal goals.
Practical Steps for Developers and Organizations
While large-scale change needs systemic shifts, there are practical steps individual developers and organizations can take. Choosing data center regions powered by renewable energy is one straightforward option. Using more efficient hardware and optimizing training algorithms to require fewer computational steps can also make a difference. Rui points out that simply being mindful of the issue and measuring energy use is a meaningful first step.
Another approach is to ask whether a large model is truly necessary for the task at hand. In some cases, a smaller, more specialized model can achieve comparable results with a fraction of the resources. Transfer learning, where a pre-trained model is fine-tuned for a specific task, can also reduce the need to train from scratch. Rui encourages developers to ask: is the marginal improvement in performance worth the additional environmental cost?
The Role of Policy and Regulation
Individual and organizational efforts matter, but Rui believes policy interventions will eventually be necessary to drive widespread change. Possible measures include requiring environmental impact disclosures for large training runs, setting efficiency standards for data centers, or incorporating environmental costs into the pricing of cloud computing services. Such policies could help internalize the externalities that are currently ignored.
Regulation could also spur innovation in energy-efficient hardware and cooling technologies. If the environmental costs of computation were reflected in its price, there would be a stronger market incentive to develop greener alternatives. Rui sees this as a natural extension of existing environmental regulations that apply to other industries.
Why This Matters for Everyone
The environmental impact of large-scale computation isn’t just a concern for researchers or tech companies. As these models become embedded in everyday services—from search engines to healthcare diagnostics—their collective footprint grows. The energy and water used to train and run these models are drawn from shared resources. Rui emphasizes that this is a collective challenge that requires collective awareness and action.
For users, being informed about the hidden costs of digital services can lead to more conscious choices. For developers, it can inspire more efficient design. For policymakers, it can highlight the need for standards and regulations. Rui’s systems-minded perspective reminds us that everything is connected: the cloud is not weightless, and every query has a cost.
Frequently Asked Questions
How much energy does training a large model actually use?
The energy consumption varies widely depending on the model size, hardware efficiency, and duration of training. Some large models have been reported to consume hundreds of megawatt-hours of electricity during training—enough to power dozens of homes for a year. However, exact figures are often not publicly disclosed, making it difficult to generalize.
Does the environmental impact stop once the model is trained?
No. Deployed models continue to consume energy every time they are used for inference. For widely used services, the cumulative energy from inference can exceed the training energy over time. Additionally, the hardware lifecycle—from manufacturing to disposal—adds further environmental costs.
What can be done to reduce the environmental footprint of large models?
Several strategies can help: using data centers powered by renewable energy, improving hardware efficiency, optimizing algorithms to require less computation, and being transparent about resource consumption. On a systemic level, policies that require disclosure of environmental impact and incentivize efficiency could drive broader change.


