Every time a new, jaw-dropping model gets announced—one that writes poetry, generates photorealistic images, or translates languages on the fly—the headlines focus on what it can do. What they almost never mention is what it took to build it. Not the cleverness of the code or the elegance of the architecture, but the physical stuff: the electricity, the water, the metals, the heat. I started pulling on this thread a while back, and the more I learned, the more it felt like we’re all driving a fleet of invisible, smoke-belching trucks through the digital world without ever seeing the exhaust.
This isn’t a story about whether the technology is good or bad. It’s a story about systems—how they connect, what they consume, and what we choose to ignore when the machinery is out of sight.
What a Single Training Run Actually Demands
Let’s get concrete. Training a large model isn’t like leaving your laptop on overnight to render a video. It’s a coordinated assault on a problem, using thousands of specialized chips—GPUs or TPUs—running flat-out for weeks or months. Each chip might draw 300-400 watts. Multiply that by 10,000 chips, then by 24 hours, then by 90 days. The numbers get absurd fast.
One of the few public estimates, from a team at the University of Massachusetts Amherst, pegged the carbon emissions of a single large training run at over 284 tonnes of CO₂ equivalent. That’s roughly the lifetime output of five average cars, including their manufacturing. And that’s just the electricity for the run itselfânot the trial runs, the failed experiments, the hyperparameter tuning that came before. The final model is the tip of an iceberg, and the hidden mass below the waterline is made of burned coal and natural gas.

It’s Not Just About Electricity
When people talk about making data centers “green,” they usually mean buying renewable energy. That matters, but it’s only one slice of the pie. The grid mix—how much of the local power actually comes from wind, solar, hydro, or nuclear versus coal and gas—varies wildly by region. A training run in a data center hooked to a coal-heavy grid leaves a much deeper carbon scar than one powered by a hydroelectric dam. Same code, same model, wildly different physical consequences.
Then there’s water. Data centers run hot, and keeping thousands of chips from melting themselves requires serious cooling. In many facilities, that means evaporative cooling towers that drink millions of gallons of freshwater. A 2023 study out of UC Riverside put the number at up to 700,000 liters for a single large training run. In places like Arizona or Spain, that water is pulled straight out of watersheds already strained by drought and agriculture. It’s a quiet competition for a resource most people don’t associate with software.
The Ghost Emissions in the Hardware
Before a server ever powers on, it’s already carrying a carbon debt. Manufacturing GPUs and other accelerators involves mining rare earth elements, smelting, chemical baths, and fabrication in energy-hungry foundries. The supply chain sprawls across continents. A single server might embody hundreds of kilograms of CO₂ before it’s even plugged in. And in this industry, hardware doesn’t age gracefully. The relentless pace of improvement means racks get swapped out every three to five years, not because they’re broken, but because they’re no longer competitive. That churn produces mountains of e-waste—circuit boards, batteries, and cables that are notoriously difficult to recycle safely. We upgrade our models and quietly landfill the old ones.

The Efficiency Trap
There’s a comforting story that goes: “Chips are getting more efficient, so the problem will solve itself.” It’s partly true. A modern GPU delivers far more calculations per watt than one from five years ago. But efficiency has a funny way of backfiring. When something gets cheaper or faster, we don’t use less of it—we use more. Economists call it the rebound effect, and it’s alive and well in the world of large-scale computing.
Every efficiency gain gets swallowed by ambition. The next model isn’t just a little bigger; it’s ten times bigger. The compute budget expands to absorb whatever headroom the engineers created. Total energy use doesn’t drop; it climbs. We’re optimizing the parts while the whole system grows hungrier. It’s a classic systems trap—one that feels familiar to anyone who’s watched traffic expand to fill a new highway lane.
Where the Data Centers Land—and Who Pays
Data centers don’t sprout randomly. They cluster where land is cheap, tax breaks are generous, and power is abundant. Often, “abundant power” means fossil fuels. Virginia’s data center alley, for instance, sits on a grid that’s only about 5% renewable. Compare that to Quebec, where hydro dominates, and the same workload can have a radically different carbon profile. The location decision—made by a cloud provider or a corporate IT team—is one of the biggest environmental levers, and it’s almost never visible to the people using the service.
Then there’s the local fallout. A data center in a desert doesn’t just use water; it takes water that might have gone to a farm or a household. Backup diesel generators cough out particulates that settle in nearby lungs. The burdens aren’t spread evenly across a spreadsheet; they land on specific communities, often ones with the least power to push back.
The Transparency Problem
Here’s where it gets genuinely frustrating. Most organizations that train large models don’t publish their energy numbers, their water consumption, or their grid mix. When they do share figures, the accounting can be creative. Buying renewable energy certificates (RECs) lets a company claim “100% renewable” even if the actual electrons powering the servers came from a gas plant. The certificates might fund a wind farm somewhere, but they don’t change the physical reality at the data center’s plug.
There are grassroots tools—CodeCarbon, ML CO2 Impact—that try to estimate emissions based on hardware type, runtime, and location. They’re useful, but they’re guesses. Without mandatory reporting, the real numbers stay locked inside corporate spreadsheets. Sunlight is scarce.

Paths Toward a Lighter Touch
If we’re going to keep building these systems—and it seems we are—the question shifts from “should we stop?” to “how do we do it with less damage?” The answers aren’t purely technical. They’re about choices, incentives, and who gets a seat at the table.
Time-shifting the workload. Training doesn’t have to happen right now. It can be scheduled for when the grid is cleanest—midday when solar is peaking, or windy nights when turbines are spinning. Some data centers already shift compute loads in real time to chase renewable availability. It requires flexible scheduling and a willingness to accept slightly longer training times, but the carbon savings can be substantial.
Keeping hardware alive longer. Extending server life through modular upgrades rather than wholesale replacement cuts embodied emissions. A few large operators are experimenting with designs where individual components—CPUs, accelerators, memory—can be swapped independently. It’s less glamorous than buying the latest rack, but it keeps materials out of the landfill longer.
Smaller, smarter models. A quiet revolution is happening in model efficiency. Techniques like distillation, pruning, and quantization can shrink a model to a fraction of its original size while keeping most of its capability. A distilled model might lose a point or two of accuracy but train in a tenth of the time and run on a fraction of the hardware. For many real-world tasks, that’s a trade worth making. It also opens the door for groups without massive compute budgets to participate.
Mandatory disclosure. The simplest, most powerful lever might be this: require any training run above a certain compute threshold to report its energy consumption, grid mix, and water usage. Make the numbers public. Once the environmental cost is visible, it becomes a factor in decision-making. Researchers might choose cleaner regions. Companies might compete on efficiency, not just capability. Transparency doesn’t fix everything, but it makes the trade-offs impossible to ignore.
Zooming Out: The Bigger Pattern
The environmental weight of training large models isn’t an isolated problem. It’s one expression of a broader habit: we build complex digital systems without accounting for their full physical lifecycle. You see the same pattern in cryptocurrency mining, in the explosion of IoT devices, in the endless expansion of streaming infrastructure. Every new service adds another layer of concrete, steel, copper, and lithium. The cloud isn’t weightless. It’s heavy, and it sits on real ground.
This isn’t a call to halt progress. It’s a call to make progress conscious. We should be asking not just “Can we build this?” but “Should we build it this way, at this scale, in this place?” Some models will be worth their footprint—the ones that accelerate drug discovery, improve climate forecasts, or design better materials. Others are trained for marginal gains on benchmarks that don’t translate to anything useful. The difference matters, and it should be deliberate, not accidental.
FAQ
Why is training a large model so much more energy-intensive than regular computing?
Training means pushing enormous datasets through networks with billions of parameters, over and over, for weeks or months, across thousands of power-hungry processors. Each chip draws hundreds of watts, and the cluster adds up to a small power plant’s worth of demand. On top of that, the cooling systems needed to keep everything from overheating consume significant energy and water on their own.
Can switching to renewable energy fix the whole problem?
Renewables slash the carbon emissions from electricity use, and that’s a big deal. But they don’t erase the embodied carbon from manufacturing the hardware, nor do they solve the water consumption issue. Clean power is a necessary piece of the puzzle, not the whole picture.
What can an ordinary person do about something so far removed from daily life?
Individual choices won’t directly change how large organizations train their models, but collective pressure can. Supporting transparency efforts, asking cloud providers about their energy sources, and pushing for policies that require emissions disclosure all add up. On a personal level, favoring services that publicly commit to environmental accountability sends a market signal, however small.
Are smaller, more efficient models really as good as the massive ones?
In many practical situations, yes. Distillation, pruning, and quantization can produce models that keep most of the original’s capability while being dramatically cheaper to train and run. You might trade a tiny bit of accuracy for a huge cut in energy use. For a lot of applications, the lean version is more than enough—and sometimes the only version that makes sense.
The environmental cost of training large models isn’t a reason to walk away from the field. It’s a reason to bring the same rigor and systems thinking we apply to the models themselves to the infrastructure that births them. Measure it. Disclose it. Optimize not just for accuracy, but for the full footprint. The smartest systems ought to rest on a foundation of responsibility—to each other, and to the planet that hosts all this computation.