AI Power in 2026: 485 TWh, a 0.3 GW Stargate, a 116 GW Turbine Backlog, and 1,800 MW That Dropped Off the Grid in Seconds
Power is the one input to my platform I never see. I buy instances, the price includes the electricity, and the closest I get to a kilowatt-hour is the scale-to-zero work that keeps idle microVMs from burning it. But every conversation about where compute will come from in 2027 is now a conversation about substations and gas turbines, and the numbers in those conversations are unusually bad: projections quoted as measurements, announced gigawatts counted as built, and per-query figures with no boundary stated. I spent the week reading the primary sources, the IEA, Lawrence Berkeley, the grid operators, the NERC alerts, the company disclosures, and this is the honest ledger.
How much, measured
The International Energy Agency's 2026 update puts global data-centre consumption at 485 TWh in 2025, up 17 percent in a year, with the AI-focused share up around 50 percent; its central projection roughly doubles that to 950 TWh by 2030, about 3 percent of world electricity, with AI-specific load tripling. For the United States, Lawrence Berkeley National Laboratory's 2025 update, dated June, revises 2024 use to 192 TWh, 4.7 percent of US electricity, and projects 649 TWh or 11.8 percent in 2030 in its reference case, with a compounded range of 521 to 843 TWh. It assumes about 35 percent of AI-server power was training in 2024, falling to 20 percent by 2030, which is the single most important assumption in the report: inference, not training, is the load that grows. EPRI's three scenarios for 2030 are 380, 515 and 790 TWh. The EIA's 2026 outlook calls data centres the dominant driver of long-term load growth and counts servers as about 7 percent of commercial-sector electricity in 2025.
| Measure | Value | Source |
|---|---|---|
| Global data-centre electricity, 2025 | 485 TWh (+17%) | IEA, April 2026 |
| Global, 2030 central projection | 950 TWh, ~3% of world electricity | IEA |
| US data-centre electricity, 2024 | 192 TWh, 4.7% | LBNL, June 2026 |
| US, 2030 reference / range | 649 TWh, 11.8% / 521–843 TWh | LBNL |
| US interconnection needed by 2030 | 148 GW at 50% utilisation | LBNL |
| US installed capacity | 118 GW by 2030, 194 GW by 2035 | BloombergNEF, July 2026 |
| Global AI-dedicated capacity, theoretical | ~30 GW, Q4 2025, doubling every ~7 months | Epoch AI, January 2026 |
| Tracked AI sites, measured IT power | 86 sites, 13.3 GW | Epoch AI, September 2026 |
Note the difference between the last two rows. Epoch's 30 GW is chip shipments times rated power times an overhead multiplier, a theoretical maximum; its site tracker, which uses satellite imagery and permits, finds 13.3 GW of IT load at the 86 largest AI sites. Both are honest; only one is a meter.
What a prompt costs
Google's August 2025 measurement remains the only per-query figure from a frontier operator with a stated method: a median Gemini Apps text prompt at 0.24 Wh, 0.03 grams of CO2 and 0.26 millilitres of water, on a boundary that includes idle and failover machines, host CPU and RAM, and a PUE of 1.09. On the narrow accelerator-only boundary most papers use it would be 0.10 Wh. It excludes training, the network and your device, and the water is on-site only. Sam Altman's 0.34 Wh has no method attached; Epoch's independent estimate for GPT-4o was about 0.3 Wh, assuming a 100-billion-active-parameter model. Mistral's lifecycle study with ADEME puts a 400-token response from Mistral Large 2 at 1.14 gCO2e and 45 mL of water.
Then the counter-trend. Hugging Face's AI Energy Score v2 measured reasoning mode at about 30 times the energy of a plain answer on average: a 70B distilled model went from 49.5 Wh to 7,627 Wh per thousand queries. The "How Hungry is AI" paper found a 65-fold spread between short queries at 0.42 Wh and long reasoning prompts above 29 Wh. Google's 33-fold per-prompt reduction in a year and a 30-fold increase from turning reasoning on are the same order of magnitude, which is why efficiency gains have not bent the demand curve; LBNL's high-inference scenario adds 20 percent for exactly this. For training, Epoch's reconstruction of Grok 4 is about 310 GWh; the frontier training-run power draw has grown 2.1 times a year, from GPT-4's roughly 19 MW to over 100 MW, toward a gigawatt-class run around 2030.
Announced versus energised
The gigawatt announcements are where I had the most to correct. Epoch's April survey of Stargate found only Abilene energised, at about 0.3 GW of a planned 1.2, with five more sites under construction and most of the more than 9 GW planned targeting late 2028; several trackers describe Abilene as "1.2 GW operational", which is the design capacity. OpenAI's supplier commitments, 10 GW with NVIDIA, 6 with AMD, 10 with Broadcom, are contracts, not capacity. The figure compares the largest campuses' announced numbers with Epoch's measured IT power this month.
Some of the rest, with what is verifiable. xAI's Colossus 2 is estimated at 946 MW; the NAACP, Earthjustice and the Southern Environmental Law Center sued in April over 27 unpermitted gas turbines at the Southaven site, the Justice Department intervened in June, and there is no ruling. Meta's Hyperion in Louisiana grew from 2 GW to 5 GW, and in March Entergy and Meta announced ten gas plants totalling 7.5 GW at about $11 billion, Meta-funded, seven of them awaiting state approval. Amazon's Rainier campus, where Anthropic trains, had 18 buildings online in October 2025 toward 32 and a 2.25 GW grid draw; Anthropic's April agreement with Amazon runs to 5 GW with about 1 GW by year-end, on top of a $50 billion Fluidstack build announced in November. Microsoft's Fairwater sites in Atlanta and Wisconsin are described as an "AI superfactory" and Microsoft publishes no megawatt figures at all. The rule of thumb that a gigawatt costs $50 billion is Jensen Huang's and covers hardware only; all-in estimates with on-site generation run $45 to 60 billion, roughly 70 percent IT and 30 percent facility.
Where the electrons come from
Gas, mostly, and later than anyone wants. GE Vernova's second-quarter results report a gas-turbine backlog plus reservations of 116 GW, heading for 125 by year-end, with output at 20 GW a year rising to 30 by 2030 and new orders booking 2031 delivery slots. Heavy-duty turbines price around $790 to $950 a kilowatt, aeroderivatives around $1,800. BloombergNEF counts 124 GW of announced on-site gas for US data centres, only 56 percent of it with commissioning dates; the IEA expects 15 to 27 GW of on-site gas globally by 2030. Oracle contracted Bloom Energy fuel cells for up to 2.8 GW, 1.2 contracted, deploying through 2027.
Nuclear is real and slow. Constellation's Crane, the former Three Mile Island unit 1, is 835 MW for Microsoft with a 2027 restart: FERC waiver in June, transformers delivered in August, NRC approval pending. Meta's January package totals up to 6.6 GW by 2035, of which 2.1 GW is existing Vistra output and the rest is TerraPower and Oklo reactors from 2030 onward. On small modular reactors, the one hard milestone is the NRC's construction permit for TerraPower's Natrium in Wyoming on 4 March 2026, the first for a commercial non-light-water reactor, targeting 2030; Kairos holds two earlier permits, X-energy filed in December, NuScale has design approval only, Oklo broke ground without a licence to operate. No SMR delivers power to any data centre in 2026, and none will before 2030. The Amazon and Talen and X-energy figures you have seen, 1.92 GW and 5 GW, I could not verify from primary pages and have left out.
The grid's response
The queue numbers are the ones I would show anyone who thinks this is a chip-supply problem. ERCOT's large-load queue reached 233 GW at the start of 2026, up roughly 300 percent in a year, 77 percent of it data centres, against a system peak near 95 GW; its own vice-president said "we have outgrown the process". PJM, the eastern grid, forecasts a summer peak of 253 GW by 2046 from 160 today, and its 2027/28 capacity auction cleared at the $333.44 per megawatt-day cap, 6.6 GW short of its reliability requirement, with about 5.1 GW of the load increase attributed to data centres. On the generation side, LBNL's Queued Up 2026 counts 1,312 GW of generation and 749 GW of storage waiting to connect, gas up 86 percent, with a median wait of over five years.
The regulatory machinery moved this year. The Department of Energy directed FERC in October 2025 to take up large-load interconnection, and on 18 June FERC issued show-cause orders to all six wholesale markets with responses due in August and November. Texas's SB 6 sets a 75 MW threshold, disclosure of backup generation and parallel requests, and mandatory remote-curtailment equipment; the governor ordered a statewide audit and pause on connections in August. New York's executive order in July paused discretionary permits over 50 MW for a year; Seattle and Indianapolis have moratoria. On prices, PJM's market monitor attributed $9.3 billion of 2025/26 capacity cost to data centres; the E3 consultancy attributes about half the capacity-price jump to load growth and finds no historical shift onto residential bills; a claim of a 267 percent rise was rated mostly false because it was wholesale, not retail. Both things are true: the wholesale capacity price hit its cap because of this load, and your bill has not yet moved much because of it.
And the incidents, which are the operational story. NERC's 2026 State of Reliability records data-centre load-loss events of about 1,800 MW in February 2025 and 1,300 MW in June, plus several in the hundreds of megawatts and nine crypto-load drops over 100 MW in ERCOT. What happens is that a voltage dip trips the sites' protective systems and they transfer to backup in seconds, a gigawatt of load vanishing from a grid that then has to shed the surplus generation. NERC escalated from a Level 2 alert in September 2025 to a Level 3 "essential actions" alert on 4 May 2026, requiring dynamic models, annual stability studies and commissioning processes for large loads. Google's answer, and the one I find most interesting as an operator, is demand response for machine-learning workloads: it has contracted about a gigawatt of curtailable capacity across five utilities, pausing or moving training jobs when the utility asks. Batch jobs are a grid asset in a way an inference fleet is not.
The rack
The engineering at the rack is where the power problem lands on people like me. From OEM specifications compiled in August, a GB200 NVL72 draws 120 to 132 kW, a GB300 NVL72 132 to 142 kW nominal and about 155 peak, with 90 percent of the heat to liquid at 45 degrees inlet and 177 litres a minute of facility water; NVIDIA publishes no official rack figure. Vera Rubin NVL72, "in full production" for second-half 2026 shipment, is over 200 kW per rack, with a claimed ten times inference performance per watt over Blackwell on vendor-chosen low-precision benchmarks. The Kyber rack for Rubin Ultra in 2027 is 600 kW, and the 800-volt DC distribution that goes with it, cutting copper by 45 percent and losses by about 5 percent, is in production for 2027 with an OCP white paper still pending. Racks that draw more than a hundred houses need a different building, which is why Fairwater is two storeys with closed-loop cooling and "almost zero water".
The rack also explains the grid incidents. Synchronous training makes tens of thousands of GPUs go from idle to full power together and back, in milliseconds, and utilities see that as an oscillating load they cannot plan for. NVIDIA's GB300 power smoothing caps the ramp, puts energy storage in the power shelf, and deliberately burns GPU power on the ramp-down to flatten the AC draw; a Stanford paper this year, EasyRider, does the filtering with rack-level storage. Google's fleet PUE is 1.09 across 2024 and 2025; LBNL's US average for AI facilities is 1.145, trending to 1.136, and Google's water use rose by a third in 2025 to 10.9 billion gallons, a figure I have only from secondary coverage.
Measuring your own
The gap between the 0.24 Wh figure and the 29 Wh figure is the boundary, so measure with a stated one. On any NVIDIA host, nvidia-smi --query-gpu=power.draw sampled during a job and integrated gives GPU-only joules; multiply by PUE and a host factor for a facility estimate. The Zeus library, version 0.16.0 from July, does it through NVML with negligible overhead, and CodeCarbon adds CPU RAPL and a carbon-intensity lookup.
# tokens_per_joule.py: integrate nvidia-smi power while a job runs. GPU only; apply PUE yourself.
import subprocess, threading, time
def gpu_power_w():
out = subprocess.check_output(["nvidia-smi", "--query-gpu=power.draw",
"--format=csv,noheader,nounits"], text=True)
return sum(float(x) for x in out.split())
def measure(fn, interval=0.1):
done, result, joules = False, None, 0.0
def run():
nonlocal done, result; result = fn(); done = True
threading.Thread(target=run).start()
t_prev, p_prev = time.time(), gpu_power_w()
while not done:
time.sleep(interval)
t, p = time.time(), gpu_power_w()
joules += 0.5 * (p + p_prev) * (t - t_prev) # trapezoid rule
t_prev, p_prev = t, p
return result, joules
tokens, J = measure(lambda: my_generate(prompt)) # my_generate returns output token count
print(f"{J/3600:.3f} Wh {tokens/J:.1f} tokens/J {J/tokens*1000:.2f} mJ/token (GPU only)")
print(f"facility estimate: {J/3600*1.15*1.5:.3f} Wh at PUE 1.15 and 1.5x host overhead")
from zeus.monitor import ZeusMonitor
mon = ZeusMonitor(gpu_indices=[0])
mon.begin_window("gen"); my_generate(prompt); m = mon.end_window("gen")
print(m.time, "s", m.total_energy, "J")
Run it in reasoning mode and in plain mode on the same model and you will reproduce the thirty-fold number yourself, which is more persuasive than any report.
What I take from all of it. The measured load is real and growing at 17 percent a year; the projections that double it by 2030 assume inference keeps growing and reasoning stays expensive, and both look right. The gigawatts on the press releases are a third to a half built, and the power to run the rest is queued behind turbines that ship in 2031 and reactors that ship in 2030 at the earliest. The near-term constraint on AI compute is not chips. It is a substation, a turbine slot, a permit for the turbines you already installed, and a grid operator who has watched a gigawatt disappear in seconds and would like a model of your load before you connect. For a small operator the only power lever is not wasting it, which brings me back to the idle VMs I sleep at fifty seconds. It is a small contribution. It is also the only one I can measure.
Related: What Scale to Zero Cost Me to Build, LLM Inference in 2026 and Quantum Hardware in 2026.
I'm Ajay Kumar — I build and operate PandaStack, an open-source Firecracker microVM cloud for AI agents. Everything above comes from running it in production.
Need this kind of infrastructure work? See what I do or email hello@ajayk.sh.
Related
AI for Science in 2026: One Phase III Trial, Two Operational Weather Models, and Erdős Proofs Nobody Wants to Read
What has actually been validated outside the lab that made the claim: rentosertib's Phase III start, ECMWF and NOAA running AI forecasts operationally, Evo 2's viable phage genomes, the Erdős-problem scoreboard and its rediscovery problem, the GNoME and MatterGen materials that turned out not to be new, and the licence and pip-install status of every open model an engineer can run this week.
16 minSep 9, 2026On-Device Inference in 2026: 160 Tokens a Second on a Laptop, a 2-Bit Model in Every iPhone, and the Bandwidth Arithmetic That Decides What Stays Local
What actually runs on the device now: Apple's 3B model trained at 2 bits and the WWDC26 framework that lets any model plug in, Gemini Nano inside AICore, Windows ML and the quiet replacement of Phi Silica, the NPU TOPS numbers and why they are the wrong metric, the small-model field from Gemma 4 E2B to Qwen3.5-2B to LFM2.5 with licences, measured decode speeds from Google's own tables and an independent phone benchmark, the runtimes (llama.cpp with multi-token prediction, LiteRT-LM, ExecuTorch 1.0, MLX), the memory-bandwidth arithmetic that predicts all of it, and the model-extraction attacks nobody has fixed. With commands to run one today.
14 minSep 9, 2026Zero-Knowledge in 2026: Ethereum Blocks Proven in 3 Seconds for a Third of a Cent, a 52-Bit Security Gap, and Google Wallet Proving Your Age
Where zero-knowledge proofs actually stand: the Ethereum Foundation's real-time proving definition and the measured numbers on ethproofs (from 16 minutes and $1.69 to 3 seconds and $0.003 a block on consumer GPUs), why the security certification is the hard part (a 52-bit gap between the proven and claimed soundness of the hash used), the zkVM bugs fuzzers found this year, the L2 landscape at $43 billion with zero Stage 2 rollups, and the non-crypto uses shipping now: Google Wallet's ZK age proofs, the EU wallet's December deadline, zkML and zkTLS. Plus a proof you can generate on a laptop.
14 min