AI Supply Chain Research · Sector 07

Datacenter, Power & Cooling

AI's real bottleneck is no longer chips but electrons, and the datacenter is being redesigned from the grid substation down to the sub-1V rail to squeeze out every token per watt.  ·  ← back to the series  ·  Background: Primer §00, §03, §07, §08, §10

At a glance
Winners
Bloom EnergyGE Vernova / Siemens Energy / Mitsubishi PowerDelta ElectronicsInfineon / Wolfspeed (power semis)SST startups: DG Matrix, Amperesand, Heron Power, Novos PowerTesla (Megapack BESS)Liquid-cooling ecosystem: Vertiv, Boyd, nVent, DanfossNeoclouds / EaaS: Crusoe, VoltaGrid
Bottlenecks
Firm dispatchable power — <10GW/yr of gas added in 2026-27, ~15GW/yr net-new ELCC vs a ~terawatt queue; grid headroom negative by 2027.Heavy-duty turbine blades — single-crystal nickel alloys (rhenium, cobalt, yttrium under Chinese export control) from just 4 Western firms; 300-500-ton cores need 24-30 months.Grain-oriented electrical steel (GOES) & large power transformers — longest lead times in the stack (3-4 years); ~80% of US high-power transformers are imported.3,300V+ SiC MOSFETs — gate the MV-input SST; still in limited production, forcing LV-input SST as the earlier-to-market variant.Quick Disconnects, CDUs & liquid-ready datacenter shells — Nvidia's ramp shorts simple parts; too few DLC-ready halls force inefficient RDHx 'bridge' solutions.Grid interconnection & transmission — ~5-7 year queues; conversion (not queue depth) now binds in PJM (24GW of executed agreements terminated since 2020).
Risks
Peak gas turbine orders — SemiAnalysis flags risk that current turbine order surge is temporary; secondary-market turbine availability is surging, pressuring OEM pricing.Datacenter delays overstated but real — SemiAnalysis debunks '50% canceled' yet flags specific slips (STACK/Oracle to 2029 on gas pipeline); demand is fine, execution slips.Grid stability / regulatory clampdown — NERC issued a Level 3 alert (May 2026) on large loads; ERCOT ride-through mandates (NOGRR282) and a Computational Load Entity registration raise the compliance bar.800VDC safety/code gap — full NEC support only ~2029; DC arc-flash has no IEEE 1584/NFPA 70E PPE table; grounding fragmentation makes each site a custom engineering project.AI-bubble / demand risk — 1GW of AI capacity implies ~$12-13B annual revenue; if token demand disappoints, the whole power/cooling capex thesis unwinds. OEMs remember 30 years of boom-bust.Silicon ceiling caps everything — AI consumes ~60% of TSMC N3 in 2026, ~86% in 2027; power/DC headroom can't be used without chips.
Catalysts
800VDC-native compute (Kyber/Vera Rubin Ultra) shipping late-2026/2027 — the inflection where 800VDC becomes mandatory, not optional, spiking penetration.NEC 2029 code + OCP/UL standards (year-end 2026 white paper, UL 857 Ed.15 to 1500VDC) unlocking colo/smaller-builder 800VDC adoption.SST UL certification — DG Matrix targets end-Q2 2026; first certification unlocks the ~$32B facility-level TAM (~2029 scale adoption).ERCOT Batch Zero / co-location rules (NPRR1325, effective Jul 2026) + FERC's large-load rulemaking (RM26-4) codifying hybrid grid+BTM structures.Behind-the-meter TAM inflection — DC BTM equipment crossing 50GW/year by 2029; new-entrant order flow (Boom Superpower, Doosan DGT6, ProEnergy retrofits).Load-flexibility / demand response reform — 20-90hrs/yr curtailment could unlock 6.5-14.7GW in ERCOT; PJM's Expedited Interconnection Track (June 2026).Microfluidic / direct-to-silicon cooling — the step beyond cold plates as chip TDP heads past 1,500W (TSMC CoWoS-R micropillar 5.3kW, Microsoft GH200 ~50% lower thermal resistance at ECTC 2026).

Overview

SemiAnalysis frames this sector through one lens: cost per token is set by how much compute you can pack densely, power, and cool. That forces three simultaneous revolutions. (1) Power supply: the US grid is 'sold out' — roughly a terawatt of load requests are queued, ELCC headroom goes negative by 2027, and interconnection now runs ~5-7 years — so behind-the-meter (BTM) onsite gas, fuel cells and engines are set to power over half of new US datacenter capacity by 2028. (2) Electrical distribution: rack density is climbing toward 600kW+ (Kyber) and eventually ~2,300W chips, breaking 48-54V distribution and forcing an 800VDC transition that migrates conversion content from grey space to white space and enables a Solid-State Transformer (SST) endgame. (3) Cooling: >1000W chips end the air-cooled era, pushing direct-to-chip liquid cooling (DLC), CDUs and quick-disconnects, while hyperscalers already run 1.08-1.15 PUE via free cooling. Layered on top are grid-stability risk from gigawatt synchronous AI loads, the political fight over household electricity bills, water-footprint myths, and the speculative frontier of space datacenters. SemiAnalysis's edge is a bottom-up, building-by-building Datacenter/Industrials/Energy model tracking 6,000+ datacenters and 40,000 generators via satellite imagery, letting it call winners and 'fake losers' before the market.

Positioning: Who Wins and Why

Bloom Energy — First called out as the biggest BTM beneficiary in Dec 2024; SOFC fuel cells (no combustion, near-instant EPA permitting, deployable in weeks) are taking share fast. Claims 2GW/yr capacity by end-2026. Premium ($3,000-4,000/kW) offset by speed and siting flexibility (usable in population centers).

GE Vernova / Siemens Energy / Mitsubishi Power — The turbine 'Big Three' — order-booked into 2028-29 with nonrefundable reservations. But SemiAnalysis flags risk of 'peak' orders and secondary-market turbine supply surging, so the durable winners may be fuel cells and RICE, not just the majors.

Delta Electronics — Leading the 800VDC transition — power racks, PSU shelves, 110kW power shelves with 80kW BBU, the first 2.4MW 800VDC In-Row CDU (GTC 2026), and 800VDC busway. Straddles power electronics AND liquid cooling, the two content-growth engines.

Infineon / Wolfspeed (power semis) — The silicon under 800VDC/SST. Infineon claims ~40x/14x SST weight/size reduction and a 12kW BBU roadmap (99.5% efficiency); Wolfspeed's 10kV SiC MOSFET (bare die since Mar 2026) opens direct MV rectification. ±400V bipolar rides the mature EV 650V-GaN supply chain.

SST startups: DG Matrix, Amperesand, Heron Power, Novos Power — >$320M funded in the year to Mar 2026. DG Matrix (ABB-backed, Infineon SiC) is the only SST in Nvidia's MGX reference; Heron building a 40GW US plant; Amperesand targets 30MW deployed in 2026. Eaton acquired Resilient Power for SST expertise.

Tesla (Megapack BESS) — The default answer to gigawatt load fluctuations, LVRT ride-through and demand response — pairs with turbines at xAI. GW-scale ~$1B TAM per site, plus the emerging role as a firm-capacity and demand-response asset class.

Liquid-cooling ecosystem: Vertiv, Boyd, nVent, Danfoss — DLC ramp shorts CDUs, cold plates, manifolds and Quick Disconnects. Danfoss Turbocor compressors already run on 700-813V DC — a rare cooling component ready for the DC-native future.

Neoclouds / EaaS: Crusoe, VoltaGrid — VoltaGrid runs full Energy-as-a-Service (e.g. 2.3GW for a 1.4GW Vantage campus, 64% overbuild) bundling RICE + synchronous condensers + BESS. Crusoe books bridge power (Abilene) and 1.2GW of Boom Superpower turbines.

Key Data

MetricValueNote
US datacenter buildout+21GW (2026) → +84GW (2030)Record pace; global tracked capacity (ex-China) ~89GW (2026) → up to 338GW (2030).
US grid load-request queue~1 terawattAgainst a 759GW US peak; ERCOT alone has ~108GW of large-load requests (with duplicates).
Net-new ELCC grid capacity added~15GW/yr (rising to 20GW+)Headroom turns negative in 2027, ~40GW aggregate deficit by 2030.
AI cloud revenue per MW~$10-13M / MW / yearWhy 'speed is the moat'; 200MW six months early ≈ $1B+ revenue / ~$400-500M NPV.
BTM share of new US DC capacity<7% (2025) → >50% (2028)DC BTM equipment TAM crosses 50GW/year by 2029; 26GW cumulative confirmed BTM by 2030.
800VDC current/loss reduction~14.8-16.7x less current, ~219-278x less I²R600kW rack: ~11.1-12.5kA at 48-54V → ~750A at 800V; enables ~2,300W chips.
HVDC power rack ASP$400-500k/unit (~$0.5M/MW)~10x a standard $40k AC power rack; sidecar TAM peaks ~$11B in 2028.
SST TAM by 2030~$32B (@ ~$1.25M/MW)>$320M flowed into SST startups in the year to Mar 2026; Eaton bought Resilient Power.
Electrical-path efficiency~82% (AC baseline) → ~87.4% (Phase 4)Phase 2 (86.5%) via UPS elimination = ~58MW saved at 1GW IT; Nvidia cites ~5% (~50MW).
Onsite generation capex$1,500-4,000/kWIGT $1,500-1,800; aero/RICE $1,700-2,000; Bloom SOFC $3,000-4,000.
Hyperscaler PUE / WUEPUE 1.08-1.15; WUE 0.20 to >2.0Meta 1.08/0.20; Google 1.10 but WUE >1.0; Microsoft ~0.3 avg but >2 in Phoenix.
GB200 NVL72 rack120kW, 72 GPUs, DLC-onlyHighest-volume Blackwell SKU; chip TDP heading to 1500W+ then ~2,300W.
PJM capacity auction (2025/26)9.3x jump to $270/MW-day~$450 in some zones; +$25-30/month on bills; total ~$16B / ~$120k/MW; 2026/27 capped $329.
Winter Storm Fern (Jan 2026)PJM lost ~21GW; Dominion $1,800/MWh15% of cleared fleet failed; DOE tapped ~35GW of DC/industrial backup; ERCOT held (~$300/MWh).
ERCOT cascade limits>2.6GW system / >2.0GW West TXLoad-loss beyond this risks 60.4Hz danger zone; Iberian blackout (2.2GW, collapse in 27s).
BESS cost (Lazard)100MW: $38-80M (2hr) / $76-157M (4hr)GW-scale ≈ ~$1B, additive to UPS/gensets; synchronous condenser $30-60k/MVAr.
Space vs terrestrial compute cost$8.64 vs $2.37 /hr/GPU (B300, 2026)LCOC $10.91 vs $2.49; parity ~2040 base case (early 2030s in 'Elon' case).
AI share of TSMC N3~60% (2026) → ~86% (2027)Silicon, not power, is the binding constraint; AI ~0.3% of global generation today → ~5% by 2030.
Microfluidic / direct-die coolingTSMC CoWoS-R micropillar cold plate: 4 kW @ 4 LPM, 5.3 kW @ 8 LPM (passes MSL4); Microsoft GH200 microfluidic: ~50% lower package thermal resistanceReliability shown as 9 potential clog events across ~4,370 observations over 6 months (ECTC 2026, SemiAnalysis).
Meta 'titan' gigawatt clusters5 clusters of 1GW+ (Prometheus >3GW, Ohio; Hyperion, Louisiana, 400MW single buildings; El Paso / Iowa / Indiana); Iowa built 1GW in one yearGigawatt single-site build cadence (Meta Superintelligence 1-yr update, SemiAnalysis).

Key Theses

1. The US grid is 'sold out' — headroom goes negative by 2027 and BTM powers half of new datacenters by 2028.
Roughly a terawatt of load requests are queued against a 759GW US peak; barely ~15GW of net-new ELCC capacity is added per year, and less than 10GW/yr of gas in 2026-27. Interconnection now runs ~5 years (7 in North Virginia). On a UCAP basis, grid headroom turns red across a growing set of subregions by 2027. Since a utility's timeline can slip from 2027 to 2029 with no penalty, the marginal buyer goes behind the meter for speed and certainty; DC BTM equipment TAM is set to cross 50GW/year by 2029.
“BTM will power well over half of new US datacenters in 2028+, and the Total Addressable Market (TAM) for DC BTM equipment to cross 50GW/year by 2029.”US Grid Constraints: Towards 40GW+ of Behind-The-Meter Datacenter by 2028? (2026-06-25), SemiAnalysis
2. Speed is the moat: onsite gas is worth billions because six months of earlier revenue dwarfs the extra cost.
An AI cloud nets ~$10-13M per MW annually, so getting 200MW online six months early is worth ~$1-1.2B (an NPV of ~$400-500M even discounted). xAI proved it — Colossus built in ~122 days by bypassing the grid with rented, truck-mounted Solar/Caterpillar turbines and VoltaGrid engines. OpenAI/Oracle then placed the largest-ever onsite order (2.3GW in Texas, Oct 2025), and 12 different suppliers now each hold >400MW of US onsite-gas orders. Power as a % of AI TCO is 'mostly insignificant,' so any power secured is worth billions.
“AI cloud revenue can net $10-12M per MW annually, meaning that getting 200 MW of datacenter powered and online even six months earlier can net $1-1.2 billion in revenue.”How AI Labs Are Solving the Power Crisis: The Onsite Gas Deep Dive (2025-12-30), SemiAnalysis
3. 800VDC is inevitable — physics, not preference — because 600kW+ racks break 48-54V distribution.
Resistive losses scale with current squared. A 600kW Kyber-class rack draws ~11.1-12.5kA at 48-54V but only ~750A at 800V — ~14.8-16.7x less current and, at fixed resistance, ~219-278x lower I²R heating. A 1MW rack at 48-54V needs ~200kg of copper busbar; at GW scale that's hundreds of tons. Today's NVL72 already uses up to 8 power shelves; at Kyber power a 48V approach would need ~64U of power hardware — an entire rack, leaving no room for compute. 800VDC also cuts facility power ~5% (~50MW saved at 1GW IT load).
“800VDC is the physics enabler for 2,300W TDP chips and 600kW racks, and those 600kW racks are the direct consequence of the push for density, because density is what drives cost per token down.”Inside the 800VDC Revolution – Part 1 (2026-05-26), SemiAnalysis
4. The 800VDC transition migrates content from grey space to white space and ends with the Solid-State Transformer.
SemiAnalysis models four phases: (1/2) 2026-28 white-space retrofit via the ~$400-500k power rack (sidecar TAM peaks ~$11B in 2028); (3) 2028/29 grey-space DC distribution with centralized rectifiers, eliminating the central UPS ($1.2M) and AC switchboards/PDUs; (4) 2029+ SSTs collapsing MV transformer + rectifier into one device (SST TAM ~$32B by 2030 at ~$1.25M/MW). Total electrical content per MW stays in a $3.6-4.8M band — it's a content migration, with cumulative path efficiency rising from ~82% (baseline AC) to ~87.4% by Phase 4.
“Total electrical content per MW stays in a $3.6-4.8M band across four of the five architectures we model. The main headline is a content migration from grey space to white space.”Inside the 800VDC Revolution – Part 1 (2026-05-26), SemiAnalysis
5. >1000W chips end the air-cooled era; liquid cooling is the fastest-evolving, highest-risk datacenter system.
The GB200 NVL72 (120kW, 72 GPUs) is DLC-exclusive and is the highest-volume Blackwell SKU; chip TDP is climbing toward 1500W+. Cooling is already the 2nd-largest capex after electrical and carries the steepest learning curve and execution/obsolescence risk. Nvidia's ramp is causing shortages of components 'that on first glance look simple such as Quick Disconnects.' Demand for liquid cooling is underestimated, driving inefficient 'bridge' solutions like RDHx (used with DLC at xAI Memphis) because there won't be enough liquid-ready datacenters.
“Nvidia's massive ramp is causing shortages on components that on first glance look simple such as Quick Disconnects!”Datacenter Anatomy Part 2 – Cooling Systems (2025-02-13), SemiAnalysis
6. Hyperscalers already run 1.08-1.15 PUE with air cooling — debunking the 'air can't be efficient' myth.
The industry-average 1.6 PUE (Uptime survey) excludes hyperscalers; Google/Meta run ~1.1, Microsoft/AWS ~1.15. They win via free cooling (air/water economizers), running server inlets above 30°C (sometimes >40°C) with custom servers, and deep CFD. Meta's 'H' hits 1.08 PUE / 0.20 WUE but takes ~2 years to build a shell (2-3x rivals) — so Meta scrapped in-progress 'H' shells for a faster AI-ready design. Google trades energy for water (PUE 1.10 but WUE >1.0 via evaporative towers).
“Hyperscalers have largely proved wrong the popular narrative that air cooling is not an energy efficient technology.”Datacenter Anatomy Part 2 – Cooling Systems (2025-02-13), SemiAnalysis
7. Gigawatt synchronous AI loads threaten cascade blackouts — an Iberian-Peninsula scenario in Texas.
ERCOT modeled a fault on a West Texas 345kV line: in all four scenarios ≥1.5GW of datacenter load disconnects near-instantly; on a duck-curve day with datacenters not equipped for LVRT, ~2.5GW (all of West Texas) trips at once. If >2.6GW system-wide (or 2.0GW in West Texas) drops, frequency exceeds ERCOT's 60.4Hz danger zone and risks cascade — exactly the April 28, 2025 Iberian blackout, where 2.2GW tripped and the grid collapsed in 27 seconds. Synchronous condensers help but still leave 1.3-1.9GW at risk and cost $10-20M per GW.
“if 2-2.5 GW of datacenter load tripped off the grid in short succession, then similar voltage and frequency fluctuations could cause cascade failures through Texas… All this, potentially started by a squirrel stepping on the wrong power line.”AI Training Load Fluctuations at Gigawatt-scale - Risk of Power Grid Blackout? (2025-06-25), SemiAnalysis
8. AI is being blamed for household electric bills, but the fault is market design (PJM), not AI.
PJM's capacity auction cleared 9.3x higher ($29→$270/MW-day, ~$450 in some zones), adding $25-30/month to bills and driving ~15% higher household costs across 67M residents — off a simulated demand curve (VRR) that PJM itself has repeatedly miscut (a methodology change made 14GW of gas capacity disappear overnight). ERCOT, an energy-only market with no capacity auction, saw equivalent AI buildout yet roughly flat prices. Winter Storm Fern proved the point: PJM's $270/MW-day 'bought' a 21GW generation failure (Dominion spiked to $1,800/MWh) while ERCOT held.
“PJM pays for forward capacity commitments, but subject to testing and event-specific Non-Performance Charges when committed capacity underperforms… PJM's 9.3x capacity price increase was supposed to buy reliability. It did not.”Are AI Datacenters Increasing Electric Bills for American Households? (2026-03-03), SemiAnalysis
9. BESS is the most flexible fluctuation fix but not a cheap one — a GW-scale system costs ~$1B.
Tesla's Megapack can charge/discharge hundreds of MW within seconds, managing fluctuations at MW/ms, MW/s and MW/min — more flexible than capacitor banks, diesel gensets or grid resources. But per Lazard, a 100MW/2hr BESS runs $38-80M and 4hr $76-157M, so a GW-scale system is 'close to a billion dollars' and is additive to (not a replacement for) UPS and gensets. BESS also unlocks demand response: 20-90hrs/yr of curtailment could free 6.5-14.7GW of new ERCOT load with no upgrades — but utility DRMS and incentives remain immature.
“a BESS suitable for a GW-scale datacenter would cost close to a billion dollars, and for that price, Tesla would not consider BESS a replacement for a UPS or a diesel generator.”AI Training Load Fluctuations at Gigawatt-scale - Risk of Power Grid Blackout? (2025-06-25), SemiAnalysis
10. 'Half of 2026 datacenter capacity is canceled' is wrong — the panic is built on a flawed denominator.
The viral claim traces to a Bloomberg/Sightline figure that only tracks large, publicly-announced megaprojects (~12GW basis, of which ~5GW 'confirmed'), skewing toward speculative builds by inexperienced developers that were never landing in 2026 anyway. SemiAnalysis's building-by-building satellite tracking shows those slip to 2028+, not cancel, and its 2026 outlook hasn't moved. The real signal: over a terawatt now sits in the US large-load queue, and leading operators are adept at navigating permitting, grid and equipment constraints.
“the data sources behind these claims of '50% of 2026 datacenters are delayed' are essentially uninformed vibe-coded datacenter forecasts that take announcements at face value.”Stop Saying Half of 2026 US Datacenter Capacity Is Canceled (2026-06-18), SemiAnalysis
11. Space datacenters stay ~4x costlier than Earth until ~2040 — terrestrial power isn't exhausted yet.
A 30.5kW B300 cluster costs $8.64/hr/GPU in space vs $2.37 terrestrial (LCOC $10.91 vs $2.49 after redundancy), largely because launch is $1.6M of the $3.1M datacenter capex and space useful life is only ~5yr vs 15yr (levelized DC capex ~18x higher). SemiAnalysis maps five terrestrial supply layers (grid → converted/crypto sites → BTM → industrial production → the semiconductor ceiling) that must exhaust before orbit makes sense. The binding constraint today is actually silicon: AI consumes ~60% of TSMC N3 in 2026, ~86% in 2027 — a limit space cannot solve.
“In our base case scenario, the cost difference between space and terrestrial datacenters starts at over 4x in 2026, before narrowing to parity in ~2040.”To Boldly Go: The Case for Space Datacenters (2026-06-03), SemiAnalysis
12. Heavy-duty turbines are the tightest supply chain; the industry is scarred by two boom-bust cycles.
GE Vernova, Siemens Energy and Mitsubishi are order-booked into 2028-29 with nonrefundable reservations beyond, yet respond with caution — GE targets just 24GW/yr (its 2007-16 level), Siemens ~30GW by 2028-30, all without expanding factory footprint. The bottleneck is turbine blades: exotic single-crystal nickel alloys (rhenium, cobalt, tantalum, yttrium — yttrium is under Chinese export control) made by just four Western firms (PCC, Howmet, CPP, Doncasters). Heavy cores are 300-500-ton logistics nightmares needing 24-30 months. Aeros, IGTs and RICE are far less constrained.
“few manufacturers appear fully 'AGI-pilled' despite surging demand… much of it reflects PTSD from 30 years of boom-bust cycles in gas generation.”How AI Labs Are Solving the Power Crisis: The Onsite Gas Deep Dive (2025-12-30), SemiAnalysis
13. ±400V bipolar won the 800VDC voltage war by borrowing the EV supply chain, but no standard has settled.
OCP Diablo 400 (Google/Meta/Microsoft) defines ±400V bipolar as base config so each rail sits only 400V from the grounded midpoint, letting operators use mature EV-grade 650V GaN FETs and 400V-class caps — Google: '400VDC allows us to leverage the supply chain established by electric vehicles.' But there's no one-size-fits-all: Nvidia sits outside Diablo 400 with a monopolar 660kW 800V design; Meta runs 600-800kW, Google pushes 900kW, Amazon lands 800kW. Full NEC code support only targets 2029, so pre-2029 sites need custom AHJ approvals — a barrier for colos, not hyperscalers.
“selecting 400 VDC as the nominal voltage allows us to leverage the supply chain established by electric vehicles, for greater economies of scale, more efficient manufacturing, and improved quality and scale.”Inside the 800VDC Revolution – Part 1 (2026-05-26), SemiAnalysis
14. The datacenter water panic targets the wrong problem — Colossus 2 uses ~2.5x one In-N-Out store.
SemiAnalysis modeled xAI's 400MW Colossus 2 (PUE 1.15, 70% utilization, ~30% wet operation) at ~346M gallons/yr (0.9M gal/day, WUE ~0.51 L/kWh). A single In-N-Out store (burgers only) consumes ~147M gallons/yr — so the whole gigawatt-class datacenter is only ~2.5:1 versus one burger joint. With 400+ In-N-Out locations and hundreds of thousands of burger restaurants, the water argument for slowing datacenters is misdirected. Cooling architecture (dry vs wet) and onsite gas generation water use dominate the honest accounting.
“Colossus 2's blue water footprint is around 346 million gallons per year, while an average In-N-Out store (yes, burgers only) comes in at around 147 million gallons. That's roughly a ~2.5 : 1 ratio.”From Tokens to Burgers: A Water Footprint Face-Off (2026-01-15), SemiAnalysis

Article Deep-Dives

The definitive map of the rack-power transition. As racks climb toward 600kW+ (Kyber) and chips toward 2,300W, 48-54V distribution physically breaks (I²R losses scale with current squared), forcing 800VDC. SemiAnalysis models four phases: (1/2, 2026-28) white-space retrofit via the ~$400-500k HVDC power rack (OCP Diablo 400, co-authored by Google/Meta/Microsoft; ±400V bipolar to reuse the EV supply chain); (3, ~2028/29) grey-space DC distribution with centralized rectifiers, killing the central UPS ($1.2M) and AC switchboards; (4, 2029+) Solid-State Transformers collapsing MV transformer + rectifier into one device. Sidecar TAM peaks ~$11B (2028); SST TAM ~$32B by 2030. Total electrical content per MW stays $3.6-4.8M — a migration from grey to white space — while path efficiency rises 82%→87.4%. Winners: Delta, Infineon, Wolfspeed (10kV SiC), DG Matrix, Amperesand, Heron Power. Barriers: NEC 2029 code, DC arc-flash safety, grounding fragmentation, and a 3,300V+ SiC supply gate.
The BYOG bible. With the grid 'sold out' (Texas approves barely >1GW/yr against tens of GW/month of requests) and interconnection ~5 years, AI labs bypass the grid entirely. xAI pioneered rented, truck-mounted 16MW Solar turbines + VoltaGrid engines to build Colossus in ~122 days; OpenAI/Oracle then placed a 2.3GW Texas order. Full taxonomy of generation: aeroderivatives (jet engines bolted down, 30-60MW, 5-10min ramp, $1,700-2,000/kW), IGTs ($1,500-1,800/kW), reciprocating engines (RICE, Jenbacher J624), and Bloom SOFC fuel cells ($3,000-4,000/kW, no combustion, near-instant permitting). Deployment configs: bridge power, islanded, N+1+1 redundancy, gas+battery hybrids, Energy-as-a-Service (VoltaGrid). Load surges need synchronous condensers, flywheels or BESS. The binding constraint is heavy-duty turbine blades — single-crystal nickel alloys from 4 Western firms — and 30 years of boom-bust PTSD keeping GE/Siemens/MHI cautious.
The comprehensive cooling primer. Modern datacenters generate >50x an office's heat/sq-ft; cooling is 60-80% of non-IT power and the 2nd-largest capex. Walks the full heat path — server fans, CRAC/CRAH/fan walls (500-600kW), RDHx (30-50kW), CDUs (>1MW), water- vs air-cooled chillers (COP ~7), evaporative vs dry cooling towers, WUE (>2 L/kWh for wet). Debunks the myth air can't be efficient: hyperscalers hit 1.08-1.15 PUE via free cooling, >30°C server inlets, custom servers, and CFD (Fan Law: -10% airflow = -27% energy). The GenAI shift: the GB200 NVL72 (120kW, DLC-exclusive) forces direct-to-chip liquid cooling; Nvidia's ramp is shorting simple parts like Quick Disconnects. Case studies of Microsoft (direct evaporative), Meta ('H', 1.08 PUE but slow), Google (energy-for-water), AWS.
Why synchronous GPU training is uniquely dangerous for the grid. AI loads swing ~15x more than a cloud datacenter (Google: 1.5MW→15MW) with MW-scale sub-second spikes from batch cycles, checkpointing and AllReduce. Inverter-based solar/BESS lack the intrinsic mechanical inertia spinning generators provide, though grid-forming inverters can supply synthetic/virtual inertia and fast frequency response. ERCOT modeled a West Texas 345kV fault: ≥1.5GW of datacenter load disconnects near-instantly; on a duck-curve day, 2.5GW (all of West Texas) trips at once. Beyond 2.6GW system-wide (2.0GW West Texas), frequency exceeds the 60.4Hz danger zone and risks cascade — mirroring the April 2025 Iberian blackout (2.2GW tripped, 27-second collapse). Solutions: synchronous condensers ($10-20M/GW, still leave 1.3-1.9GW at risk), and BESS — the most flexible fix (MW/ms→MW/min) but ~$1B at GW-scale, additive to UPS/gensets. Demand response could free 6.5-14.7GW in ERCOT but utility DRMS is immature.
Models the exact grid shortfall BTM must fill. From 40,000 tracked generators, only ~15GW/yr of net-new ELCC capacity is added; <10GW/yr of gas in 2026-27 as slow CCGTs dominate utility orders and transformer/turbine lead times stretch to 3-4 years. On a UCAP basis, headroom turns red across a growing set of subregions by 2027 (PJM's 2027/28 auction cleared ~6.5GW short of its 20% target). BTM wins on speed and certainty — utility timelines slip with no penalty (Switch posted a multi-billion LOC facility) while AI labs accept lower redundancy (Meta targets 2 nines, no gensets). Details ERCOT's Batch Zero co-location framework (NMA, BYOG, WLPUN, PCLR) codifying hybrid grid+onsite structures, with ~2,885MW of announced co-location deals (Crusoe, AWS/Comanche Peak, CyrusOne).
A market-design autopsy. PJM's capacity auction (Base Residual Auction) cleared 9.3x higher ($29→$270/MW-day, ~$450 some zones), adding $25-30/month to 67M residents' bills — off a simulated VRR demand curve PJM keeps miscutting (a methodology change made 14GW of gas 'disappear' overnight). The IMM found removing all datacenters would cut capacity payments $9.33B (64%). But ERCOT, an energy-only market, saw equivalent AI buildout and roughly flat prices. Winter Storm Fern (Jan 2026) was the proof: PJM's record capacity price 'bought' a 21GW generation failure (Dominion $1,800/MWh) while ERCOT held. Verdict: the bill shock is government/market-design driven, not AI — and ERCOT's single-state regulatory agility (SB6) is 'durable and investable' vs PJM's 13-state FERC gridlock.

Reference: Value Chain

Power Generation & Grid SupplyGE Vernova, Siemens Energy, Mitsubishi Power (heavy turbines); Caterpillar/Solar, Wärtsilä, Bergen, Everllence, Jenbacher/INNIO (engines); Bloom Energy (fuel cells); Doosan Enerbility, ProEnergy, Boom Supersonic (new entrants); IPPs Vistra, Constellation, Talen, NRG — Supplies the electrons — increasingly onsite (BYOG) because the grid can't add firm capacity fast enough. The 'Big Three' turbine makers are order-booked into 2028-29; a dozen suppliers each hold >400MW of US datacenter orders.
Grey-Space Electrical (facility-level)ABB, Schneider Electric, Eaton, Vertiv, Legrand (switchgear/UPS/transformers); SST startups DG Matrix, Amperesand, Heron Power, Novos Power; Wolfspeed, Infineon (SiC/GaN); LS Electric, EPEC (DC breakers) — Steps utility MV down to usable voltage and distributes it. The 800VDC transition eliminates central low-voltage UPS ($1.2M), migrates conversion to DC rectifiers/SSTs, and forces new DC-rated busway, switchgear and solid-state circuit breakers.
White-Space Power (rack-level)Delta, Advanced Energy, TE Connectivity, Infineon (power racks/PSU/BBU); Rittal (cabinets); supercapacitor & BBU vendors; Nvidia (660kW 800V reference) — The power rack / battery rack — the headline new-equipment category. Rectifies AC→800VDC, houses BBUs and supercapacitors, steps down to ~50V then sub-1V at the GPU VRM. Content per MW ~$0.5M (power rack) to ~$0.2M (battery rack).
CoolingVertiv, Boyd, Cooler Master, nVent, Delta (CDUs/cold plates); Danfoss (Turbocor compressors, VFDs); Schneider, Mitsubishi Heavy (chillers); quick-disconnect & manifold vendors — Rejects the heat: air (CRAH/fan wall + chillers/cooling towers) for legacy, but >1000W chips force DLC (cold plates + CDUs). Cooling is the fastest-evolving, highest-execution-risk system and the 2nd-largest capex after electrical.
Datacenter Developers & OperatorsHyperscalers (Google, Meta, Microsoft, AWS, Oracle); AI labs (OpenAI, xAI, Anthropic); neoclouds/EaaS (Crusoe, VoltaGrid, Fluidstack, Nebius); crypto-converters (Core Scientific, IREN, Cipher, Applied Digital, TeraWulf); colos (Switch, CyrusOne, Vantage) — Integrate the whole stack and take the timeline risk. Vertically-integrated hyperscalers (Google/Meta) push aggressive designs (distributed UPS, no gensets, 2-nines uptime); AI labs prioritize speed and accept lower redundancy.
Grid Stability & Storage EnablersTesla (Megapack BESS); flywheel & synchronous-condenser vendors (Piller, Baldor); ABB HiPerGuard MV UPS; supercapacitor makers; EPC/interconnection specialists (Aran Industries) — Absorb sub-second load swings and prevent cascade blackouts. BESS handles MW/ms→MW/min fluctuations, provides fast frequency response and demand-response; synchronous condensers/flywheels add inertia. A GW-scale BESS can cost ~$1B.

Reference: Core Concepts

PUE / WUE (Power & Water Usage Effectiveness). PUE = total facility power / IT power; industry-average PUE is ~1.6 (Uptime survey excluding hyperscalers), but Google/Meta run ~1.1 and Microsoft/AWS ~1.15. WUE (liters/kWh) can exceed 2 for water-cooled chillers; Meta ~0.20, Microsoft ~0.3 company-wide but >2 in dry Phoenix. Non-IT power is 60-80% cooling, 15-30% electrical losses. Fan Law: a 10% airflow cut yields ~27% fan-energy cut (cubic).

Behind-the-Meter (BTM) / Bring Your Own Generation (BYOG). Generating power onsite (gas turbines, engines, fuel cells) rather than waiting years for a grid connection. 'Speed is the moat': an AI cloud nets ~$10-13M revenue per MW annually, so bringing 200MW online six months early is worth ~$1B+. BTM is expected to power over half of new US datacenter capacity by 2028, up from <7% in 2025.

800VDC / ±400VDC & the Power Rack (Sidecar). Distributing ~800V DC to the rack instead of 48-54V. A 600kW rack at 48-54V needs ~11.1-12.5kA vs ~750A at 800V (~14.8-16.7x less current, I²R losses ~219-278x lower). The 'power rack' (OCP Diablo 400, co-authored by Google/Meta/Microsoft) disaggregates AC-DC rectification, BBUs and supercaps into a dedicated cabinet (~$400-500k ASP, ~10x a standard AC power rack). ±400V bipolar leverages the mature EV supply chain (650V GaN, 400V caps).

Solid-State Transformer (SST). Replaces the iron-core transformer + rectifier with high-frequency (20kHz+) semiconductor conversion, going directly from medium-voltage AC (13.8-45kV) to 800VDC. Infineon claims ~40x weight and ~14x size reduction; targets >97% efficiency (best public benchmark: ETH Zurich 98% at 400kW). Depends on 3,300V+ SiC MOSFETs still in limited production. The '800VDC endgame' (Phase 4, ~2029+).

Direct-to-Chip Liquid Cooling (DLC) & CDU. Cold plates carry coolant directly against the silicon; the GB200 NVL72 (120kW, 72 GPUs) is DLC-exclusive. A Coolant Distribution Unit (CDU, typically >1MW) exchanges heat between the IT loop and facility water. Nvidia's ramp is causing shortages of seemingly simple parts like Quick Disconnects. Rear-Door Heat Exchangers (RDHx, 30-50kW) are a bridge; xAI Memphis combined RDHx + DLC.

ELCC (Effective Load Carrying Capability) & Grid Headroom. ELCC = the 'true' firm-capacity value of a plant to the system, discounting intermittent solar/wind steeply and thermal forced-outage risk. Incremental 4hr BESS adds little marginal ELCC once <4hr risk is saturated. Grid headroom = accredited supply − peak demand − required reserves; on a UCAP basis it turns red (negative) across a growing set of US subregions by 2027, an aggregate ~40GW deficit by 2030.

Load Fluctuations, Inertia & LVRT. Synchronous GPU training swings power ~15x more than a cloud datacenter (Google: 1.5MW→15MW), with MW-scale spikes sub-second. Synchronous generators provide intrinsic mechanical inertia; inverter-based solar/BESS provide none intrinsically, but grid-forming inverters can provide synthetic/virtual inertia and fast frequency response. The disturbance is a transient voltage sag (30ms-5s); low-voltage ride-through (LVRT) is the capability/requirement to stay connected through it. If datacenters trip during one, ERCOT modeled 1.5-2.5GW of West Texas load disconnecting near-simultaneously — echoing the April 2025 Iberian blackout (collapse in 27 seconds).

Capacity Auction (PJM BRA) vs Energy-Only (ERCOT). PJM pays plants to stay on standby via a forward capacity auction (Base Residual Auction), priced two years out off a simulated demand curve (VRR). The 2025/26 clearing jumped 9.3x to $270/MW-day (some zones ~$450), driving ~15% higher household bills. ERCOT has no capacity market — real-time scarcity pricing only — and Texas power prices stayed roughly flat despite equivalent AI buildout. SemiAnalysis's verdict: the bill shock is government/market-design driven, not AI.

Aeroderivative / IGT / RICE / SOFC (onsite generation types). Aeroderivatives = jet engines bolted down (30-60MW, 5-10min ramp, $1,700-2,000/kW). Industrial Gas Turbines (IGTs, 5-50MW, ~20min ramp, $1,500-1,800/kW). Reciprocating engines (RICE, 3-20MW, 10min ramp, $1,700-2,000/kW) — Jenbacher J624 is a favorite. Solid-oxide fuel cells (Bloom Energy, no combustion, $3,000-4,000/kW, near-instant permitting). Heavy-duty CCGTs are most efficient (>60%) but slowest (4-6yr COD).

Open Questions

Sources (SemiAnalysis)

Independence & sourcing. This is independent analysis by Yicheng Yang, distilled from publicly accessible SemiAnalysis articles (free posts and free previews; no paywall circumvention) and verified against the underlying text. It is not affiliated with, endorsed by, or a substitute for SemiAnalysis — subscribe there for the full research. All referenced claims are sourced and linked per SemiAnalysis's attribution terms. No SemiAnalysis images are reproduced. Nothing here is investment advice.