AI Supply Chain Research · Sector 08

Hyperscalers & AI Economics

The trillion-dollar buildout, decoded: who pays for GPUs, who captures the margin, and why a token is now the unit of economic value.  ·  ← back to the series  ·  Background: Primer §00, §04, §08, §09, §10

At a glance
Winners
Amazon / AWSAnthropicOracleGold-tier neoclouds (CoreWeave, Nebius, Crusoe, Fluidstack)Nvidia & TSMC
Bottlenecks
Power / multi-GW PPAs — watts gate GPU deployment; securing power is now a share-determining act.TSMC N3 wafer allocation — every accelerator roadmap converged on N3; tightest node in the system.HBM / memory — prices up ~6x in a year; a once-in-four-decades shortage.Fault-tolerant training software — no framework is both free and overhead-free; concentrates reliable scale training.Frontier-model access — only AWS/Azure/Google can sell all three frontier families; a distribution moat.
Risks
Depreciation debate: if GPU useful life proves shorter than 5-6yr, hyperscaler earnings are overstated (Burry thesis).Circular financing: Amazon/MSFT/Google invest in labs that then buy their compute — revenue that could reverse.Token budgets tightening: enterprise caps could slow API growth if the s-curve stalls (SemiAnalysis says overblown).Nvidia/TSMC repricing: 'central bank' restraint could end, venting margin away from labs/neoclouds.Overbuild / ROIC: Oracle/neocloud AI ROIC (~20%) well below legacy cloud; a demand air-pocket compresses returns.
Catalysts
New OpenAI–Microsoft deal reaccelerating Azure; 100% of OpenAI API inference through 2032.Vera Rubin (VR NVL72) ramp: step-jump in perf/TCO resetting the value-capture math.Next enterprise verticals (cyber via Mythos re-release, white-collar via Cowork/Copilot/Codex).Trainium3 ramp + Project Rainier scaling as a proof point for hyperscaler ASIC economics.ClusterMAX 3.0 (B300/GB300, 800Gb) re-rating neocloud tiers and TCO leadership.GPU-backed debt markets maturing (pre-agreed price backstops, DSCR-able) unlock non-hyperscaler compute financing.

Overview

SemiAnalysis treats the AI buildout not as a hardware story but as an economics story stacked on top of hardware. Their unifying frame is the 'AI Token Economic Stack' — Energy → Datacenter → Accelerator → Neocloud/Hyperscaler → Model Lab → Application — with a token as the atomic unit that ties watts to revenue. Two proprietary engines anchor almost every note: the Datacenter Industry Model (building-by-building MW tracking, the best leading indicator of capex) and the Tokenomics / AI Cloud TCO Model (every major compute contract, its cost, margin, ROIC and RPO). The core, repeated theses: (1) headline GPU $/hr is a bad proxy for true cost — real TCO is dominated by goodput (useful work), reliability, storage, networking, setup and debugging, which SemiAnalysis quantifies via ClusterMAX and a public Goodput calculator; (2) value capture has violently rotated — 2023-25 all the money went to infra (Nvidia, power, memory), but from ~December 2025 'agentic AI began to really work' and the model labs suddenly captured the value, with Anthropic inference gross margins going from 38% to >70% and ARR from $9B to $30-44B in months; (3) the hyperscalers diverge on business model — AWS's token-as-a-service (Bedrock) + vertical silicon (Trainium/Graviton) is structurally higher-margin than Azure's/GCP's IaaS mix, while Microsoft's 'Big Pause' handed Oracle ~$150B of OpenAI gross profit; (4) Nvidia and TSMC deliberately under-price to scarcity, acting as the 'central bank of AI' to keep the ecosystem expanding rather than extract maximum rent. Underneath sit older but foundational pieces (GPU cloud economics, the AWS 'cloud crisis', LLM search cost, DeepSeek tokenomics) that established the vocabulary the market now uses.

Positioning: Who Wins and Why

Amazon / AWS — Only CSP with token-as-a-service (Bedrock) as its dominant AI mix + vertical silicon (Trainium >50% of Bedrock tokens, leading Graviton) — structurally highest margin. Anthropic offtake circular but real. Building more capacity than anyone in 2025-27.

Anthropic — The purest value-capture play: ARR $9B→$30-44B, inference GM 38%→70%+, 90%+ B2B/API driven by Claude Code. Frontier pricing power means margins won't be competed away.

Oracle — ~20% structural CapEx advantage (networking + ODM + IG debt), ~$300B OpenAI backlog (>$420B = total RPO across all customers), ByteDance/Johor engine. But ROIC ~20% (below MSFT 35-40%) and margins disappointed the market.

Gold-tier neoclouds (CoreWeave, Nebius, Crusoe, Fluidstack) — Reliability/support translates into a real TCO discount that commands price premium; rental prices +40% off Oct-2025 bottom lift returns.

Nvidia & TSMC — Still own the scarcest layers (GPUs, N3) and deliberately hold pricing below scarcity — dry powder to reprice later. GB300/Rubin step-change in perf/TCO keeps them indispensable.

Key Data

MetricValueNote
Gold vs Silver-tier cluster TCO gap (equal $/GPU-hr)Silver ~15% more expensive (1.15x); Hyperscaler 1.10-1.61xDriven by goodput, setup, storage — not headline price.
Goodput expense, large LLM pretrain (Gold/Hyperscaler/Silver)6.14% / 10.53% / 20.91%Collapses to 0.23-0.96% for many small jobs.
Anthropic ARR trajectory$9B → $30B (added $21B net-new in 1Q26) → ~$44B YTD; model sees >$100B by year-endDriven by Claude Code / enterprise API (90%+ B2B).
Anthropic inference gross margin-94% (2024) → 38% (2025) → mid-60s / >70% (2026)WSJ: operating-income profitable in 2Q26 ex-SBC.
Oracle OpenAI contract value / gross profit~$300B OpenAI contract → ~$150B gross profit (>$420B = Oracle total RPO, all customers)~$30B/yr would have lifted MSFT FY25 gross profit ~15%.
AWS AI revenue mix & margin moveAI = 10% of AWS (from 2% in 1Q24); EBIT margin +213bp Q/QBedrock now 37% of AWS AI vs GCP/Azure IaaS 80%+.
Bedrock run-rate & growth~$5.5B run-rate; +170% Q/Q (1Q26), +60% Q/Q (4Q25); 80-90% AnthropicTrainium powers >50% of Bedrock token usage.
GB300 NVL72 vs H100 throughput / TCO~17x (FP8) to 32x (FP4) throughput at only ~70% higher TCO/GPUSoftware alone can 14x B300 throughput (wideEP+disagg+MTP).
Opus effective blended token price~$0.99/MTok vs $5/$25 sticker~300:1 input:output, 90%+ cache hit, cached input $0.50/MTok.
SemiAnalysis's own token spend~30% of employee comp; ~5B tokens/mo/employee (5x Meta); peak $10.95M/yr on ClaudePower-law: some staff run 100B+ tokens/mo.
Enterprise token budgets (per employee/mo)$250 (aerospace) / $500 (pharma) → ~$2,000 (Stripe/Workday) → tens of thousandsNo convergence; Uber capped at $1,500/mo after 4-month burn.
Ramp AI spend distribution99th pct ~$90k/yr/emp; 90th ~$7,300; median $136Median Fortune 500 still <$100/employee.
GPU vs CPU server TCO splitGPU: $7,025 capital vs $1,871 hosting/mo; CPU: $301 vs $220GPU cloud is capital-dominated → capital is the only real barrier.
H100 all-in cost vs rental~$1.525/hr all-in (13% debt) vs ~$2/hr best deals, $3+ when fleeced1-yr H100 rentals +40% off Oct-2025 bottom.
Microsoft datacenter freeze>2GW non-binding LOIs dropped; 1.5GW near-term self-build frozen; ~5GW binding remainsMSFT drove >60% of new leased turnkey capacity Q1'23-Q2'24.
Microsoft Fairwater datacenter scale~300MW GPU buildings (>150k GB200 each); 3rd phase 600MW+ buildingsLargest individual datacenters on Earth if built on time.
Coding share of frontier-lab ARR>70% of OpenAI + Anthropic ARRAnthropic 90%+ B2B vs OpenAI ~60% consumer.
Project Rainier (Anthropic × AWS)400k Trainium2 chipsFunded by Amazon's $4B+ circular Anthropic investment.
AI financing scale>$7T AI debt by 2029; ~$11.1T cumulative AI capex 2024-29; >$2T/yr by 2028Nvidia GPU debt backstop framing (SemiAnalysis).
Meta compute contracting>5GW of Cloud+Colo contracted in H1 2026; 2.5GW under construction across two largest campusesHyperscalers turning into neoclouds; Meta in final talks with Anthropic for private Claude instances (Bedrock-type token service) (Meta Compute).
Anthropic 3Q26 economicsprofit >$1B in 3Q26; confidentially filed for IPO 2026-06-01; Anthropic + OpenAI ≈ $100B combined ARRSemiAnalysis floats a $6T valuation as a base-case possibility (Anthropic 3Q26).

Key Theses

1. GPU $/hr is a trap: real cluster TCO is 10-60% higher once you count goodput, setup and storage.
Holding GPU pricing equal, SemiAnalysis's TCO+Goodput calculators show a Gold-tier cloud is ~5-15% cheaper on TCO than Silver-tier across large training workloads, and Hyperscalers run >10% more expensive (support + EFA tuning). In a large LLM pretrain, goodput expense alone was 6.14% (Gold) vs 10.53% (Hyperscaler) vs 20.91% (Silver). But for fault-tolerant single-node inference the gap collapses to near zero — reliability only matters when your job is big relative to the cluster.
“even in scenarios where pricing per GPU-hour is equal, there are always hidden costs across Storage, Network, Control Plane, Support, Goodput, Setup, and Debugging expenses.”How Much Do GPU Clusters Really Cost? (2026-04-20), SemiAnalysis
2. Value capture violently rotated from infra to the model labs in December 2025 when agentic AI 'began to really work'.
2023-25 all AI money went to infra (Nvidia +25% AH in May 2023; power names led 2024; memory led 2025). Then agentic AI crossed an inflection: tasks worth thousands of person-hours now cost a few dollars of tokens. Anthropic ARR exploded $9B→$44B and inference gross margins went 38%→>70% in the same window. The labs went from capturing almost none of the value to capturing all of it.
“The age of low gross margins for frontier model providers is over. Real agentic AI has permanently increased the market-clearing price per token, and there's no going back.”AI Value Capture - The Shift To Model Labs (2026-05-01), SemiAnalysis
3. Anthropic's Bedrock deal structure — not just volume — is what inflected AWS margins above every other CSP.
On Bedrock, Anthropic is seller-of-record (books full token revenue) while AWS collects BOTH an infra fee AND a distribution/revenue-share fee — a margin stack IaaS can't match. AWS EBIT margins rose 213bp Q/Q while Oracle, CoreWeave and Azure margins fell, despite AWS running the shortest (5yr) depreciation. Bedrock is a ~$5.5B run-rate business (80-90% Anthropic) and 37% of AWS AI revenue vs IaaS still 80%+ of Azure/GCP AI.
“Our AWS Trainium chips, designed in-house for AI workloads, now power more than 50% of Amazon Bedrock token usage.”Anthropic Growth and Bedrock Mix Drive AWS Margins Higher While Peers Lag (2026-05-27), SemiAnalysis
4. Microsoft's 'Big Pause' handed Oracle ~$150B of OpenAI gross profit — a self-inflicted competitive gift.
Microsoft drove >60% of all new leased turnkey capacity Q1'23-Q2'24, then froze — walking away from >2GW of non-binding LOIs and 1.5GW of near-term self-build. OpenAI diversified to Oracle, CoreWeave, Nscale, Amazon, Google. Oracle signed ~$300B of OpenAI contract value (~$150B gross profit; the >$420B figure is Oracle's total RPO across all customers, not the OpenAI deal alone); at ~$30B/yr that would have lifted Microsoft's $194B FY25 gross profit by ~15%. SemiAnalysis frames it as partly deliberate (OpenAI would be ~50% of Azure at worse ROIC) but mostly a fumble.
“They potentially have just allowed a competitor to fund their own entry into the AI factory business!”Microsoft's AI Strategy Deconstructed - From Energy to Tokens (2025-11-12), SemiAnalysis
5. Nvidia and TSMC deliberately under-price to scarcity — acting as the 'central bank of AI'.
With compute demand far above supply, both could reprice hard but don't. N3 is the tightest node (Nvidia, Broadcom, Annapurna, MediaTek, AMD all fighting for wafers) yet pricing is stable; Nvidia won't fully reprice Rubin to capture its performance/memory-cost gains. Motive: avoid antitrust scrutiny, prevent customer diversification, and keep downstream labs profitable so total demand keeps growing. They 'take the oxygen out of the room' to remain the protagonist.
“By taking the oxygen out of the room – Nvidia aims to ensure it remains the main protagonist in the AI era for the foreseeable future.”AI Value Capture - The Shift To Model Labs (2026-05-01), SemiAnalysis
6. A Blackwell GB300 NVL72 delivers ~17x (FP8) to 32x (FP4) the throughput of an H100 for only ~70% more TCO — the real driver of falling token cost.
Token production cost has plummeted because accelerator price increases are more than offset by throughput gains. On the same B300, software alone (wideEP + disagg + MTP) takes DeepSeek R1 from ~1k to ~14k tok/s/GPU — a 14x software-only swing. Stack hardware on top and GB300 NVL72 hits ~17x an H100 in FP8, 32x in FP4 (Hopper lacks native FP4), while costing only ~70% more per GPU. This is why labs can cut sticker prices (Opus 4.5 at $5/$25, down from $15/$75) and still expand margins.
“One can 14x throughput with software improvements alone.”AI Value Capture - The Shift To Model Labs (2026-05-01), SemiAnalysis
7. Sticker price per token is meaningless: blended agentic price is a fraction of it because of caching and input/output ratios.
SemiAnalysis estimates the true blended price for running Opus on agentic tasks at ~$0.99/MTok despite a $5/$25 sticker. Agentic workloads have ~300:1 input:output ratios and 90%+ cache hit rates; cached input costs only $0.50/MTok, so most tokens land in the cheapest tier. Analysts constantly get this wrong (I/O ratios, cached tokens, pricing). This is also why DeepSeek's cheap sticker ($0.55/$2.19) didn't commoditize the market — token price is an output of KPI tuning, not a fixed input.
“a model's price per token is an OUTPUT of these KPI decisions which can be tuned based on the model providers' hardware and model setup.”DeepSeek Debrief: >128 Days Later (2025-07-03), SemiAnalysis
8. AWS's 'cloud crisis' (2023) and 'AI resurgence' (2025) are the same call: business-model, not chips, decides who wins the GPU cloud era.
In 2023 SemiAnalysis warned AWS would lose the next compute wave despite owning the most servers — culture and business choices, not tech, would hamper it (Titan/Olympus models failed, Trainium1 uncompetitive). By Sept 2025 they made the out-of-consensus reversal: Anthropic (revenue 5x YTD to $5B annualized) + Trainium's best-perf/TCO-per-memory-bandwidth for inference/RL + token-as-a-service would reaccelerate AWS beyond 20% YoY. Both calls pivot on the same axis: does the provider have the model access and margin structure, not just the silicon?
“Technological prowess, culture, and/or business decisions will hamper them from capturing the next wave of cloud computing like they have the last two.”Amazon's Cloud Crisis: How AWS Will Lose The Future Of Computing (2023-03-20), SemiAnalysis
9. Oracle out-competes on structural cost, not software: ~20% capex advantage from networking, ODM sourcing and investment-grade debt.
Oracle's OCI, latecomer to cloud, wins large GPU contracts via three cost levers: (1) optimized RoCEv2 networking (built on Sun/Mellanox HPC heritage) cutting cluster CapEx ~20% vs a neocloud giant; (2) buying servers ODM-direct from Foxconn, bypassing Dell/Supermicro margin; (3) the market's lowest cost of capital via investment-grade debt vs neoclouds' double-digit debt. The ByteDance relationship made Johor, Malaysia the world's #2 AI hub. Larry's >$130B booking / >$420B RPO surge validated SemiAnalysis's early capex ramp call.
10. Trainium is a targeted weapon, not a Nvidia-killer: it wins on perf/TCO where memory bandwidth is king (high-batch inference, RL).
Trainium2 (~500W, 667 TFLOP/s BF16, 96GB HBM3e) course-corrected after uncompetitive Trainium1/Inferentia2. AWS's 400k-Trainium2 'Project Rainier' cluster for Anthropic is effectively funded by Amazon's own $4B+ Anthropic investment circling back. SemiAnalysis is blunt: Amazon is a 'distant second' to Google in custom silicon, Trainium is unproven for frontier training, and most volume will be inference. But its natural offtake via Bedrock (where hardware is abstracted, porting takes days) makes it structurally valuable.
“Amazon's new $4 billion investment in Anthropic is effectively going to find its way back into the Project Rainier 400k Trainium2 cluster, and there are no other major customers yet.”Amazon's AI Self Sufficiency | Trainium2 Architecture & Networking (2024-12-03), SemiAnalysis
11. 'Tokenmaxxing' headlines are overblown; enterprise token budgets are real but the adoption s-curve is still early.
After Meta's leaderboard (60T+ tokens/30 days, top user ~280B) and Uber's $1,500/mo cap made news, SemiAnalysis talked to 50+ enterprises. Reality: even Meta (~$50k/yr/employee at list) is only a 3-5% Anthropic customer; the 90th+ percentile drives most revenue and is at little risk. Ramp data: 99th pct spends ~$90k/yr/employee, 90th ~$7,300, but the MEDIAN Ramp customer just $136 and the median Fortune 500 <$100. Budgets range wildly ($250 aerospace → $2,000 Stripe/Workday → tens of thousands). Coding is the first vertical (>70% of OpenAI+Anthropic ARR); cyber and white-collar are next.
“Even Meta, who was burning through 70T tokens per month in February... is only a 3-5% customer for Anthropic per our estimates.”TokenBudgeting: Our Conversations with Enterprises on Token Spend (2026-06-30), SemiAnalysis
12. Extended server depreciation is defensible, not accounting fraud — GPUs stay economically useful past the product cycle.
Michael Burry alleged hyperscalers boost earnings by stretching useful life from 3-5yr to 5-6yr on a 2-3yr Nvidia cycle. SemiAnalysis calls this a 'fatal flaw': supercomputers prove longevity (IBM Summit ran 6.5yr; Fugaku/Sierra still on Top500; V100 from 2017 still rented on AWS/Lambda). Warranties now extend 6-7yr; the driver is reliability + incentives, not fiction. Notably AWS runs the SHORTEST depreciation (5yr) of all CSPs yet shows the only rising margin trend — the opposite of what an earnings-management story would predict.
“Dr. Burry's claim is predicated on an assumption that the NVIDIA product cycle is now 2-3 years... We believe this is a fatal flaw in the argument.”Microsoft's AI Strategy Deconstructed - From Energy to Tokens (2025-11-12), SemiAnalysis
13. Fault-tolerant training is a moat: it's a 'secret sauce' available only to frontier labs or those who pay for a software license.
At 1k+ GPU scale you can't rely on the provider to catch every failure; job restarts cost 10-15 min plus lost compute. The three fault-tolerant frameworks all carry tradeoffs: TorchFT (open, but GLOO comms cost >10% perf and whole-replica-group blast radius), AWS HyperPod Checkpointless (recovery in 1m45s vs 15min, but ~5% memory overhead, AWS/Megatron-only), Clockwork TorchPass (zero perf overhead but licensed + idle-node cost). Because none is both free and overhead-free, reliable large-scale training stays concentrated among well-resourced players.
“This leaves fault tolerant training as a secret sauce available to frontier labs or those willing to pay for a software license.”How Much Do GPU Clusters Really Cost? (2026-04-20), SemiAnalysis
14. Selling enterprise tokens is still tiny and hard — even Google's flagship disclosure implies <0.5% of GCP.
Converting tokens to revenue is riddled with errors (I/O ratios, cached tokens, pricing). Sundar Pichai's Q3'25 flex — ~150 Google Cloud customers each processed ~1 trillion tokens over 12 months — actually implies Gemini token sales to those 150 enterprises are less than 0.5% of GCP. The TaaS enterprise business is nascent; the frontier-model access (Claude/GPT/Gemini) that AWS/Azure/Google monopolize matters far more than model-library breadth or headline throughput.
“This disclosure suggests that sales of Gemini tokens to these 150 enterprises is less than 0.5% of GCP's business.”Microsoft's AI Strategy Deconstructed - From Energy to Tokens (2025-11-12), SemiAnalysis
15. GPU-backed debt is the missing financing leg: Nvidia's backstop completes the 'AI Project Trinity' of capital, offtake and datacenters.
With ~$11.1T of cumulative AI capex modeled for 2024-29 and >$7T of AI debt by 2029, financing — not just silicon — gates the buildout. Nvidia's ~6-year backstop pre-agrees GPU-rental price floors and shares upside above them with neoclouds, giving lenders a debt-service-coverage floor that makes GPU-backed debt underwritable. Early deals: SharonAI ($4.88B backstop / 40,000 GB300 / 72MW Australia) and Firmus (360MW, Batam).

Article Deep-Dives

SemiAnalysis's definitive methodology for cluster TCO, turning gut-feel into dollars. It decomposes cost into GPUs + Storage + Networking + Control Plane + Support + Goodput + Setup + Debugging, then introduces the 'Grand Unifying Theory of Goodput' — quantifying downtime as an implicit % tax that scales with job size / cluster MTBF. Across three scenarios (large LLM pretrain, multimodal RL, inference endpoints) with equal $/GPU-hr, Gold-tier is baseline, Hyperscaler runs 1.10-1.61x and Silver 1.15x. It benchmarks three fault-tolerant frameworks (TorchFT, AWS Checkpointless at 1m45s recovery, Clockwork TorchPass) and ships free public TCO/Goodput calculators, plus a ClusterMAX 2.1 tier update.
A full-stack teardown of Microsoft across the AI Token Economic Stack (Apps → Models → PaaS → IaaS → Chips). Narrative arc: 2023-24 all-in (Fairwater, world's largest DCs for OpenAI) → the 'Big Pause' → a 2025 return that finds MSFT out of near-term capacity and forced to rent from neoclouds and resell at thin margin. Key claims: MSFT stepped away from ~$150B of OpenAI gross profit (Oracle won it); Azure risks a ClusterMAX downgrade (weak CycleCloud/AKS for managed clusters); but MSFT still holds 100% of OpenAI API inference compute through 2032 and OpenAI IP rights via Foundry. Also defends extended depreciation against Michael Burry, and dissects GitHub Copilot's eroding moat and the struggling Maia ASIC.
The single most important framing piece: where AI profit pools sit and where they're rotating. 2023-25 infra captured everything; from Dec-2025 agentic AI works and the labs capture it. Anchored by two mechanisms — tokens getting cheaper to produce (GB300 ~17-32x an H100 at only ~70% more TCO; 14x from software alone) and rising market-clearing token value (Opus repriced down to $5/$25 yet margins UP by shifting volume to pricier SKUs like Mythos $25/$125). Argues profits won't be competed away (frontier pricing power + compute scarcity). Introduces 'Nvidia as the central bank of AI' — Nvidia/TSMC deliberately under-pricing scarcity — and the 'One Chart to Rule Them All' GPU-rental value-split framework built on the Vera Rubin VR NVL72 vs GB300 economics.
The clearest case study of business-model-as-margin. Powered by the Tokenomics 2.0 model, it explains why AWS EBIT margin inflected +213bp Q/Q while Oracle, CoreWeave and Azure margins fell. Three levers: (1) Bedrock's token-as-a-service mix (vs 80%+ IaaS at GCP/Azure); (2) the Anthropic distribution deal where Anthropic is seller-of-record but AWS collects infra + revenue-share; (3) vertical silicon (Trainium >50% of Bedrock tokens, industry-leading Graviton4/5 CPUs). It also quantifies Anthropic's blowout ($21B net-new ARR in 1Q26 to $30B; inference GM mid-60s) and frames Google as the ultimate supply-constrained business trying to serve cloud + hardware + models + ads simultaneously.
Ground-truth from 50+ enterprise conversations (Databricks AI Summit, calls, Slack) that de-hypes the 'tokenmaxxing → budget wall' narrative. Findings: the Meta/Uber blowout stories stem from bad incentives not a broad slowdown; even Meta is only a 3-5% Anthropic customer; the 90th+ percentile drives most revenue at little risk. Budgets are now universal but with no convergent number ($250 to tens of thousands/mo). Companies conserve by downgrading defaults (Opus→Sonnet), turning off Fast-Mode, and employees gaming M365 Copilot allowances. Crucially: median Fortune 500 still spends <$100/employee — the adoption s-curve has enormous runway, with cyber and white-collar knowledge work the next verticals after coding.
How a cloud latecomer became an AI-compute powerhouse by blending three archetypes: traditional US hyperscaler + AI-native neocloud + fast-moving Chinese developers. Oracle's edge is structural cost, not software: ~20% cluster CapEx advantage from optimized RoCEv2 networking (Sun/Mellanox HPC heritage), Foxconn ODM sourcing (bypassing Dell/Supermicro margin), and the market's lowest cost of capital via investment-grade debt. The most underappreciated growth engine is ByteDance (making Johor, Malaysia the #2 global AI hub), alongside the ~880MW Abilene 'Stargate' OpenAI campus already on Oracle's books. Validated by Larry's >$130B booking forecast and >100% RPO growth.

Reference: Value Chain

Energy & PowerVistra, GE Vernova, Constellation, SB Energy, on-site turbines — The first bottleneck. In 2024 power names (Vistra +265%, GE Vernova +146%) led the S&P as the market realized watts gate GPU deployment. Securing multi-GW PPAs is now a share-determining act.
Datacenter shells & constructionQTS, Crusoe/Lancium, Vantage, Khazna, self-build hyperscaler teams — MW of pre-leased and self-built capacity is the earliest capex signal. Microsoft's Fairwater sites (~300MW GPU buildings, planned 600MW+) are the largest on Earth; Abilene TX ('Stargate') anchors OpenAI/Oracle.
Accelerators (merchant + custom)Nvidia (GB200/GB300/Rubin), AMD MI300X, AWS Trainium, Google TPU, Marvell/Broadcom ASICs — The cost core: GPU capex dwarfs all hosting cost. Custom ASICs (Trainium, TPU) are the hyperscalers' lever to strip Nvidia margin — Trainium now powers >50% of Bedrock token usage.
Neoclouds & IaaSCoreWeave, Nebius, Crusoe, Fluidstack (Gold); Lambda, Together, Vultr, Voltage Park (Silver) — Rent GPUs by the hour. Low software moat (capital is the barrier) but wide quality spread: Gold-tier's reliability/support makes it cheaper on TCO. Rental prices are surging (+40% off the Oct-2025 H100 bottom).
Hyperscaler PaaS / TaaSAWS Bedrock, Azure Foundry, Google Vertex/Gemini Enterprise — Managed token endpoints. The margin battleground: distributing frontier models (Claude, GPT, Gemini) earns infra + revenue-share. Only AWS, Azure, Google can sell all three frontier families.
Model LabsOpenAI, Anthropic, Google DeepMind, xAI, DeepSeek — From value-destroyers (neg. gross margin in 2024) to value-captors. Frontier pricing power + compute scarcity means margins won't get competed away by open-source. Anthropic ARR $9B→$30-44B, inference GM 38%→70%+.
Applications & Enterprise spendGitHub Copilot, Cursor, M365 Copilot, Claude Code, Codex, enterprise buyers — Where tokens meet ROI. Agentic coding is the first mass vertical (>70% of OpenAI+Anthropic ARR). Enterprises now budget per-employee ($250–$4,000/mo caps) — the s-curve of adoption is early.

Reference: Core Concepts

AI Token Economic Stack. SemiAnalysis's master framework: Energy → Datacenter → Accelerator (GPU/ASIC) → IaaS (bare-metal rental) → PaaS/Token-as-a-Service → Models → Applications. A token is the atomic unit that flows through and lets you convert watts into dollars at every layer. Whoever controls the scarce layer at a given moment captures the margin.

Goodput (vs Throughput). The share of raw GPU throughput that is actually useful work. A cluster can burn full power while producing garbage — GPU fell off the bus, NCCL stalled, an OOM at the next checkpoint, or a job restart eating 10-15 minutes plus all compute since the last checkpoint. 'Goodput expense' is the hidden % tax you pay for a less reliable provider; it dominates large-job TCO.

Cluster TCO ≠ GPU $/hr. True cost = GPUs + Storage + Networking + Control Plane + Support uplift + Goodput expense + Setup expense + Debugging expense, with setup amortized over contract term. Two clouds at the same $/GPU-hr can differ 10-60% in TCO once you count EFA tuning weeks, poor storage tiers, and downtime.

Token-as-a-Service (TaaS) vs IaaS mix. Two ways a hyperscaler sells AI. IaaS = rent bare-metal GPUs on multi-year deals (Azure/GCP heavy). TaaS = sell tokens via a managed endpoint (AWS Bedrock, Azure Foundry, GCP Vertex/Gemini). When you distribute someone else's model (e.g. Claude on Bedrock) you earn both an infra fee AND a revenue-share/distribution fee — the mix that lifts margin.

RPO / Pre-leasing as leading indicators. Remaining Performance Obligation (contracted-but-unbilled backlog) and datacenter pre-leasing MW are SemiAnalysis's earliest, hardest signals of a hyperscaler's true capex intent — visible quarters before the P&L. Microsoft's pre-leasing dwarfed all rivals combined in 2023-24; the collapse in that signal called the 'Big Pause'.

Server depreciation / useful life. Hyperscalers extended assumed useful life from 3-5yr (2020) to 5-6yr, boosting reported earnings. Michael Burry called it artificial; SemiAnalysis disagrees, arguing HPC supercomputers (Summit ran 6.5yr, V100 still rented) prove GPUs stay economically useful far past the 2-3yr product cycle. AWS runs the shortest (5yr) yet still shows rising margins.

Tokenomics: price is an OUTPUT. For any model, $/token is not fixed but solved for by tuning latency (TTFT), interactivity (tok/s/user), context window and batch size against your hardware. This is why DeepSeek's $0.55/$2.19 sticker didn't win share — third-party hosts served the same weights faster; and why 'blended' agentic prices (Opus at ~$0.99/MTok effective vs $5/$25 sticker) look nothing like the list price.

Neocloud / ClusterMAX tiers. 'Neoclouds' are GPU-pure-play clouds (CoreWeave, Nebius, Crusoe, Lambda...). SemiAnalysis ranks all providers with ClusterMAX (Gold/Silver/Bronze) based on hands-on testing of 80+ clouds — networking, storage, reliability, support. Gold-tier commands a real price premium because its lower goodput/setup tax makes it cheaper on TCO even at higher $/hr.

Nvidia/TSMC as 'central bank of AI'. Despite controlling the scarcest layers (GPUs, N3 wafers) at full utilization, both deliberately hold pricing below scarcity-clearing levels. Motive: avoid antitrust attention, keep customers from diversifying, and let downstream labs/neoclouds stay profitable so total demand keeps expanding. They vent value downstream on purpose.

Open Questions

Sources (SemiAnalysis)

Independence & sourcing. This is independent analysis by Yicheng Yang, distilled from publicly accessible SemiAnalysis articles (free posts and free previews; no paywall circumvention) and verified against the underlying text. It is not affiliated with, endorsed by, or a substitute for SemiAnalysis — subscribe there for the full research. All referenced claims are sourced and linked per SemiAnalysis's attribution terms. No SemiAnalysis images are reproduced. Nothing here is investment advice.