Overview
SemiAnalysis's coverage frames the accelerator sector as a systems war, not a chip war. Its recurring thesis — 'systems matter more than microarchitecture' — explains why Nvidia's GB200/GB300 NVL72 rack (72-GPU NVLink scale-up domain) and its 4-million-developer CUDA moat let it keep ~75% gross margins even as Google's TPUv7 Ironwood, AWS Trainium3, and AMD's MI450X close the raw-silicon gap. The firm's edge is empirical: it runs its own benchmarks (InferenceX/InferenceMAX across ~1,000 GPUs, MI300X-vs-H100 training), rates every GPU cloud (ClusterMAX), models the full bill-of-materials and TCO, and tracks GPU rental prices and per-accelerator HBM supply. Three structural facts dominate. First, the CUDA moat is real but is being eroded on the ASIC flank: Gemini 3 was trained entirely on TPUs, Claude Opus 4.5 was trained on multiple hardware types including TPUs (the majority of Anthropic's training and inference infrastructure sits on Google TPUs and Amazon Trainium), Anthropic committed to ~1M TPUs (~$10B of Broadcom racks plus ~$42B of GCP RPO), and OpenAI cut ~30% off its Nvidia fleet cost merely by threatening to buy TPUs. Second, inference has fractured into prefill vs decode phases, spawning specialized silicon (Rubin CPX, the Groq LPU that Nvidia licensed (a non-exclusive inference-tech + talent deal; Groq stayed independent; ~$20B reported)) and disaggregated serving. Third, the 2026 market is supply-constrained: H100 1-year rentals rose ~40% off their October-2025 low, and all rental capacity through ~September 2026 is booked. The winners are those who own the system — the rack, the scale-up fabric, the compiler, and the developer ecosystem — not just the die.
Positioning: Who Wins and Why
Nvidia (NVDA) — Still 'King of the Jungle': ~75% GM, the 72-GPU NVLink rack moat, the CUDA developer flywheel, and it keeps out-innovating (Rubin CPX, $20B Groq LPU, CPO). Even three simultaneous challengers only nibble at the flanks — as long as Jensen keeps accelerating.
Broadcom (AVGO) — Co-designs the TPU (largest BOM item, fat margin) and is central to the merchant-TPU externalization; the biggest beneficiary of the Google/Anthropic ASIC ramp and custom-silicon supercycle.
TSMC — Every leading accelerator (Nvidia, TPU, Trainium, MI450X, LPU4) is on N3-class + CoWoS. N3 capacity is the binding constraint on total industry output — pure toll-taker on the whole war.
CoreWeave (CRWV) — Sole ClusterMAX Platinum, premium pricing power via SUNK; $22.4B OpenAI + $14.2B Meta contracts; +200% in 6 months post-IPO. The reference neocloud.
AMD (AMD) — The credible #2 — but the bet is on the H2-2026 MI450X rack matching VR200 NVL144, plus OpenAI's equity rebate and sweetheart pricing. Software (ROCm) execution is the swing factor; 'wartime mode' is real but Nvidia is still sprinting.
Astera Labs (ALAB) / interconnect suppliers — PCIe retimers/switches into AWS Trainium (with equity 'rebate'); scale-up/scale-out connectivity is the new decisive layer — a rising tide across ASIC and GPU builds.
Crypto-miner-to-AI pivots (IREN, Terawulf, Cipher) — Own scarce power + PPAs; the 'hyperscaler backstop' financing template (Google's off-balance-sheet IOU) unlocks a NeoCloud growth wave; IREN scored a 200MW GB300 deal with Microsoft.
Key Data
| Metric | Value | Note |
|---|---|---|
| Nvidia gross margin / markup | ~75% GM, ~4x markup | The margin umbrella that leaves room to invest in labs rather than cut price; Broadcom takes a chunk on TPU BOM. |
| GB200 NVL72 scale-up world size | 72 GPUs (NVLink) | vs AMD MI355X's 8; the gap that keeps AMD off frontier MoE reasoning inference until MI450X. |
| TPUv7 Ironwood max pod size | 9,216 TPUs (ICI 3D torus) | Built from 144 4x4x4 cubes via 48 144x144 OCSs; ~8,000 is the practical training block due to slice availability. |
| H100 1-year rental price | $1.70→$2.35/hr/GPU (+~40%, Oct-25→Mar-26) | Defied the consensus of a Hopper price collapse; on-demand sold out across all SKUs. |
| Cost to train GPT-3 175B (FP8, H100) | 72¢→54.2¢ per 1M tokens (2024, software gains) | = $216k→~$163k for a 300B-token run, purely from CUDA-stack MFU improvements at ~$1.42/hr/GPU. |
| H100 MFU improvement (software only) | BF16 34%→54%, FP8 29.5%→39.5% in 12 mo | +~59% BF16 throughput from cuDNN/cuBLAS/NCCL kernels alone — the CUDA moat compounding. |
| GB200 NVL72 vs H100 TCO | ~1.6x higher | Must be ≥1.6x faster to win perf-per-TCO; rack-level reliability is the hidden tax. |
| GB300 NVL72 inference uplift vs H100 | up to 100x (FP8→FP4), 65x (FP8→FP8) | InferenceX v2 across ~1,000 GPUs; H100→GB200 shows up to 55x at 75 tok/s/user. |
| Anthropic TPU commitment | ≥1M TPUs (400k Ironwood ≈ $10B racks + ~$42B GCP RPO) | 400k sold directly via Broadcom; 600k rented via GCP — most of GCP's $49B Q3 backlog jump. |
| OpenAI Nvidia-fleet cost cut from TPU threat | ~30% | Achieved before deploying a single TPU — the perf-per-TCO threat alone extracts the discount. |
| Trainium3 spec deltas vs Trn2 | 2x MXFP8 FLOPS; 144GB 12-Hi HBM3E; +70% bandwidth | On N3P; HBM switched from sub-par Samsung to Hynix/Micron (9.6Gbps pins, highest seen). |
| Rubin CPX prefill chip | 20 PFLOPS dense FP4, 128GB GDDR7, 2TB/s | HBM-free, >50% lower memory cost/GB vs R200's 288GB HBM/20.5TB/s — attacks the memory-cost wall. |
| Nvidia–Groq deal | $20B IP license + team hire | Structured to dodge antitrust; LPU3 (LP30) has 500MB SRAM, 1.2 PFLOPS FP8, runs on Samsung SF4. |
| CoreWeave contract wins | $22.4B OpenAI + $14.2B/6yr Meta | IPO'd on NASDAQ ($CRWV), stock +200% in 6 months. Its proposed ~$9B all-stock acquisition of Core Scientific (~1.3GW) was voted down by Core Scientific shareholders on 2025-10-30 and terminated; the two remain a customer relationship, not merged. |
| ClusterMAX top-cloud RPO booked | ~$400B since v1.0 (Mar 2025) | 84 providers reviewed; CoreWeave sole Platinum; AMD cloud offerings rated worse than same firm's Nvidia offering. |
| Astera Labs (ALAB) equity 'rebate' to AWS | ~23% effective discount (strike $20.34) | AWS earns warrants for hitting PCIe-retimer purchase milestones on Trainium3 — a component 'rebate.' |
| Opus 4.6 fast-mode economics | 6x price for ~2.5x (now ~1.75x) interactivity | Revealed preference for fast tokens over smart tokens; ~80% of SemiAnalysis's ~$10M AI spend was on it. |
| Cerebras–OpenAI deal | 750MW compute (by 2028) | WSE-3 wafer-scale chip wins on speed (fast tokens); Cerebras filed to IPO on the strength of it. |
Key Theses
“AMD and custom silicon competitors may have made a small step forward in emulating Nvidia's 72-GPU rack scale design, but Nvidia has just made another Giant Leap, again leaving competitors very distant objects in the rear-view mirror.”— Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack (2025-09-10), SemiAnalysis
“As fast as AMD tries to fill in the CUDA moat, NVIDIA engineers are working overtime to deepen said moat with new features, libraries, and performance updates.”— MI300X vs H100 vs H200 Benchmark Part 1: Training - CUDA Moat Still Alive (2024-12-22), SemiAnalysis
“The more (TPU) you buy, the more (NVIDIA GPU capex) you save! OpenAI hasn't even deployed TPU yet and already increased perf per TCO by getting ~30% off their compute fleet due to competitive threats.”— TPUv7: Google Takes a Swing at the King (2025-11-28), SemiAnalysis
“AMD needs to invest significantly more GPUs, they have less than 1/20th of Nvidia's total GPU count.”— AMD 2.0 - New Sense of Urgency | MI450X Chance to Beat Nvidia | Nvidia's New Moat (2025-04-23), SemiAnalysis
“Only with hardware specialized to the very different phases of inference, prefill and decode, can disaggregated serving achieve its full potential.”— Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack (2025-09-10), SemiAnalysis
“Rack scale Blackwell NVL72 is framemogging hopper and makes hopper looks like it is jestermaxxing.”— InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper (2026-02-16), SemiAnalysis
“TCO for the GB200 NVL72 is about 1.6x higher than TCO for the H100. This means that the GB200 NVL72 needs to be at least 1.6x faster than the H100 in order to have an performance per TCO advantage.”— H100 vs GB200 NVL72 Training Benchmarks (2025-08-20), SemiAnalysis
“Trying to find GPU compute in early 2026 has been like trying to book airplane tickets on the last flight out, high prices, and almost no availability.”— The Great GPU Shortage – Rental Capacity (2026-04-02), SemiAnalysis
“Nvidia paid Groq $20B to license their IP and hire most the team… if this transaction were structured as a full acquisition and were put to anti-trust review, such a transaction would likely not go through.”— Nvidia – The Inference Kingdom Expands (2026-03-24), SemiAnalysis
“CoreWeave retains top spot as the only member of the Platinum tier… the only cloud to consistently command premium pricing in our interviews with end users.”— ClusterMAX™ 2.0: The Industry Standard GPU Cloud Rating System (2025-11-06), SemiAnalysis
“Nvidia aims to protect its dominant position at the foundation labs by offering equity investment rather than cutting prices, which would lower Gross margins and cause widespread investor panic.”— TPUv7: Google Takes a Swing at the King (2025-11-28), SemiAnalysis
Article Deep-Dives
Reference: Value Chain
Reference: Core Concepts
Scale-up world size (NVLink / ICI / NeuronLink / UALink). The number of accelerators participating in one high-bandwidth scale-up (collective) domain. It is the single most decisive competitive metric today. Nvidia's GB200 NVL72 = 72 GPUs; TPUv7 Ironwood scales via ICI to a 9,216-chip 3D-torus pod; AMD's MI355X is still stuck at 8, which is why SemiAnalysis says it cannot compete head-to-head with GB200 NVL72 on frontier MoE reasoning inference.
Prefill vs Decode (disaggregated serving). LLM inference has two phases with opposite hardware needs. Prefill processes the whole prompt in parallel — compute-bound, hungry for FLOPS. Decode emits one token at a time, reloading the KV cache from HBM each step — memory-bandwidth-bound. Running them on the same GPU means prefill batches constantly disrupt decode. 'Disagg' splits them onto separate GPU pools, each tuned independently — the technique used in production at OpenAI, Anthropic, xAI, DeepSeek.
Performance per TCO. SemiAnalysis's core yardstick: not FLOPS or price/chip, but realized throughput divided by all-in total cost of ownership (capex + power + reliability downtime + engineering time). It is why on-paper specs mislead: the GB200 NVL72 has ~1.6x the TCO of an H100 cluster, so it must be ≥1.6x faster just to break even on perf-per-TCO. 'The hardware you cannot use has infinite TCO.'
CUDA moat. Nvidia's durable advantage isn't the chip — it's the ~4 million external developers who write kernels, file bugs, and ship day-one CUDA implementations (FlashAttention, Mamba, vLLM all launched CUDA-first). SemiAnalysis's sharp framing: the moat isn't dug by Nvidia's engineers but by the millions of outside developers. AMD's RCCL/ROCm libraries are largely forks of Nvidia's, so every NCCL refactor forces AMD to burn engineering hours just to keep pace.
Systolic array / Tensor Core / MXU. The matrix-multiply engine at the heart of every AI chip. Nvidia's Tensor Core (introduced in Volta 2017) does a fused matrix-multiply-accumulate to amortize instruction overhead — a plain FP op costs ~1.5pJ but issuing the instruction costs ~30pJ, a 20x overhead. Google's TPU uses a systolic array; for Trillium (TPUv6) Google quadrupled the array from 128x128 to 256x256 tiles, doubling FLOPS on the same N5 node.
Number formats: FP8, FP4, MXFP, NVFP4. Lower-precision formats trade numerical range/mantissa for throughput and memory. Each generation adds a narrower format: Blackwell added FP4, which roughly doubles throughput vs FP8. Microscaling formats (MXFP8/MXFP4, and Nvidia's NVFP4) attach a shared scale to blocks of values to preserve accuracy. InferenceX shows GB300 delivering up to 100x on FP8-vs-FP4 relative to an H100 baseline — most of Blackwell's inference gain is the format, not just the silicon.
Optical Circuit Switch (OCS) scale-up. Google's alternative to electrical packet switches. An OCS is a reconfigurable mirror-based 'patch panel' that routes whole optical fibers with near-zero latency and no optical-electrical-optical conversion, making it lower-power. It lets Google carve arbitrary 3D-torus slices out of a 9,216-TPU pod, route around faults, and give complete cube 'fungibility.' Building the max world size needs 48 144x144 OCSs. This is the backbone of Nvidia's only real scale-up rival.
Neocloud / ClusterMAX. Neoclouds are GPU-focused clouds (CoreWeave, Nebius, Crusoe, Lambda) that rent GPUs to labs. ClusterMAX is SemiAnalysis's independent rating system (Platinum→Underperforming) grading 84 providers on 10 criteria (security, orchestration, storage, networking, reliability, monitoring, pricing…). CoreWeave is the sole Platinum. The rating has teeth: top-rated neoclouds have booked ~$400B in RPO since v1.0.
HBM as % of BOM / the memory wall. High-Bandwidth Memory has gone from 80GB/3.4TB/s (H100) to 288GB/8TB/s (GB300) in under three years and is now the single largest component of the GB300 package BOM. Because prefill wastes HBM's expensive bandwidth, Nvidia built Rubin CPX with cheap GDDR7 (>50% lower cost/GB) instead — a direct attack on the memory-cost wall.
Open Questions
- Will Google open-source XLA:TPU, the runtime and MegaScaler? SemiAnalysis argues it's the single missing ingredient for TPU to threaten the CUDA moat at scale — but Google has resisted for years.
- Can AMD's MI450X rack actually match VR200 NVL144 in H2-2026, and — more importantly — will ROCm software composability be there? Hardware parity without software parity has failed before.
- Is the ~40% H100 rental surge structural (insatiable inference/agentic demand) or a memory-price-driven supply squeeze that reverses once OEM server pricing normalizes?
- Does prefill-decode specialization (Rubin CPX, LPU) become the dominant inference architecture, forcing every competitor to build two chip families — and can anyone match Nvidia's system integration to exploit it?
- How durable is the 'hyperscaler backstop' financing template that papers over the 4-5yr-GPU vs 15yr-lease duration mismatch if a frontier lab's demand or funding falters?
Sources (SemiAnalysis)
- ClusterMAX™ 2.0: The Industry Standard GPU Cloud Rating System (2025-11-06)
- The GPU Cloud ClusterMAX™ Rating System | How to Rent GPUs (2025-03-26)
- AWS Trainium3 Deep Dive | A Potential Challenger Approaching (2025-12-04)
- AI Neocloud Playbook and Anatomy (2024-10-03)
- AMD 2.0 - New Sense of Urgency | MI450X Chance to Beat Nvidia | Nvidia's New Moat (2025-04-23)
- MI300X vs H100 vs H200 Benchmark Part 1: Training - CUDA Moat Still Alive (2024-12-22)
- InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper (2026-02-16)
- TPUv7: Google Takes a Swing at the King (2025-11-28)
- InferenceMAX™: Open Source Inference Benchmarking (2025-10-09)
- Cerebras — Faster Tokens Please (2026-05-13)
- AMD Advancing AI: MI350X and MI400 UALoE72, MI500 UAL256 (2025-06-13)
- AMD vs NVIDIA Inference Benchmark: Who Wins? - Performance & Cost Per Million Tokens (2025-05-23)
- Nvidia – The Inference Kingdom Expands (2026-03-24)
- Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack (2025-09-10)
- H100 vs GB200 NVL72 Training Benchmarks - Power, TCO, and Reliability Analysis (2025-08-20)
- NVIDIA Tensor Core Evolution: From Volta To Blackwell (2025-06-23)
- The Great GPU Shortage – Rental Capacity (2026-04-02)
- NVIDIA GTC 2025 - Built For Reasoning, Vera Rubin, Kyber, CPO, Dynamo Inference (2025-03-19)
- AI Dark Output: The Visible Cost of Invisible Output (2026-05-29)
- Nvidia's Christmas Present: GB300 & B300 - Reasoning Inference, Amazon, Memory (2024-12-25)