AI Supply Chain Research · Sector 01

AI Accelerators (GPU / ASIC)

Nvidia is still King of the Jungle, but for the first time three credible challengers — Google TPU, AWS Trainium, and AMD MI450X — have arrived at once, and the real battlefield has moved from chip specs to rack-scale systems, software, and performance-per-TCO.  ·  ← back to the series  ·  Background: Primer §00, §01, §02, §03, §04, §05, §06, §07, §08, §09, §10

At a glance
Winners
Nvidia (NVDA)Broadcom (AVGO)TSMCCoreWeave (CRWV)AMD (AMD)Astera Labs (ALAB) / interconnect suppliersCrypto-miner-to-AI pivots (IREN, Terawulf, Cipher)
Bottlenecks
TSMC N3 wafer capacity — caps total accelerator output; N3P defect density is improving slower than expected, forcing re-spins/queues.HBM supply — the largest single BOM item on GB300; constrained across Hynix/Micron/Samsung and repriced parabolically (DDR5 ~5x YoY).Power / datacenter capacity — Google's true bottleneck is power (3-yr MSA cycles); all rental capacity through ~Sept-2026 is booked.Advanced packaging (CoWoS-R/-L) — the reticle-scale interposers and 20-layer ABF substrates that gate chiplet + HBM integration.Software talent — AMD is uncompetitive on AI-software-engineer comp; ROCm/RCCL forever chasing NCCL refactors.
Risks
ASIC substitution accelerates: if Google open-sources XLA:TPU/runtime and AWS open-sources NKI, the CUDA moat erodes on the software flank faster than expected.GB200/GB300 rack reliability: NVLink copper backplane + cross-rack ACC cables have persistent signal-integrity failures; the 72-GPU failure domain raises the hidden TCO tax.GPU terminal-value / rental-price reversal: the ~40% H100 rental rally could unwind if inference demand growth slows, hitting neocloud IRRs and 6-yr depreciation assumptions.'Circular economy' financing: Nvidia/hyperscaler equity investments and off-balance-sheet backstops concentrate counterparty risk if a marquee lab's demand disappoints.Duration mismatch: 4-5yr GPU life vs 15-yr datacenter leases (~8yr payback) — solved for now by the hyperscaler backstop, but a fragile structure.
Catalysts
H2-2026 rack-scale showdown: AMD MI450X (IF64/IF128) and Nvidia VR200 NVL144 both hit production — the first apples-to-apples 72+ rack comparison.TPUv7 Ironwood + Trainium3 added to InferenceX — the first independent, apples-to-apples ASIC-vs-GPU inference numbers.More TPU externalization deals (Meta, xAI, SSI, OpenAI) + Cerebras IPO — re-rating the ASIC/merchant-challenger supply chain.Rubin CPX / Vera Rubin ramp — validates prefill-decode specialization and disaggregated serving at scale.GPU rental index inflection — SemiAnalysis's public H100 1yr index is the real-time barometer for the whole cycle.

Overview

SemiAnalysis's coverage frames the accelerator sector as a systems war, not a chip war. Its recurring thesis — 'systems matter more than microarchitecture' — explains why Nvidia's GB200/GB300 NVL72 rack (72-GPU NVLink scale-up domain) and its 4-million-developer CUDA moat let it keep ~75% gross margins even as Google's TPUv7 Ironwood, AWS Trainium3, and AMD's MI450X close the raw-silicon gap. The firm's edge is empirical: it runs its own benchmarks (InferenceX/InferenceMAX across ~1,000 GPUs, MI300X-vs-H100 training), rates every GPU cloud (ClusterMAX), models the full bill-of-materials and TCO, and tracks GPU rental prices and per-accelerator HBM supply. Three structural facts dominate. First, the CUDA moat is real but is being eroded on the ASIC flank: Gemini 3 was trained entirely on TPUs, Claude Opus 4.5 was trained on multiple hardware types including TPUs (the majority of Anthropic's training and inference infrastructure sits on Google TPUs and Amazon Trainium), Anthropic committed to ~1M TPUs (~$10B of Broadcom racks plus ~$42B of GCP RPO), and OpenAI cut ~30% off its Nvidia fleet cost merely by threatening to buy TPUs. Second, inference has fractured into prefill vs decode phases, spawning specialized silicon (Rubin CPX, the Groq LPU that Nvidia licensed (a non-exclusive inference-tech + talent deal; Groq stayed independent; ~$20B reported)) and disaggregated serving. Third, the 2026 market is supply-constrained: H100 1-year rentals rose ~40% off their October-2025 low, and all rental capacity through ~September 2026 is booked. The winners are those who own the system — the rack, the scale-up fabric, the compiler, and the developer ecosystem — not just the die.

Positioning: Who Wins and Why

Nvidia (NVDA) — Still 'King of the Jungle': ~75% GM, the 72-GPU NVLink rack moat, the CUDA developer flywheel, and it keeps out-innovating (Rubin CPX, $20B Groq LPU, CPO). Even three simultaneous challengers only nibble at the flanks — as long as Jensen keeps accelerating.

Broadcom (AVGO) — Co-designs the TPU (largest BOM item, fat margin) and is central to the merchant-TPU externalization; the biggest beneficiary of the Google/Anthropic ASIC ramp and custom-silicon supercycle.

TSMC — Every leading accelerator (Nvidia, TPU, Trainium, MI450X, LPU4) is on N3-class + CoWoS. N3 capacity is the binding constraint on total industry output — pure toll-taker on the whole war.

CoreWeave (CRWV) — Sole ClusterMAX Platinum, premium pricing power via SUNK; $22.4B OpenAI + $14.2B Meta contracts; +200% in 6 months post-IPO. The reference neocloud.

AMD (AMD) — The credible #2 — but the bet is on the H2-2026 MI450X rack matching VR200 NVL144, plus OpenAI's equity rebate and sweetheart pricing. Software (ROCm) execution is the swing factor; 'wartime mode' is real but Nvidia is still sprinting.

Astera Labs (ALAB) / interconnect suppliers — PCIe retimers/switches into AWS Trainium (with equity 'rebate'); scale-up/scale-out connectivity is the new decisive layer — a rising tide across ASIC and GPU builds.

Crypto-miner-to-AI pivots (IREN, Terawulf, Cipher) — Own scarce power + PPAs; the 'hyperscaler backstop' financing template (Google's off-balance-sheet IOU) unlocks a NeoCloud growth wave; IREN scored a 200MW GB300 deal with Microsoft.

Key Data

MetricValueNote
Nvidia gross margin / markup~75% GM, ~4x markupThe margin umbrella that leaves room to invest in labs rather than cut price; Broadcom takes a chunk on TPU BOM.
GB200 NVL72 scale-up world size72 GPUs (NVLink)vs AMD MI355X's 8; the gap that keeps AMD off frontier MoE reasoning inference until MI450X.
TPUv7 Ironwood max pod size9,216 TPUs (ICI 3D torus)Built from 144 4x4x4 cubes via 48 144x144 OCSs; ~8,000 is the practical training block due to slice availability.
H100 1-year rental price$1.70→$2.35/hr/GPU (+~40%, Oct-25→Mar-26)Defied the consensus of a Hopper price collapse; on-demand sold out across all SKUs.
Cost to train GPT-3 175B (FP8, H100)72¢→54.2¢ per 1M tokens (2024, software gains)= $216k→~$163k for a 300B-token run, purely from CUDA-stack MFU improvements at ~$1.42/hr/GPU.
H100 MFU improvement (software only)BF16 34%→54%, FP8 29.5%→39.5% in 12 mo+~59% BF16 throughput from cuDNN/cuBLAS/NCCL kernels alone — the CUDA moat compounding.
GB200 NVL72 vs H100 TCO~1.6x higherMust be ≥1.6x faster to win perf-per-TCO; rack-level reliability is the hidden tax.
GB300 NVL72 inference uplift vs H100up to 100x (FP8→FP4), 65x (FP8→FP8)InferenceX v2 across ~1,000 GPUs; H100→GB200 shows up to 55x at 75 tok/s/user.
Anthropic TPU commitment≥1M TPUs (400k Ironwood ≈ $10B racks + ~$42B GCP RPO)400k sold directly via Broadcom; 600k rented via GCP — most of GCP's $49B Q3 backlog jump.
OpenAI Nvidia-fleet cost cut from TPU threat~30%Achieved before deploying a single TPU — the perf-per-TCO threat alone extracts the discount.
Trainium3 spec deltas vs Trn22x MXFP8 FLOPS; 144GB 12-Hi HBM3E; +70% bandwidthOn N3P; HBM switched from sub-par Samsung to Hynix/Micron (9.6Gbps pins, highest seen).
Rubin CPX prefill chip20 PFLOPS dense FP4, 128GB GDDR7, 2TB/sHBM-free, >50% lower memory cost/GB vs R200's 288GB HBM/20.5TB/s — attacks the memory-cost wall.
Nvidia–Groq deal$20B IP license + team hireStructured to dodge antitrust; LPU3 (LP30) has 500MB SRAM, 1.2 PFLOPS FP8, runs on Samsung SF4.
CoreWeave contract wins$22.4B OpenAI + $14.2B/6yr MetaIPO'd on NASDAQ ($CRWV), stock +200% in 6 months. Its proposed ~$9B all-stock acquisition of Core Scientific (~1.3GW) was voted down by Core Scientific shareholders on 2025-10-30 and terminated; the two remain a customer relationship, not merged.
ClusterMAX top-cloud RPO booked~$400B since v1.0 (Mar 2025)84 providers reviewed; CoreWeave sole Platinum; AMD cloud offerings rated worse than same firm's Nvidia offering.
Astera Labs (ALAB) equity 'rebate' to AWS~23% effective discount (strike $20.34)AWS earns warrants for hitting PCIe-retimer purchase milestones on Trainium3 — a component 'rebate.'
Opus 4.6 fast-mode economics6x price for ~2.5x (now ~1.75x) interactivityRevealed preference for fast tokens over smart tokens; ~80% of SemiAnalysis's ~$10M AI spend was on it.
Cerebras–OpenAI deal750MW compute (by 2028)WSE-3 wafer-scale chip wins on speed (fast tokens); Cerebras filed to IPO on the strength of it.

Key Theses

1. Systems, not chips, decide the war — and Nvidia's GB200/GB300 NVL72 rack opened a canyon-sized gap in scale-up.
SemiAnalysis's founding thesis is that a strong system beats a strong die: even when TPU silicon lagged Nvidia on paper, Google's system engineering matched Nvidia on perf and cost. The GB200 NVL72's 72-GPU NVLink domain is why AMD's 8-GPU MI355X 'cannot compete head on' on frontier MoE reasoning; competitors only reach a 72-class rack a year later (MI450X, Trainium3).
“AMD and custom silicon competitors may have made a small step forward in emulating Nvidia's 72-GPU rack scale design, but Nvidia has just made another Giant Leap, again leaving competitors very distant objects in the rear-view mirror.”Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack (2025-09-10), SemiAnalysis
2. The CUDA moat is real and still widening on training — AMD's software, not its hardware, is the problem.
After a five-month benchmarking quest, SemiAnalysis found the MI300X's on-paper spec advantage evaporated: out-of-box AMD PyTorch training was 'impossible' without workarounds, and public-release performance trailed H100/H200 despite lower TCO. Because RCCL is a fork of NCCL, every Nvidia refactor forces AMD to burn thousands of engineer-hours just to sync. The takeaway: the moat isn't the chip — it's the ~4M external developers.
“As fast as AMD tries to fill in the CUDA moat, NVIDIA engineers are working overtime to deepen said moat with new features, libraries, and performance updates.”MI300X vs H100 vs H200 Benchmark Part 1: Training - CUDA Moat Still Alive (2024-12-22), SemiAnalysis
3. Google's TPU is the most dangerous merchant challenger — and you get its benefit even before turning one on.
Gemini 3 was trained entirely on TPUs; Claude Opus 4.5 was trained on multiple hardware types including TPUs — the majority of Anthropic's training/inference infrastructure runs on Google TPUs and Amazon Trainium. Anthropic committed to ≥1M TPUs (400k Ironwoods = ~$10B of Broadcom racks + 600k via ~$42B of GCP RPO). Meta, xAI, SSI and OpenAI are in the pipeline. Crucially, the mere competitive threat lets labs extract price cuts: 'OpenAI hasn't even deployed TPUs yet and they've already saved ~30% on their entire lab-wide NVIDIA fleet.'
“The more (TPU) you buy, the more (NVIDIA GPU capex) you save! OpenAI hasn't even deployed TPU yet and already increased perf per TCO by getting ~30% off their compute fleet due to competitive threats.”TPUv7: Google Takes a Swing at the King (2025-11-28), SemiAnalysis
4. 'Systems matter more than microarchitecture' — but TPU silicon has now nearly closed the gap too.
Historically Google under-specced TPU FLOPS/memory versus Nvidia (prioritizing RAS/uptime, and because RecSys had low arithmetic intensity). That changed post-LLM: Trillium (TPUv6, same N5 node) doubled FLOPS by quadrupling the systolic array to 256x256; TPUv7 Ironwood nearly matches GB200 on FLOPS, bandwidth and 8-Hi HBM3E capacity — arriving only ~a year after Blackwell.
5. AWS Trainium3 is a credible third front, built on an 'Amazon Basics' perf-per-TCO obsession.
Trainium3 moves to N3P, doubles MXFP8 FLOPS, upgrades to 12-Hi HBM3E (144GB) and — by switching from sub-par Samsung to Hynix/Micron HBM — lifts bandwidth 70%. It adds an all-to-all switched scale-up fabric (Nvidia Oberon-like), making AWS the first outside Nvidia to ship a 72-class switched rack. AWS deliberately drives suppliers hard: 'Annapurna places greater emphasis on driving down the TCO denominator.'
6. AMD is finally in 'wartime mode'; its real shot at Nvidia is the H2-2026 MI450X rack, not the MI355X.
After the December-2024 article, Lisa Su personally engaged; AMD stood up devrel (Anush Elangovan), added MI300 to PyTorch CI/CD, and is launching a dev cloud. But the MI355X's 8-GPU world size relegates it to competing with air-cooled HGX B200/B300, not GB200 NVL72. SemiAnalysis's window: the rack-scale MI450X IF64/IF128 (UALink/Infinity-Fabric-over-Ethernet) could match Nvidia's VR200 NVL144 in H2 2026 — with 'sweetheart pricing' + OpenAI's equity rebate as sweeteners.
“AMD needs to invest significantly more GPUs, they have less than 1/20th of Nvidia's total GPU count.”AMD 2.0 - New Sense of Urgency | MI450X Chance to Beat Nvidia | Nvidia's New Moat (2025-04-23), SemiAnalysis
7. Inference has split into prefill and decode — spawning specialized silicon like Rubin CPX and the Groq LPU.
Because prefill is compute-bound and decode is bandwidth-bound, running both on one HBM-rich GPU wastes expensive memory. Rubin CPX answers with 20 PFLOPS dense FP4 but only 128GB of cheap GDDR7 (2TB/s) — HBM-free, >50% lower memory cost/GB. SemiAnalysis calls this second only to the GB200 NVL72 in significance, and says it sends every competitor 'back to the drawing board' to build their own prefill chip.
“Only with hardware specialized to the very different phases of inference, prefill and decode, can disaggregated serving achieve its full potential.”Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack (2025-09-10), SemiAnalysis
8. On real inference, rack-scale Blackwell 'framemogs' Hopper — and Nvidia dominates energy efficiency.
InferenceX v2 (formerly InferenceMAX) benchmarks ~1,000 GPUs across all SKUs from the past 4 years. GB300 NVL72 hits up to 100x FP8-vs-FP4 and 65x FP8-vs-FP8 over an H100 disagg+wideEP baseline; H100→GB200 NVL72 shows up to 55x at 75 tok/s/user. Jensen 'under-promised and over-delivered' on his GTC-2024 30x claim. AMD is competitive on subsets, but its weakness is composability — combining disagg+wideEP+FP4 breaks down.
“Rack scale Blackwell NVL72 is framemogging hopper and makes hopper looks like it is jestermaxxing.”InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper (2026-02-16), SemiAnalysis
9. GB200 NVL72 costs ~1.6x an H100 cluster's TCO — and its rack-scale reliability is the hidden tax.
Factoring capex + opex, GB200 NVL72 TCO is ~1.6x an H100 cluster, so it must be ≥1.6x faster to win on perf-per-TCO. The catch: with Blackwell the failure domain jumped from the node to the entire 72-GPU rack — a single faulty component can drain a whole rack. NVLink copper backplane and cross-rack ACC cables have persistent signal-integrity/reliability issues; top providers now offer 99% rack-level SLAs. Software also matters enormously: H100 MFU rose from 34%→54% in 12 months purely from CUDA-stack improvements.
“TCO for the GB200 NVL72 is about 1.6x higher than TCO for the H100. This means that the GB200 NVL72 needs to be at least 1.6x faster than the H100 in order to have an performance per TCO advantage.”H100 vs GB200 NVL72 Training Benchmarks (2025-08-20), SemiAnalysis
10. The 2026 GPU market is acutely supply-constrained: H100 rentals rose ~40% and everything through ~Sept-2026 is booked.
Contrary to the consensus that Hopper would collapse in price as Blackwell ramped, H100 1-year rental prices rose ~40% from a $1.70/hr low in Oct 2025 to $2.35/hr by March 2026. On-demand is sold out across every GPU type; AWS p6-b200 spot fetches $14/hr; all capacity online through Aug–Sept 2026 is booked. Drivers: open-weight model adoption, agentic/multi-step workloads (Claude Code), native media generation, and a memory-price spike (DDR5 ~5x YoY) that made OEMs reprice servers and withhold supply.
“Trying to find GPU compute in early 2026 has been like trying to book airplane tickets on the last flight out, high prices, and almost no availability.”The Great GPU Shortage – Rental Capacity (2026-04-02), SemiAnalysis
11. Nvidia's $20B Groq deal is really about building an SRAM 'fast-token' front for disaggregated decode.
Nvidia paid Groq $20B to license IP and hire the team (structured to dodge antitrust). Standalone, a Groq LPU is uneconomic at scale — its ~500MB SRAM saturates fast and leaves little for KV cache — but it serves tokens extremely fast, commanding a premium. Nvidia will fuse the LPU into Vera Rubin via Attention-FFN Disaggregation (AFD): FFN on the SRAM-heavy LPU, memory-hungry attention on the HBM-rich GPUs. Bonus: the LPU runs on Samsung SF4, sidestepping constrained TSMC N3/HBM — 'true incremental revenue and capacity that no one else can access.'
“Nvidia paid Groq $20B to license their IP and hire most the team… if this transaction were structured as a full acquisition and were put to anti-trust review, such a transaction would likely not go through.”Nvidia – The Inference Kingdom Expands (2026-03-24), SemiAnalysis
12. CoreWeave is the only 'Platinum' GPU cloud — and ClusterMAX ratings now steer hundreds of billions in compute contracts.
ClusterMAX 2.0 reviewed 84 providers (from 26); CoreWeave alone holds Platinum, the only cloud that can consistently command a pricing premium because its all-in TCO is better even at a higher $/GPU-hr. Its differentiator SUNK (Slurm-on-Kubernetes) is the only viable way to run Slurm and Kubernetes jobs on one cluster. CoreWeave signed $22.4B with OpenAI and a $14.2B/6-year Meta deal; ClusterMAX's top-rated clouds have collectively booked ~$400B in RPO.
“CoreWeave retains top spot as the only member of the Platinum tier… the only cloud to consistently command premium pricing in our interviews with end users.”ClusterMAX™ 2.0: The Industry Standard GPU Cloud Rating System (2025-11-06), SemiAnalysis
13. Nvidia's investments in AI startups are a margin-defense strategy, not a fake 'circular economy.'
Skeptics claim Nvidia props up cash-burning startups to book revenue. SemiAnalysis disagrees: Nvidia offers equity investment rather than cutting prices, because a price cut would compress its ~75% gross margin and trigger investor panic. It's how Nvidia keeps foundation labs on GPUs while labs use the TPU/AMD threat to negotiate. The competitive pressure is real — hence Nvidia's own defensive PR and finance-team responses reproduced by SemiAnalysis.
“Nvidia aims to protect its dominant position at the foundation labs by offering equity investment rather than cutting prices, which would lower Gross margins and cause widespread investor panic.”TPUv7: Google Takes a Swing at the King (2025-11-28), SemiAnalysis

Article Deep-Dives

The definitive 46,000-word GPU-cloud rating: 84 providers hands-on tested against 10 criteria, 209 tracked, 140+ end users interviewed. CoreWeave is the sole Platinum (its SUNK Slurm-on-Kubernetes and premium pricing power stand out); Nebius/Oracle/Azure lead Gold; Google/AWS lead Silver. Key trends: the shift to Slurm-on-Kubernetes, GB200 NVL72 rack-level failure domains forcing 99% rack-level SLAs, crypto-miners (IREN, Terawulf) pivoting to AI hosting, and — pointedly — that any given firm's AMD cloud offering is materially worse than its Nvidia one. Rating carries commercial weight: top clouds have booked ~$400B in RPO.
The 10K-word case that Google is Nvidia's most dangerous merchant challenger. Gemini 3 trained entirely on TPUs; Opus 4.5 trained on multiple hardware types including TPUs (most of Anthropic's training/inference infra sits on TPUs and Trainium); Anthropic committed ≥1M TPUs; Meta, xAI, SSI, OpenAI are lining up. Ironwood nearly closes the silicon gap to GB200 (FLOPS/bandwidth/HBM), and Google's real edge is the ICI 3D-torus + Optical Circuit Switch scale-up fabric — the only true NVLink rival — reaching a 9,216-TPU pod with reconfigurable, fungible slices. Google earns superior EBIT margins even after Broadcom's cut. The missing ingredient vs the CUDA moat: Google still won't open-source XLA:TPU, the runtime, and the MegaScaler multi-pod code.
SemiAnalysis's most detailed accelerator teardown, covering silicon (N3P, CoWoS-R dual-die), the new all-to-all switched scale-up fabric (making AWS the first outside Nvidia to ship a 72-class switched rack), a 'don't decide' multi-vendor switch strategy, and software (open-sourcing NKI + a native PyTorch backend to seed a developer ecosystem). Trainium3 doubles MXFP8 FLOPS, moves to 144GB 12-Hi HBM3E with +70% bandwidth (Samsung→Hynix/Micron), and is architected for perf-per-TCO with cableless serviceability and redundant scale-up lanes. Alchip beat Marvell for the backend; Astera Labs gives AWS an equity 'rebate' (~23%) on PCIe retimers. Trainium4 splits into UALink and NVLink-fusion tracks.
The open-source (Apache 2.0) benchmark run across ~1,000 GPUs — every western Nvidia SKU of the past 4 years and every AMD SKU of the past 3 — measuring the full throughput-vs-interactivity Pareto frontier with production techniques (disagg prefill + wideEP + FP4/FP8). Results: rack-scale Blackwell dominates (GB300 up to 100x FP8→FP4 over H100), and Nvidia leads energy-per-token everywhere. AMD is competitive on subsets and doubled its DeepSeek-R1 FP4 throughput in <2 months, but its Achilles' heel is composability — stacking disagg+wideEP+FP4 breaks. Widely reproduced/validated by Google Cloud, Azure, Oracle, OpenAI.
The five-month training-benchmark quest that reset the AMD narrative and triggered Lisa Su's personal response. Despite the MI300X's superior on-paper memory and lower TCO, out-of-box AMD PyTorch training was broken and needed workarounds; public-release performance trailed H100/H200. Root causes: weak QA culture, forked libraries (RCCL from NCCL), too few internal GPUs for CI/CD, and poor scale-out. SemiAnalysis open-sourced its benchmarks and issued a detailed roadmap to AMD leadership — the article that arguably kick-started AMD's 'wartime mode.'

Reference: Value Chain

Merchant GPU / accelerator designNvidia, AMD, Intel (Gaudi), Cerebras, Groq (→Nvidia), Tenstorrent — Design the silicon + increasingly the full rack. Nvidia leads with ~75% gross margin, ~4x markup; AMD is the #2 GPU; Cerebras/Groq are specialized inference challengers.
Hyperscaler custom ASIC (in-house silicon)Google TPU, AWS Trainium/Inferentia, Meta MTIA, Microsoft Maia — Vertically-integrated accelerators to escape Nvidia's margin. Google TPU + AWS Trainium are the only ASICs that have trained/served frontier models at scale; Microsoft's Maia program is 'struggling.'
ASIC design partners / IPBroadcom, Marvell, Alchip, Synopsys, Cadence, Annapurna (AWS in-house) — Co-design hyperscaler ASICs and supply SerDes/interface IP. Broadcom earns fat margins co-designing the TPU (largest BOM item); Alchip beat Marvell for the Trainium3 backend socket.
Scale-up / scale-out interconnectNvidia NVLink/NVSwitch, Google ICI+OCS, AMD/UALink consortium, Broadcom/Arista Ethernet, Astera Labs (retimers/PCIe switches) — The fabric that binds chips into a rack and racks into a pod — now the decisive moat. Astera Labs (ALAB) supplies PCIe retimers/switches to AWS Trainium and gets equity 'rebates.'
Foundry + advanced packaging + HBMTSMC (N3P/N5, CoWoS-R/-L), SK Hynix, Micron, Samsung (HBM) — The physical bottleneck. TSMC N3 capacity and HBM supply cap total accelerator output; Trainium3 moves to N3P (defect-density issues), and AWS switched HBM from sub-par Samsung to Hynix/Micron for a 70% bandwidth gain.
GPU clouds (Neoclouds) + hyperscalersCoreWeave, Nebius, Oracle, Crusoe, Lambda, Fluidstack, Nscale; AWS/Azure/GCP; crypto-miner pivots (IREN, Terawulf, Cipher) — Buy accelerators and rent them to labs. CoreWeave alone signed $22.4B with OpenAI + $14.2B with Meta; ex-Bitcoin miners are pivoting into AI hosting via 'hyperscaler backstop' financing.
Software / compiler / inference stackCUDA/cuDNN/NCCL/TensorRT-LLM/Dynamo (Nvidia), ROCm/RCCL (AMD), XLA/JAX (Google), NKI/Neuron (AWS); vLLM, SGLang, PyTorch — Where the moat lives. Open engines vLLM & SGLang are CUDA-first; challengers are racing to open-source their compilers (AWS open-sourcing NKI; Google pressured to open XLA:TPU) to seed an external developer ecosystem.

Reference: Core Concepts

Scale-up world size (NVLink / ICI / NeuronLink / UALink). The number of accelerators participating in one high-bandwidth scale-up (collective) domain. It is the single most decisive competitive metric today. Nvidia's GB200 NVL72 = 72 GPUs; TPUv7 Ironwood scales via ICI to a 9,216-chip 3D-torus pod; AMD's MI355X is still stuck at 8, which is why SemiAnalysis says it cannot compete head-to-head with GB200 NVL72 on frontier MoE reasoning inference.

Prefill vs Decode (disaggregated serving). LLM inference has two phases with opposite hardware needs. Prefill processes the whole prompt in parallel — compute-bound, hungry for FLOPS. Decode emits one token at a time, reloading the KV cache from HBM each step — memory-bandwidth-bound. Running them on the same GPU means prefill batches constantly disrupt decode. 'Disagg' splits them onto separate GPU pools, each tuned independently — the technique used in production at OpenAI, Anthropic, xAI, DeepSeek.

Performance per TCO. SemiAnalysis's core yardstick: not FLOPS or price/chip, but realized throughput divided by all-in total cost of ownership (capex + power + reliability downtime + engineering time). It is why on-paper specs mislead: the GB200 NVL72 has ~1.6x the TCO of an H100 cluster, so it must be ≥1.6x faster just to break even on perf-per-TCO. 'The hardware you cannot use has infinite TCO.'

CUDA moat. Nvidia's durable advantage isn't the chip — it's the ~4 million external developers who write kernels, file bugs, and ship day-one CUDA implementations (FlashAttention, Mamba, vLLM all launched CUDA-first). SemiAnalysis's sharp framing: the moat isn't dug by Nvidia's engineers but by the millions of outside developers. AMD's RCCL/ROCm libraries are largely forks of Nvidia's, so every NCCL refactor forces AMD to burn engineering hours just to keep pace.

Systolic array / Tensor Core / MXU. The matrix-multiply engine at the heart of every AI chip. Nvidia's Tensor Core (introduced in Volta 2017) does a fused matrix-multiply-accumulate to amortize instruction overhead — a plain FP op costs ~1.5pJ but issuing the instruction costs ~30pJ, a 20x overhead. Google's TPU uses a systolic array; for Trillium (TPUv6) Google quadrupled the array from 128x128 to 256x256 tiles, doubling FLOPS on the same N5 node.

Number formats: FP8, FP4, MXFP, NVFP4. Lower-precision formats trade numerical range/mantissa for throughput and memory. Each generation adds a narrower format: Blackwell added FP4, which roughly doubles throughput vs FP8. Microscaling formats (MXFP8/MXFP4, and Nvidia's NVFP4) attach a shared scale to blocks of values to preserve accuracy. InferenceX shows GB300 delivering up to 100x on FP8-vs-FP4 relative to an H100 baseline — most of Blackwell's inference gain is the format, not just the silicon.

Optical Circuit Switch (OCS) scale-up. Google's alternative to electrical packet switches. An OCS is a reconfigurable mirror-based 'patch panel' that routes whole optical fibers with near-zero latency and no optical-electrical-optical conversion, making it lower-power. It lets Google carve arbitrary 3D-torus slices out of a 9,216-TPU pod, route around faults, and give complete cube 'fungibility.' Building the max world size needs 48 144x144 OCSs. This is the backbone of Nvidia's only real scale-up rival.

Neocloud / ClusterMAX. Neoclouds are GPU-focused clouds (CoreWeave, Nebius, Crusoe, Lambda) that rent GPUs to labs. ClusterMAX is SemiAnalysis's independent rating system (Platinum→Underperforming) grading 84 providers on 10 criteria (security, orchestration, storage, networking, reliability, monitoring, pricing…). CoreWeave is the sole Platinum. The rating has teeth: top-rated neoclouds have booked ~$400B in RPO since v1.0.

HBM as % of BOM / the memory wall. High-Bandwidth Memory has gone from 80GB/3.4TB/s (H100) to 288GB/8TB/s (GB300) in under three years and is now the single largest component of the GB300 package BOM. Because prefill wastes HBM's expensive bandwidth, Nvidia built Rubin CPX with cheap GDDR7 (>50% lower cost/GB) instead — a direct attack on the memory-cost wall.

Open Questions

Sources (SemiAnalysis)

Independence & sourcing. This is independent analysis by Yicheng Yang, distilled from publicly accessible SemiAnalysis articles (free posts and free previews; no paywall circumvention) and verified against the underlying text. It is not affiliated with, endorsed by, or a substitute for SemiAnalysis — subscribe there for the full research. All referenced claims are sourced and linked per SemiAnalysis's attribution terms. No SemiAnalysis images are reproduced. Nothing here is investment advice.