Overview
SemiAnalysis's model/lab coverage is built around one contrarian spine: the 'scaling laws are dead' narrative is wrong because scaling is not one curve but a stack of them. When pre-training hit the data wall, the labs pivoted to post-training — supervised fine-tuning, RLHF/RLAIF, and above all reinforcement learning on verifiable rewards — and then to test-time (inference) compute via reasoning models. OpenAI rode the same GPT-4o base model for 18 months through o1, o3 and GPT-5, proving post-training alone drives capability. The consequence is a different industrial structure than pre-training: where everyone once trained on the same internet, RL requires bespoke environments, verifiers, and expert-authored data that must be built from scratch — so 'data is the moat.' An entire shovel-selling economy has erupted: 35+ RL-environment startups, UI-gym website clones at ~$20k each, and data foundries (Surge ~$1B ARR, Mercor, Handshake) that recruit PhDs to write tasks and rubrics. On the product side, Claude Code is SemiAnalysis's declared inflection point — coding is the beachhead into a $15T information-work economy, the six-player coding-assistant market already clears $30B+ ARR heading to $100B, and Anthropic is now adding more revenue per quarter than OpenAI. The technical throughline is that RL fuses inference into training (rollouts are inference-heavy), which reshapes hardware demand toward memory-bound inference chips (Trainium), CPUs for RL environments, and decentralized rather than mega-centralized clusters. Below the frontier, DeepSeek and other Chinese open-weights labs (MLA cutting KV cache ~93%, V4 hitting 1M context with a 90% KV-cache cut) commoditize the low end, but still trail on agentic coding — 'data is the moat' keeps the frontier ahead. Benchmarks are broken (SWE-bench contamination, GDPval single-turn), so the real north-star metric is cost-per-task, not cost-per-token.
Positioning: Who Wins and Why
Anthropic — SemiAnalysis's clearest structural winner: Claude Code is the declared agentic inflection, ~4% of GitHub commits and climbing, and quarterly ARR additions have overtaken OpenAI's — growth capped by compute, not demand. Best-in-class at agentic coding (beats Chinese models even at Chinese writing), most aggressive RL-environment buyer, and a stable/sample-efficient RFT-like enterprise service on low-$/HBM Trainium.
OpenAI — Distribution king: ChatGPT is the #5 website with 700M+ free users, and the router lays the groundwork to monetize them via agentic purchasing/take-rates. GPT-5.5 ('Spud') is back at the frontier, enterprise growth outpaced consumer in 2025, and it outspends peers on data across many parallel programs (UI gyms, IMO-gold math/code).
Data foundries & RL-env startups (Surge, Mercor, Handshake, Prime Intellect, Mechanize) — The picks-and-shovels of the RL boom. 'During an RL scaling boom, sell environments.' Surge ~$1B ARR, Scale did ~$870M before Meta absorbed it; demand for coding environments is so high defunct startups are bought just for private GitHub repos. Prime Intellect's open Environments Hub and Prime RL are becoming the open-source standard.
Google DeepMind — Uniquely positioned if capital/compute is king: owns TPUs (lowest cost per inference) and the apps to post-train on (Sheets, Docs, Drive, Maps) with PM visibility into how users actually use them. Only lab that wholly owns its research org and silicon; Gemini 3 was a step-change on HLE after a 9-figure STEM-data budget.
Nvidia (GB200/GB300 NVL72) — RL's memory-heavy, inference-fused profile rewards rack-scale NVLink world size: GB300 mogs 8-GPU islands at reasoning ($0.156/Mtok on DeepSeek V4) by keeping MoE all-to-all on NVLink. Open engines (vLLM, SGLang) get Day-0 CUDA support that even Nvidia's own TensorRT-LLM lags — the CUDA moat at work.
DeepSeek & Chinese open-weights labs — Set the efficiency frontier (MLA ~93% KV-cache cut; V4 1M context at ~10% KV cache) and keep American open-source alive via DeepEP/DeepGEMM/FlashMLA. Kimi K2.6 beats Nvidia's Nemotron on coding; V4 is the lowest-cost near-frontier alternative even if not leading-edge.
Key Data
| Metric | Value | Note |
|---|---|---|
| Claude Code share of GitHub commits | ~4% now → projected 20%+ by end-2026 | SemiAnalysis's headline datapoint for the agentic inflection; 'Pretty much 100% of our code is written by Claude Code + Opus 4.5' — Boris Cherny, Claude Code creator. |
| Coding-assistant market ARR | $30B+ across 6 players, on track past $100B by year end | 'The greatest B2B SaaS application the world has ever seen,' per the Tokenomics model. |
| Information-work TAM at risk | $15T economy, 1B+ workers (1/3 of the 3.6B global workforce) | Coding is only the beachhead; the READ/THINK/WRITE/VERIFY workflow generalizes to all information work. |
| GDPval win rate (GPT-5.2) | ~71% tie-or-preferred vs human experts across 44 occupations | OpenAI eval of real economically-valuable tasks; the models ran ~100x faster and ~100x cheaper, but experts kept the quality/reliability edge (best model lost-or-tied ~half its head-to-heads) — augmentation, not automation. |
| MLA KV-cache reduction | ~93.3% vs standard attention | DeepSeek's Multi-Head Latent Attention — the core reason its inference is so cheap. |
| DeepSeek V4 architecture | Pro 1.6T total / 49B active; Flash 284B / 13B; 128k → 1M context | Claims only 27% of single-token FLOPs and 10% of KV cache vs V3.2 at 1M tokens (a ~90% KV-cache cut). |
| GB300 NVL72 cost per Mtok | $0.156 per 1M output tokens (DeepSeek V4, 50 tok/s/user, 8k in/1k out) | Rack-scale NVLink world size keeps MoE all-to-all on NVLink — why GB300 mogs 8-GPU islands at reasoning. |
| MI355X software gains (DeepSeek V4) | >100x throughput improvement Day 0 → Day 26 | Pure software (AITER/Triton/TileLang kernels replacing PyTorch fallbacks) — shows the CUDA-moat / ecosystem-maturity gap. |
| Inference cost-down for GPT-3 quality | ~1200x cheaper | Algorithmic progress ~4x/year (Dario Amodei argues ~10x); Jevons paradox reinvests savings into bigger models. |
| UI-gym / environment cost | ~$20,000 per cloned website; OpenAI bought hundreds | One-time purchases reused across models; trajectories fed back into mid-training. |
| Data-foundry revenue | Surge ~$1B ARR; Scale AI ~$870M (2024), ~$1.5B exit run-rate before Meta absorbed it | Mercor, Handshake, Aboda.ai fill the gap after labs fled Scale post-Meta acquisition. |
| Opus 4.7 tokenizer price impact | up to +35% token usage (implicit +35% price) | New tokenizer trades finer counting for more total tokens — why cost-per-task beats cost-per-token. |
| ChatGPT reach & free base | #5 website globally, 700M+ mostly-unmonetized free users | The router is the groundwork for monetizing them via agentic purchasing / take-rates. |
| GPT-5.5 API pricing | $5 / 1M input, $30 / 1M output (2x GPT-5.4, ~Opus 4.7 level) | 'Trained on a 100k GB200 NVL72 cluster' was post-training (RL) only; priority tier is 2.5x standard. |
| Qwen post-training compute | ~5% of pre-training compute | Most Chinese labs are early to RL scaling; homegrown data foundries would accelerate the transition. |
| Autonomous task-horizon doubling | every ~7 months for coding (accelerating to ~4 months in 2024-25) | METR data; each doubling unlocks more of the automation pie (30 min → refactor a module → multi-day audit). |
| Anthropic RFT vs young startups | YC RL-as-a-service serves at ~5x lower cost than OpenAI's RFT | OpenAI's RFT is unstable/expensive; long-run the labs capture most revenue via 'Strategic Deployment' teams. |
| Scale of coding-env task generation | DeepSeek used 24,667 coding tasks for V3.2; labs instantiate 10,000+ sandboxes simultaneously | SWE-rebench yields only 21,336 valid tasks from 450k PRs — strict execution-based verification. |
| Meta Muse Spark (Apr 2026) | ≈ Opus 4.6 / GLM 5.2 for general agentic use; priced just under GLM 5.2; lagged DeepSeek v4 Pro and Kimi K2.6 on most benchmarks at launch | Meta's post-Llama-4 flagship from Meta Superintelligence Labs; ~3,000 engineers work full-time authoring RL tasks (Meta Superintelligence 1-yr update, SemiAnalysis). |
Key Theses
“It is possible to stack 'scaling laws' - pre-training will become just one of the vectors of improvement, and the aggregate 'scaling law' will continue scaling just like Moore's Law has over last 50+ years.”— Scaling Laws – O1 Pro Architecture, Reasoning Training Infrastructure, Orion and Claude 3.5 Opus 'Failures' (2024-12-11), SemiAnalysis
“RL for language models is one of the first cases where inference has truly become intertwined into the training process. Inference performance now directly impacts training speed.”— Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Data (2025-06-08), SemiAnalysis
“Ultimately what Qwen signals is that high quality data is a uniquely important resource for scaling RL... This paradigm is not like pre-training where everyone had access to the same data.”— Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Data (2025-06-08), SemiAnalysis
“During a Gold Rush, sell shovels. In an RL Scaling Boom, sell RL environments. More than 35 companies have popped up whose goal is to do exactly this across a variety of domains.”— RL Environments and RL for Science: Data Foundries and Multi-Agent Architectures (2026-01-06), SemiAnalysis
“4% of GitHub public commits are being authored by Claude Code right now... While you blinked, AI consumed all of software development.”— Claude Code is the Inflection Point (2026-02-05), SemiAnalysis
“Notably, our forecast shows that Anthropic's quarterly ARR additions have overtaken OpenAI's. Anthropic is adding more revenue every month than OpenAI.”— Claude Code is the Inflection Point (2026-02-05), SemiAnalysis
“cost per task, not cost per token, is the true north star metric that determines model pricing.”— The Coding Assistant Breakdown: More Tokens Please (2026-04-24), SemiAnalysis
“Non-verifiable domains include areas like writing or strategy, where no clear right answer exists. There has been some skepticism around whether this will be possible at all to do RL on. We think it is. In fact, it's already been done.”— Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Data (2025-06-08), SemiAnalysis
“Claude for Excel effectively is what Copilot for Excel should have been, but it was launched by an external party on their own first party product.”— Claude Code is the Inflection Point (2026-02-05), SemiAnalysis
“The $6M cost in the paper is attributed to just the GPU cost of the pre-training run, which is only a portion of the total cost of the model... We are confident their hardware spend is well higher than $500M over the company history.”— DeepSeek Debates: Chinese Leadership On Cost, True Training Cost (2025-01-31), SemiAnalysis
“We believe that the Router is the groundwork for the next leg of ChatGPT's story, and that's monetization of free users.”— GPT-5 Set the Stage for Ad Monetization and the SuperApp (2025-08-13), SemiAnalysis
“RL environments are spilling into the physical world. These environments are no longer just a docker container... It is an experiment that needs to be run by a human or a robot, with a real cost to materials, electricity, equipment, and lab space.”— RL Environments and RL for Science: Data Foundries and Multi-Agent Architectures (2026-01-06), SemiAnalysis
“Anthropic is bringing on large volumes of Trainium, which is economically attractive for RL in part due to low $/HBM... RL workloads are mostly inference, which is memory bound.”— RL Environments and RL for Science: Data Foundries and Multi-Agent Architectures (2026-01-06), SemiAnalysis
“RL training system efficiency is a matter of queue health... in-flight weight updates allow trainer and generator to operate at different rates.”— RL Systems Mind the Gap: Matching Trainer and Generator Throughput (2026-06-16), SemiAnalysis
Article Deep-Dives
Reference: Value Chain
Reference: Core Concepts
The scaling stack (pre-train → post-train → test-time). SemiAnalysis's core framing: overall AI 'scaling' is the sum of three compounding vectors, not just pre-training. Pre-training (next-token prediction on the internet) hit a data wall; post-training (SFT/RLHF/RLAIF/RL on verifiable rewards) unlocked reasoning; test-time compute lets a model 'think' longer for a better answer. Like Moore's Law surviving the end of clock-speed scaling via multi-core, the aggregate curve keeps climbing even as one vector slows.
RL on verifiable rewards (GRPO/PPO). RL worked best where a reward is cleanly checkable — code (does it pass tests?) and math (is the answer right?). The model generates many 'rollouts' (attempts) per prompt, each scored; GRPO (DeepSeek's algorithm) drops PPO's critic model to save memory by scoring each rollout against the group average. This makes RL inference-heavy: hundreds of answers per question, so production-grade inference is now inside the training loop.
RL environments & sandboxes. An 'environment' is the simulated world (a code sandbox, a cloned website, a browser, a spreadsheet app) where an agent acts and receives a reward. Engineering robust, low-latency, exploit-resistant sandboxes (Firecracker micro-VMs to full QEMU VMs) is a first-class systems problem — sandbox startup latency and scaling to thousands of concurrent rollouts are real bottlenecks. Reward hacking (the model gaming the reward without doing the task) is the recurring failure mode.
Synthetic data, RLAIF & LLM-as-judge. As human data stopped scaling, labs turned to synthetic data (rejection sampling, model-generated tasks) and replaced human feedback (RLHF) with AI feedback (RLAIF). A model or 'LLM judge' scores outputs against a rubric — the key that opens non-verifiable domains like writing, medicine and law. Better models are better judges, creating a self-improving loop. Anthropic's Constitutional AI is the canonical RLAIF example.
Agent / harness (Claude Code). An agent wraps a model in a loop of READ → THINK → WRITE → VERIFY with tools, memory and sub-agents. Claude Code is a terminal-native agent ('Claude Computer') that reads a codebase, plans, and executes. The 'harness' (tools, prompts, context management, caching) is now part of the product — it drives cost-per-task as much as the model does, which is why same-harness benchmark comparisons miss the point.
Cost-per-task vs cost-per-token. SemiAnalysis's north-star pricing metric. A model priced 5x higher per token (e.g. Mythos vs Opus) can be cheaper in practice if it solves the task in far fewer tokens ('token efficiency'). Anthropic's Opus 4.7 tokenizer change alone raised token usage up to 35% — an implicit 35% price hike. The harness (Codex ~80:1 input/output vs Claude Code ~100:1) is a major driver.
Data foundry / RL-as-a-service. Firms that supply the labs with RL environments, expert-authored tasks, and grading rubrics. Environment startups (Mechanize, Turing, Prime Intellect, HUD) clone apps and wrap software in dockerized sandboxes; data foundries (Surge, Mercor, Handshake) recruit domain experts. RL-as-a-service startups (RunRL, Osmosis, Applied Compute, Thinking Machines Tinker) post-train custom models (often Qwen) for enterprises at ~5x lower cost than OpenAI's RFT.
Multi-Head Latent Attention (MLA) & MoE. DeepSeek's signature efficiency innovations. MLA compresses the KV cache ~93% versus standard attention, slashing the memory (and thus cost) per query. Combined with fine-grained Mixture-of-Experts (many small experts, few active per token) and FP8 training, this is how a Chinese lab reached near-frontier quality at a fraction of the compute — and forced Western labs to copy the ideas almost immediately.
Open Questions
- Is the path to AGI just environments stacked on top of each other, or does RL hit a wall in non-verifiable, long-horizon, sparse-reward domains (computer use, biology) that rubrics can't fully paper over?
- Does the frontier stay ahead of DeepSeek and Chinese open-weights labs on agentic coding, or does 'data is the moat' erode as homegrown Chinese data foundries stand up and Qwen scales past 5% post-training compute?
- Can Microsoft resolve its conundrum — accelerate Azure (renting GPUs to the barbarians) without letting Claude Code and agents gut its seat-based Office 365 terminal value?
- Will LLM automation be genuine job replacement or mostly augmentation? GDPval and radiology suggest the latter for expert work — but maybe not for shorter-horizon repetitive tasks like call centers.
- How decentralized can RL get? If synthetic-data generation and training can live in different datacenters, does the mega-cluster arms race matter less than everyone assumes — and who captures the 'free' idle-inference compute edge?
Sources (SemiAnalysis)
- Scaling Laws – O1 Pro Architecture, Reasoning Training Infrastructure, Orion and Claude 3.5 Opus 'Failures' (2024-12-11)
- Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Data (2025-06-08)
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architectures (2026-01-06)
- Claude Code is the Inflection Point (2026-02-05)
- The Coding Assistant Breakdown: More Tokens Please (2026-04-24)
- RL Systems Mind the Gap: Matching Trainer and Generator Throughput (2026-06-16)
- DeepSeekV4 1.6T Day 0 to Day 43 Performance Over Time - Huawei, GB300 NVL72, MI355X, B200 (2026-06-09)
- DeepSeek Debates: Chinese Leadership On Cost, True Training Cost, Closed Model Margin Impact (2025-01-31)
- GPT-5 Set the Stage for Ad Monetization and the SuperApp (2025-08-13)
- CPUs are Back: The Datacenter CPU Landscape in 2026 (2026-02-09)
- OpenAI Is Doomed? - Et tu, Microsoft? (2024-05-07)
- Microsoft Swallows OpenAI's Core Team – GPU Capacity, Incentive Structure, Intellectual Property (2023-11-20)
- Amazon Anthropic: Poison Pill or Empire Strikes Back (2023-10-02)
- Google 'We Have No Moat, And Neither Does OpenAI' (2023-05-04)
- Google Gemini Eats The World – Gemini Smashes GPT-4 By 5X, The GPU-Poors (2023-08-28)
- GPT-4 Architecture, Infrastructure, Training Dataset, Costs, Vision, MoE (2023-07-10)
- How Nvidia's CUDA Monopoly In Machine Learning Is Breaking - OpenAI Triton And PyTorch 2.0 (2023-01-16)
- TPUv5e: The New Benchmark in Cost-Efficient Inference and Training for <200B Parameter Models (2023-09-01)
- The Future of Meta Superintelligence: A 1 Year Progress Update (2026-07-09)