Method demand data → phase model → scenario arithmetic · Reads with Neoclouds & Inference Platforms, our deep-dive on the supply side · Next review Q4 2026
Part 01
Demand: the token curve is vertical
Start with the only demand data that can't be argued with, disclosed token volumes. Google processed ~9.7 trillion tokens a month across its products in 2024, ~480 trillion by May 2025, and over 3.2 quadrillion by May 2026, roughly 7x in the last year alone. Microsoft processed over 100 trillion tokens in a single quarter of early 2025, up 5x year-on-year. The largest independent routing layer went from negligible volume to 25 trillion tokens a week. And the composition is shifting as fast as the volume: reasoning models went from a rounding error to more than half of all routed tokens in about a year, and average tokens-per-request has roughly quadrupled since early 2024. Agents compound this, credible estimates put agentic workloads at 5-30x the tokens of a chat interaction, with complex multi-tool tasks far above that.
Part 02
How the market evolves: three phases
Phase 1, the training-capex era (2023-2025). Value pooled at silicon and scarce compute; the neocloud tier was born in the gap between hyperscaler capacity and frontier-lab demand. This phase built the tier's revenue, and its cost structure.
Phase 2, the token price war (2025-2027, now). Inference overtakes training as the dominant workload, by 2026 this is industry consensus, and the chip vendor now describes the entire buildout as "token manufacturing." Raw compute commoditises, the terrain of our companion deep-dive, Neoclouds & Inference Platforms; token prices collapse; and the collapse itself drives adoption: enterprise generative-AI spend went $1.7B → $11.5B → $37B across 2023-25, tripling in the last year, with ~$12.5B of that flowing through model APIs. Consensus forecasts put the inference market at ~$255B by 2030, with the hosted-platform slice around $105B, but note that these forecasts assume ~17-19% annual growth in a market whose observed spend is tripling annually. Every credible forecast of cloud in 2013 was wrong in the same direction.
Phase 3, the consumption economy (2027-2030). Agents turn inference from a product feature into a metered operating expense of ordinary business, closer to electricity or payroll than to software licensing. When token consumption is a board-level cost line, cost-per-token stops being a nice-to-have and becomes procurement's problem: workloads migrate toward whoever serves equivalent quality cheapest at enterprise grade. That is the moment the open-model platform layer either consolidates into two or three AWS-shaped winners, efficiency edge, workflow gravity, owned capacity, or dissolves into the hyperscalers' bundles. Everything in our subsector stance is about telling those two endings apart early.
Part 03
The precedent: what "winning a commodity" is worth
The rejoinder to every commodity critique of this layer is a single company. In 2013, AWS did $3.1B of revenue renting a commodity, compute, at margins everyone called structurally thin, in a market everyone knew the incumbents would eventually dominate. It ended 2025 at $128.7B of revenue with ~35% operating margins, generating more operating income than the rest of Amazon combined in some years. The mechanism was precisely the one the platform archetype is attempting: win developers first with price and ergonomics, layer sticky services on the commodity, let enterprises follow the developers, and convert scale into both cost leadership and margin. The leading open-model platforms today, at roughly $1B of revenue, sit where AWS sat around 2012-13, with demand growing faster than cloud's ever did.
Part 04
The prize: scenario arithmetic for the platform archetype
What is the winning open-model platform worth in 2030? We refuse to answer with a point estimate, but the arithmetic of three coherent scenarios is knowable, and it is the honest way to hold both our discipline and the upside in one frame. All three use the same skeleton: market size × independent-platform share × leader's share of that layer × margin-appropriate multiple. Every input is stated so you can attack it.
| Scenario | What has to be true | Leader economics, 2030 | Indicative value |
|---|---|---|---|
| Commodity end | Open serving stacks equalise the software edge; enterprises stay on closed APIs (open share stays ~11%); platform layer becomes routing + resale | ~$2-3B rev · ~40% GM | ~$8-12B ≈ today's marks |
| Category platform | Enterprise open-model adoption inflects toward the developer curve; hosted platforms reach ~$105B; independents hold ~15% of it; leader takes ~a third; margins prove 50%+ | ~$5-6B rev · ~55% GM | ~$45-60B 5-7x today |
| The AWS of open inference | Agent economics make cost-per-token a procurement line; spend tracks the observed curve (~$280B+); leader converts kernel edge + owned power into durable cost leadership and software attach | ~$10-15B rev · 60%+ GM | ~$120-180B 15-20x today |
The data contains a genuine contradiction, and it is the whole game. On developer-routed traffic, open-weight models are ~30% of tokens and rising, with Chinese open models spiking to 30% of weekly volume after each release. In enterprise surveys, open-model adoption fell from 19% to 11% last year, with closed labs holding ~88% of API share. Bulls read this as cloud circa 2012, developers lead, enterprises follow five years later, and the platform serving the developer curve inherits the enterprise wave. Bears read it as evidence the enterprise never comes, and the open-model layer stays a startup-serving niche squeezed between closed labs and hyperscaler bundles. Every quarterly datapoint on enterprise open-model share is therefore worth more to this thesis than any funding announcement.
Our selective stance on this layer carries a stated upside, not just a stated discipline. The market-evolution view says: demand is not the risk, share and margin are. If the evidence lands (margin structurally above the compute-resale floor, enterprise mix rotating toward the developer curve, retention holding through token price cuts), the leading platform is not a fairly-priced infrastructure business but an early claim on a $45-180B outcome, one worth owning at size. If the evidence doesn't land, the commodity ending values it at roughly today's marks: dead money rather than disaster. That asymmetry, bounded downside to the mark, order-of-magnitude upside on proof, is what makes this layer worth the underwriting effort at all.
Part 05
Sources
Token demand: Google I/O CEO disclosures 2024-26 (9.7T → 480T → 3.2Q monthly tokens); Microsoft FY25 Q3 earnings call (100T quarterly tokens); OpenRouter State of AI report and Series B disclosures (25T weekly, reasoning-model share, tokens-per-request); agentic-multiplier estimates (industry sources, 2026, wide-ranged and flagged as such).
Market sizing & adoption: Menlo Ventures State of Generative AI in the Enterprise (2025, $37B spend, API share, open-model adoption decline); Grand View Research and MarketsandMarkets inference-market forecasts to 2030 (hardware-inclusive ~$255B; hosted-platform PaaS ~$105B); a16z/industry open-model usage coverage.
Precedent & comps: Amazon segment reporting via compiled sources (AWS revenue 2013-2025, operating margins); CNBC AWS margin coverage; 2026 private-round multiples of leading inference platforms (7-17.5x ARR, press-reported); listed wholesale-leader EV/revenue (~8x, Jul 2026); infrastructure-SaaS median multiples (2026).
Data as of 27 July 2026. Scenario values are Orien arithmetic on stated assumptions, illustrations of coherent endings, not forecasts or price targets. All public sources; no confidential information. Companies referenced by archetype are illustrative of the subsector. Not investment advice.