Essay · Market evolution · AI Infrastructure

The inference economy: sizing the prize

Token demand is compounding faster than any input cost is falling. This essay sizes the inference economy, maps how it evolves through 2030, and prices the three coherent endings for the platforms serving it. If one converts demand into an AWS-shaped position on open models, today's multi-billion marks are not the ceiling; they are the entry fee.

Essay by Edouard Fayard v1.1 · July 2026 ← All perspectives

Method demand data → phase model → scenario arithmetic  ·  Reads with Neoclouds & Inference Platforms, our deep-dive on the supply side  ·  Next review Q4 2026

The view, in four lines
DemandTokens are going vertical: one hyperscaler's monthly volume went from 9.7 trillion to 3.2 quadrillion in two years (~330x), and agentic workloads multiply tokens-per-task by 5-100x again. Deflation is not killing this market; deflation is what's growing it.
EvolutionThree phases: the training-capex era (2023-25) → the token price war (now) → the consumption economy (2027-30), where inference becomes a permanent operating expense of every business and a few full-stack platforms consolidate the open-model workload.
The prizeScenario arithmetic on the winning open-model platform: ~$10B outcome if the layer commoditises, ~$45-60B if it becomes the category platform, $120B+ if it traces the AWS path. Against today's ~$8-18B leading marks, that is the actual bet being underwritten.
The swingOne variable decides which scenario: whether the developer world's ~30% open-model token share becomes the enterprise's (today ~11%, and falling). Developers led, enterprises followed is exactly how cloud played out, but it is a bet, not a fact.

Part 01

Demand: the token curve is vertical

Start with the only demand data that can't be argued with, disclosed token volumes. Google processed ~9.7 trillion tokens a month across its products in 2024, ~480 trillion by May 2025, and over 3.2 quadrillion by May 2026, roughly 7x in the last year alone. Microsoft processed over 100 trillion tokens in a single quarter of early 2025, up 5x year-on-year. The largest independent routing layer went from negligible volume to 25 trillion tokens a week. And the composition is shifting as fast as the volume: reasoning models went from a rounding error to more than half of all routed tokens in about a year, and average tokens-per-request has roughly quadrupled since early 2024. Agents compound this, credible estimates put agentic workloads at 5-30x the tokens of a chat interaction, with complex multi-tool tasks far above that.

~330x
Growth in one hyperscaler's monthly token volume, 2024 → mid-2026
Google I/O disclosures
>50%
Share of routed tokens now from reasoning models, near zero in early 2025
OpenRouter State of AI
5-30x
Token multiplier of agentic workloads vs single-turn chat (wide range; some estimates far higher)
Industry estimates, 2026
~1.75x
Revenue effect of prices −75% × volume 7x, deflation and growth coexisting
Orien arithmetic
Tokens processed per month: the demand side of the price war
One hyperscaler's disclosed monthly token volume, trillions.
Source: Google I/O CEO disclosures (2024, May 2025, May 2026). This resolves the token-deflation treadmill described in our neoclouds deep-dive: prices fell ~75% while volume rose ~7x, revenue still grew. Token deflation is the market's growth engine, exactly as compute deflation was for cloud. The question is never whether the market grows; it is who keeps margin while serving it.

Part 02

How the market evolves: three phases

Phase 1, the training-capex era (2023-2025). Value pooled at silicon and scarce compute; the neocloud tier was born in the gap between hyperscaler capacity and frontier-lab demand. This phase built the tier's revenue, and its cost structure.

Phase 2, the token price war (2025-2027, now). Inference overtakes training as the dominant workload, by 2026 this is industry consensus, and the chip vendor now describes the entire buildout as "token manufacturing." Raw compute commoditises, the terrain of our companion deep-dive, Neoclouds & Inference Platforms; token prices collapse; and the collapse itself drives adoption: enterprise generative-AI spend went $1.7B → $11.5B → $37B across 2023-25, tripling in the last year, with ~$12.5B of that flowing through model APIs. Consensus forecasts put the inference market at ~$255B by 2030, with the hosted-platform slice around $105B, but note that these forecasts assume ~17-19% annual growth in a market whose observed spend is tripling annually. Every credible forecast of cloud in 2013 was wrong in the same direction.

Phase 3, the consumption economy (2027-2030). Agents turn inference from a product feature into a metered operating expense of ordinary business, closer to electricity or payroll than to software licensing. When token consumption is a board-level cost line, cost-per-token stops being a nice-to-have and becomes procurement's problem: workloads migrate toward whoever serves equivalent quality cheapest at enterprise grade. That is the moment the open-model platform layer either consolidates into two or three AWS-shaped winners, efficiency edge, workflow gravity, owned capacity, or dissolves into the hyperscalers' bundles. Everything in our subsector stance is about telling those two endings apart early.

The market, sized honestly: forecasts vs the spend curve
US$ billions. Consensus 2030 forecasts sit oddly beside a spend line that tripled last year.
Source: Menlo Ventures enterprise GenAI spend (2023-25 actuals); Grand View / MarketsandMarkets 2030 forecasts (inference market ~$255B, hardware-inclusive; hosted inference platforms ~$105B). The dashed comparison: enterprise spend compounding at even 50%/yr from 2025 reaches ~$280B by 2030, above the "whole market" consensus forecast. We treat the forecasts as a floor, not a ceiling; sell-side CAGRs have never kept pace with a platform shift.

Part 03

The precedent: what "winning a commodity" is worth

The rejoinder to every commodity critique of this layer is a single company. In 2013, AWS did $3.1B of revenue renting a commodity, compute, at margins everyone called structurally thin, in a market everyone knew the incumbents would eventually dominate. It ended 2025 at $128.7B of revenue with ~35% operating margins, generating more operating income than the rest of Amazon combined in some years. The mechanism was precisely the one the platform archetype is attempting: win developers first with price and ergonomics, layer sticky services on the commodity, let enterprises follow the developers, and convert scale into both cost leadership and margin. The leading open-model platforms today, at roughly $1B of revenue, sit where AWS sat around 2012-13, with demand growing faster than cloud's ever did.

The AWS path: commodity rental to $129B at 35% margins
AWS annual revenue, US$ billions, 2013-2025. The platform archetype today ≈ the left edge of this chart.
Source: Amazon segment reporting (compiled). Operating margin roughly ~25% mid-decade → 37% by 2024. The honest differences from the archetype's position: AWS had no deflating unit price remotely as steep as tokens, faced no open-source substitute for its core service, and its parent funded a decade of buildout. The parallel is real; it is not a guarantee.

Part 04

The prize: scenario arithmetic for the platform archetype

What is the winning open-model platform worth in 2030? We refuse to answer with a point estimate, but the arithmetic of three coherent scenarios is knowable, and it is the honest way to hold both our discipline and the upside in one frame. All three use the same skeleton: market size × independent-platform share × leader's share of that layer × margin-appropriate multiple. Every input is stated so you can attack it.

ScenarioWhat has to be trueLeader economics, 2030Indicative value
Commodity endOpen serving stacks equalise the software edge; enterprises stay on closed APIs (open share stays ~11%); platform layer becomes routing + resale~$2-3B rev · ~40% GM~$8-12B ≈ today's marks
Category platformEnterprise open-model adoption inflects toward the developer curve; hosted platforms reach ~$105B; independents hold ~15% of it; leader takes ~a third; margins prove 50%+~$5-6B rev · ~55% GM~$45-60B 5-7x today
The AWS of open inferenceAgent economics make cost-per-token a procurement line; spend tracks the observed curve (~$280B+); leader converts kernel edge + owned power into durable cost leadership and software attach~$10-15B rev · 60%+ GM~$120-180B 15-20x today
What the winner is worth: three coherent endings
Indicative 2030 value of the leading open-model platform, US$ billions (midpoints).
Source: Orien scenario arithmetic (inputs in table above; multiples: ~4x revenue for commodity economics, ~8-10x for platform economics, ~12x for software economics, against 2026 private marks of 7-17.5x ARR for the leading platforms and ~8x EV/revenue for the listed wholesale leader). The red line marks today's leading private marks (~$8-18B). This is arithmetic, not forecast: its purpose is to show that the bull case does not require heroic market assumptions, it requires share and margin assumptions, exactly what the underwriting checklist in our neoclouds deep-dive is built to test.
The swing variable: whose adoption curve wins

The data contains a genuine contradiction, and it is the whole game. On developer-routed traffic, open-weight models are ~30% of tokens and rising, with Chinese open models spiking to 30% of weekly volume after each release. In enterprise surveys, open-model adoption fell from 19% to 11% last year, with closed labs holding ~88% of API share. Bulls read this as cloud circa 2012, developers lead, enterprises follow five years later, and the platform serving the developer curve inherits the enterprise wave. Bears read it as evidence the enterprise never comes, and the open-model layer stays a startup-serving niche squeezed between closed labs and hyperscaler bundles. Every quarterly datapoint on enterprise open-model share is therefore worth more to this thesis than any funding announcement.

Orien's read

Our selective stance on this layer carries a stated upside, not just a stated discipline. The market-evolution view says: demand is not the risk, share and margin are. If the evidence lands (margin structurally above the compute-resale floor, enterprise mix rotating toward the developer curve, retention holding through token price cuts), the leading platform is not a fairly-priced infrastructure business but an early claim on a $45-180B outcome, one worth owning at size. If the evidence doesn't land, the commodity ending values it at roughly today's marks: dead money rather than disaster. That asymmetry, bounded downside to the mark, order-of-magnitude upside on proof, is what makes this layer worth the underwriting effort at all.

Part 05

Sources

Token demand: Google I/O CEO disclosures 2024-26 (9.7T → 480T → 3.2Q monthly tokens); Microsoft FY25 Q3 earnings call (100T quarterly tokens); OpenRouter State of AI report and Series B disclosures (25T weekly, reasoning-model share, tokens-per-request); agentic-multiplier estimates (industry sources, 2026, wide-ranged and flagged as such).

Market sizing & adoption: Menlo Ventures State of Generative AI in the Enterprise (2025, $37B spend, API share, open-model adoption decline); Grand View Research and MarketsandMarkets inference-market forecasts to 2030 (hardware-inclusive ~$255B; hosted-platform PaaS ~$105B); a16z/industry open-model usage coverage.

Precedent & comps: Amazon segment reporting via compiled sources (AWS revenue 2013-2025, operating margins); CNBC AWS margin coverage; 2026 private-round multiples of leading inference platforms (7-17.5x ARR, press-reported); listed wholesale-leader EV/revenue (~8x, Jul 2026); infrastructure-SaaS median multiples (2026).

Data as of 27 July 2026. Scenario values are Orien arithmetic on stated assumptions, illustrations of coherent endings, not forecasts or price targets. All public sources; no confidential information. Companies referenced by archetype are illustrative of the subsector. Not investment advice.