Subsector Deep-Dive · GPU Clouds & Open-Model Serving

Renting intelligence: margin, or a margin illusion?

A new cloud tier grew from nothing to $25B in revenue in three years by renting the world's scarcest asset. Then the price of a GPU-hour fell 64%. What's left, and who keeps it, is the whole investment question.

The call, in four lines
OwnLayers where scarcity is durable: contracted capacity anchored in power (the subject of our forthcoming powered-land deep-dive), and, selectively, the thin set of serving platforms that can prove software economics on top of commodity FLOPs.
WatchThe first public disclosure of an inference platform's real gross margin; Blackwell-generation rental pricing; hyperscalers and Meta entering the neocloud's own market.
AvoidMerchant GPU rental without committed contracts, and paying software multiples for margins that are still infrastructure-grade.
WhyRaw compute is a depreciating commodity financed with 9-10% debt: rental prices decay ~15-18% a year and two-thirds of a GPU's lifetime cash flow lands in its first three years. Any durable moat has to live somewhere other than the GPU itself.

01The shortage that built a cloud tier

The neocloud exists because the hyperscalers could not build fast enough. When frontier-model training and then inference demand outran every capacity plan, a new tier of specialist GPU clouds grew in the gap, and grew violently: the segment cleared $25B of revenue in 2025, up 223% year-on-year in the fourth quarter, with credible forecasts approaching $400B by 2031. Six names have scale; behind them, more than 300 providers entered the H100 rental market in 2025 alone. That last number is the tell. This is a business with almost no barrier to entry at the commodity end, anyone with capital and a Nvidia allocation can stand up GPUs, and the price curve shows it. The demand side of this market, token volumes, agentic workloads, and how the curve evolves to 2030, is sized in our companion note, The Inference Economy; this piece is about who keeps margin while serving it.

$25B+
Neocloud segment revenue, FY2025, Q4 up 223% YoY
Synergy Research, Apr 2026
~$400B
Forecast segment size by 2031 (~58% CAGR)
Synergy Research
300+
New H100 rental providers entered in 2025, the commodity long tail
Introl, Dec 2025
−64%
H100 $/hr decline from the 2023 peak
Silicon Data index
The price of a GPU-hour is a melting asset
H100 rental, US$ per GPU-hour, from 2023 scarcity peak to late-2025 reality.
Source: Silicon Data index / Introl market survey, Dec 2025. Peak scarcity pricing ~$8; on-demand at scale providers ~$2.49; one-year reserved ~$2.00; estimated provider break-even ~$1.65; aggressive spot as low as $0.99, below break-even. Prices for the raw commodity now sit within sight of cost, three years into the boom.

02The framework: scarcity decays at different speeds

The organising mechanism for the whole AI Infrastructure sector is simple: value pools wherever scarcity is durable, and every scarcity in this stack decays at a different speed. Power and grid connections decay slowest (a decade). Data-centre shells are next (years). GPUs decay fastest of all, not just because supply catches up, but because the asset itself depreciates: each new silicon generation devalues the last, and each price cut reprices the installed base. The neocloud's tragedy is structural: it sits on the fastest-decaying scarcity in the stack, holding the depreciating asset, funded by debt that outlives the asset's pricing power.

Within the tier, two distinct business models have emerged, and they should be underwritten as different asset classes, because they are.

The wholesale contract model
Sell committed capacity, multi-year, to giants
  • Revenue is real and contracted: the listed leader carries a $99.4B backlog on ~5-year weighted terms
  • Moat = Nvidia allocation, deployment speed, scale financing
  • Risks: one customer at 67% of revenue; $24.9B of debt; interest ≈ 46% of EBITDA; renewal at reset prices
  • In truth a credit instrument wearing a growth multiple
The inference platform model
Sell tokens and APIs on open models, to thousands
  • Serves open-weight models (Llama / DeepSeek / Qwen class) to AI-native builders; usage tripled in a year
  • Moat claim = research-grade serving software (kernels, compilers, speculative decoding) + workflow and data gravity + buying power from demand density
  • Risks: token prices collapsing (one leading lab cut output prices 75% in one move); customers are burn-funded startups
  • Priced as software; margins disclosed so far are infrastructure-grade

The market is paying software prices for the second model ahead of public proof. The archetypes re-rated violently through 2025-26: one serving platform went from a $4B valuation to $17.5B in nine months on ~$1B of ARR, and the leading platforms now carry multi-billion marks on revenue around the billion-dollar scale. What the tier has not yet shown publicly is a software gross margin: where numbers have surfaced, they sit in the mid-40s, solid infrastructure economics, not yet the 70-80% the multiples imply. That gap is the central diligence question, and it is answerable rather than rhetorical, because the bull mechanics are real. The best platforms employ the researchers who wrote the defining inference-efficiency kernels, compound that into a genuine cost-per-token lead (customers report multi-fold serving-cost reductions on switching), and are now vertically integrating into owned power and capacity at hundreds-of-megawatts scale to convert the cost lead into margin. That is the AWS trajectory in miniature. Whether it holds turns on whether the serving-software edge stays proprietary faster than open-source stacks absorb it, the moat is the research team's velocity as much as any artifact.

The margin ladder: headline, reported, and real
Gross margin, % of revenue, what the tier claims vs what survives depreciation.
Source: McKinsey (via press, 2026); company reporting. Neocloud headline gross margins of 55-65% compress to 14-16% after depreciation, and returns flatline below ~80% utilisation. The leading open-model serving platforms report gross margins in the mid-40s, better, but still infrastructure, not software. The entire bull case for the platform layer is that software attach lifts that number; it is not yet proven in public.

03The economics: depreciation is the whole debate

Every argument about this subsector reduces to one accounting line, so we tested it against market data rather than opinion, in a five-generation study of rental-price decay. Rental prices across five chip generations decay at a consistent ~15-18% per year against launch-era rates; cash-flow life runs to roughly eight or nine years (Azure retired its V100 fleet at 7.5); and about two-thirds of a GPU's lifetime cash flow arrives in its first three years. That evidence cuts both ways. The sceptics' claim that a rented GPU is economically dead in two-to-three years, the basis for a much-cited $176B understated-depreciation estimate, is not supported: old chips still rent at multiples of their running cost. But the books do not get off clean either: straight-line depreciation over five-to-six years charges ~17% of the asset per year while the economics consume ~28% in year one, so early-year reported profits, which is all this young sector has yet shown, are flattered by roughly 1.6x on the depreciation line. Amazon's own move (cutting server lives from six to five years, a $677M nine-month hit) points the direction. The sharpest finding is about vintage rather than life: re-based to the 2023 scarcity peak, the same chips show ~24%/yr decay. The worst outcomes in this tier come not from machines dying young but from capacity bought, leased or debt-financed at panic prices.

The decay curve: value retained vs age
Marketplace rental rate as % of own launch-era rate, by generation age (mid-2026). Line = fitted ~18%/yr decay.
Source: Orien analysis of launch-era vs current marketplace rates (AWS/GCP launch pricing; Lambda, RunPod, Hyperstack medians, Jul 2026). The hollow point re-bases the H100 against its 2023 scarcity peak rather than normal launch pricing, anchoring to panic prices makes decay look ~24%/yr. H200 rents actually rose ~28% YoY on memory-bound demand: decay is a trend, not a law, and scarcity cycles overlay it.
Old chips still earn their keep, until year eight or nine
US$ per GPU-hour, mid-2026 marketplace/neocloud rates vs the cash cost of operating the machine.
Source: Cross-provider medians (Jul 2026); CloudRift colo TCO model for the cash-cost floor (H100-class; older generations draw less power, so their true floor is lower). A generation dies economically when its bar crosses the red line, on the current curve a year 8-9 event, not year 3. Azure retired its V100 fleet at 7.5 years, right on schedule.
A GPU is a salmon run: the catch comes early
Share of lifetime net cash flow, modelled H100-class buyer at normal launch pricing.
Source: Orien model (18%/yr price decay, 80% utilisation, $0.80/hr cash floor). Years 1-3 deliver ~66% of lifetime cash flow, ~74% in present-value terms. The tail through year nine is real but thin, and it is the part of the curve that services the later years of 9%+ debt.

The financing stack compounds the problem. The listed wholesale leader added roughly $21.9B of debt in the first half of 2026 alone, at 8.5-9.75% on unsecured notes, against an asset whose rental price has fallen 64% from peak. Interest expense now consumes nearly half of adjusted EBITDA. And the capital keeps coming partly because the silicon vendor recycles it: a $2B direct investment into that same leader, stakes across the tier, over $540B of announced partnership and financing arrangements in 2026, and a discussed $250B lease-payment guarantee for its largest end-customer. The circularity is not a scandal; it is a structure, but it means headline demand signals (backlogs, bookings, round sizes) are partially manufactured by the supply chain that benefits from them, and should be discounted accordingly.

One quarter in the life of a wholesale neocloud
US$ billions, the listed leader, Q1 2026. The shape of the model in four numbers.
Source: Company Q1 2026 earnings release (SEC). A $99.4B contracted backlog is genuinely enormous, and it is serviced by $2.1B of quarterly revenue, $7.7B of quarterly capex and $24.9B of debt. Equity here is a levered call option on utilisation staying high and prices staying up. The public market has noticed: the stock is roughly half its 52-week high.
The token deflation undertow

Beneath the GPU price curve runs a second, faster one: the price of the output. One leading open-model lab cut its flagship output pricing 75% in a single move in mid-2026, $3.48 to $0.87 per million tokens. Token deflation is wonderful for AI adoption and brutal for anyone whose revenue is a markup on tokens. A serving platform only escapes it by charging for something the token price war can't touch: latency guarantees, fine-tuning workflows, enterprise deployment, data gravity. That is the moat test. And the treadmill has a second face: the same deflation is what is growing the market, prices fell ~75% while one hyperscaler's token volume rose ~7x in a year, so revenue still grew. Why deflation and demand compound rather than cancel is the subject of The Inference Economy.

04How to underwrite a name in this subsector

This is the checklist we apply to any neocloud or inference platform that crosses our desk, private allocation or public position. It is designed to separate the contracted-cashflow story from the commodity-rental story, and the software claim from the software reality.

  • Contract cover. What share of deployed capacity is under committed, take-or-pay contract, at what weighted duration, and what share is merchant? Merchant exposure is exposure to the $0.99 spot print.
  • Customer quality and concentration. Who actually pays, hyperscalers and profitable enterprises, or burn-funded AI startups whose own runway is a venture-market variable? A backlog is only as good as its weakest counterparty. Anything above ~30% single-customer concentration is a structural, not incidental, risk.
  • Margin truth. Demand gross margin after depreciation on an economics-matched schedule (declining-balance at the observed ~15-18%/yr rental decay, not straight-line). For platforms: is software priced and reported separately from compute resale, or blended into one number the multiple flatters? And what is the margin by layer, token API vs dedicated endpoints vs cluster rental?
  • Price-basis risk. What happens to the model at renewal, when today's contracted $/GPU-hour resets to the then-current curve? A 5-year contract signed at 2024 prices is an asset; its renewal is the business.
  • Capital stack vs the cash-flow curve. A GPU delivers two-thirds of its lifetime cash in its first three years; 9%+ debt serviced from the thin end of that curve in later years demands utilisation and pricing hold simultaneously. And check the vintage: capacity bought at scarcity peaks decays ~24%/yr, not 15-18%.
  • Vendor circularity. Is the chip vendor in the cap table, in the backlog, or both? Discount demand signals that the supply chain itself is financing.

05Orien's verdict

Position

Selective, underwrite margins, not bookings. The demand is real, the revenue growth is real, and most of the tier's economics are still a commodity-rental business dressed in software multiples. We do not own raw GPU rental. We will selectively own the platform layer, but only where a name can evidence software economics (separately-priced software attach, retention through token price cuts, margin structurally above the compute-resale floor). At the commodity end, the best risk-reward may come when the cycle reprices the long tail, we keep capacity to act if it does.

SegmentReadStance
Open-model inference platformsThe one candidate software moat in the tier; unproven margins, violent marks, diligence the margin ladder, not the bookingsSelective
Committed-contract wholesale neocloudsReal backlogs, brutal capital intensity; equity = levered utilisation bet; public entry points exist and have repriced ~50%Selective (public)
Merchant / spot GPU rental300-provider long tail selling below break-even; no moat, no contracts, first casualty of the cycleAvoid
Vendor circularity across the tierChip-maker capital in rounds and backlogs is not disqualifying, but demand signals the supply chain finances get discounted, not taken at face valueDiscount
Hyperscaler & mega-cap entry (incl. Meta)The tier's existential competitor; one report of Meta building competing AI cloud cut 14-17% off listed neoclouds in a dayTrack

Signposts we track

  • A platform discloses real gross margin. The first credible public print of an inference platform's software-layer margin settles the subsector's central question, in either direction.
  • Blackwell-generation pricing. If GB200-class rental holds price materially better than the H100 curve did, the depreciation bear case softens; if it repeats the curve, it confirms.
  • A credit event in the long tail. Sub-break-even spot pricing plus 9%+ debt across a 300-provider long tail should eventually produce one; it would reprice the whole tier and reset entry marks.
  • Enterprise mix rotation. The platform bull case shows up as lengthening contracts and a rising share of enterprise revenue versus AI-native startup spend. Flagship customer lists still skew to the latter; whether that rotates is the software claim materialising in the revenue mix.
  • Renewal prints. The first big wholesale contracts signed in 2023-24 come up for renewal from 2028; the re-contracting price is the truth serum.

06What breaks this call

The software margin proves real. If the leading platforms demonstrate durable 60%+ gross margins as token prices fall, genuine software economics on commodity FLOPs, the layer re-rates as software and our caution will have cost us the best names at pre-proof prices. This is the key upside risk, and the reason our stance is selective rather than avoid.

Scarcity lasts longer than we model. If power constraints, the subject of our forthcoming powered-land deep-dive, keep effective GPU supply tight through 2028+, rental prices stabilise, utilisation stays high, and the levered wholesale model works long enough to deleverage. The neocloud bear case is a supply forecast; supply forecasts have been wrong in both directions.

The vendor keeps underwriting the tier. $540B+ of announced vendor financing can defer the cycle for years. Circularity is a risk that pays out slowly, being right early here has the same P&L as being wrong.

Demand goes vertical again. A step-change in inference demand (agents, video, robotics workloads) could outrun even the current buildout, restoring pricing power across the tier. In that world the commodity is scarce again and everything in it re-rates. The token data in The Inference Economy suggests this is the upside scenario to take most seriously.

07Sources

Market structure: Synergy Research neocloud forecast (Apr 2026); Introl H100 market survey (Dec 2025); Silicon Data GPU rental index.

Wholesale economics: CoreWeave Q1 2026 earnings release and FY2025 10-K data (SEC), backlog, concentration, debt, capex; 2026 debt issuance terms; McKinsey neocloud margin analysis (via press); Amazon useful-life change (10-Q); depreciation-scepticism coverage (Nov 2025).

Platforms & tokens: 2025-26 funding rounds and Series C / capacity-commitment announcements of open-model serving platforms (TechCrunch, Fortune, Businesswire, company announcements); DeepSeek V4-Pro pricing change (InfoWorld, May 2026); open-model usage growth (2026 coverage); Orien Research, The Inference Economy (Jul 2026, companion note).

Comps & competitive threat: Nebius Q1 2026 shareholder letter; Oracle RPO disclosures; Meta AI-cloud report and neocloud selloff coverage (Jul 2026); Nvidia partnership/financing announcements (2026).

Data current as of 27 July 2026; private-round figures are as press-reported. All public sources; no confidential information. Companies referenced by archetype are illustrative of the subsector, not recommendations. Not investment advice.