The model labs, the chipmakers, and the investors who back them are all watching the same number fall. The price of a token — the atomic unit an AI model reads and writes — has dropped about a thousandfold between 2023 and 2026. The reflex is to treat that as unambiguous good news: intelligence is becoming free, so anyone building on top of it wins.

The reflex is wrong, and the number that matters is the one going up. In May 2026, Google alone reported processing about 3.2 quadrillion tokens a month — seven times as many as a year earlier (Google I/O 2026). And over the same period that per-token prices collapsed, enterprise spending on generative AI more than tripled in a single year, from $11.5 billion in 2024 to $37 billion in 2025 (Menlo Ventures). Both facts — price down a thousandfold, spending up threefold — are true at once, and reconciling them is the whole game. You cannot value the AI economy, or decide where in it to put capital, on the price of a token, because price is only one leg of the model. This is a demand story wearing a deflation costume.

We build alongside founders and co-invest with capital partners in AI companies, so this is not an academic question for us. It is the question we answer before we write a cheque: when the unit cost of intelligence falls this fast, where does the value actually go? Here is the model we use, and the verdict it produces.

Why the token is the right unit

Start with the unit itself, because it is what makes AI newly legible as an economy. A token is a chunk of text — very roughly three-quarters of a word — and every interaction with a large language model is metered in tokens read (input) and tokens written (output). A sentence is tens of tokens; a document, thousands; an autonomous "agent" grinding through a task, millions.

That means machine intelligence, for the first time, has an atomic unit of consumption with a published price — the way electricity has the kilowatt-hour and data has the gigabyte. And once something is metered, it can be modelled: total spend is simply the price per token multiplied by the number of tokens consumed. Almost all the public commentary obsesses over the first term. The money is in the second.

The supply side: intelligence is getting cheaper — but not everywhere

The deflation is real and historically fast. For a fixed level of capability, the cost of inference has fallen roughly 10× per year for three years running — what a16z labelled "LLMflation." Measured against specific capability benchmarks, Epoch AI puts the fall at anywhere from 9× to 900× per year depending on the task, with a median around 50× and the frontier of hardest problems deflating fastest. The drivers are stacked and durable: better silicon, more efficient algorithms, brutal competition, and a wave of open-weight models that reset the price floor.

But that headline — intelligence roughly 1,000× cheaper — hides the single most important nuance for an investor, and Exhibit 1 makes it visible: the frontier barely deflates; only the commodity tier collapses.

Exhibit 1
Two speeds of deflation — the frontier holds its price while "good-enough" intelligence craters
2023 (GPT-4 launch) 2026
Frontier tier — $/M tokens (output)
Mar 2023 — GPT-4$60
2026 — frontier flagship~$15–25
~4× in three years. The best intelligence stays expensive.
Commodity tier — GPT-4-class quality
Mar 2023 — GPT-4$60
2026 — e.g. Flash / mini / open-weight~$0.15
~200–400× in three years. Yesterday's intelligence is nearly free.
Representative API prices, US dollars per million tokens. Frontier 2026 = flagship reasoning models (OpenAI / Anthropic / Google). Commodity = GPT-4-class output from small/fast or open-weight models. Sources: OpenAI, Anthropic, Google, DeepSeek pricing; a16z LLMflation; Epoch AI. Directional.

This matters because the workloads that actually generate return — the reasoning and agentic tasks we come back to below — run on the frontier that isn't getting cheaper. The near-free tokens are last year's capability. Any business plan that assumes it will ride commodity pricing into fat margins is quietly betting that its hardest, most valuable work can run on the cheap tier. It usually can't.

The demand side: consumption is exploding

Now the leg everyone underweights. As price falls, the number of tokens consumed is rising far faster — and it is doing so for a structural reason, not a cyclical one. Two forces compound.

The first is a shift in what a task costs. A single-turn chat answer is a few hundred tokens. But the industry has moved decisively toward models that "think" before they answer, and toward agents that loop through a task autonomously. Reasoning tokens crossed 50% of all output tokens in 2025; agentic workloads consume anywhere from 10× to 100× the tokens of a simple chat. A single deep-research-style run can burn 300,000 tokens. Exhibit 2 shows the multiplier.

Exhibit 2
The same question, many times the tokens — as work goes agentic, consumption per task explodes
Single-turn chat
Chat with conversation history~3–5×
Reasoning ("thinking") query~10–85×
Agentic workflow~10–100×
Deep-research run (~300k tokens)~300–600×
Tokens consumed per task, relative to a single-turn chat baseline (~500–1,000 tokens). Sources: OpenRouter State of AI; agent-workload analyses, 2026. Illustrative ranges.

The second force is price elasticity, and it is the linchpin of the entire model. Demand for tokens is not merely elastic; it is super-elastic. Simulation work puts the elasticity above 1 — roughly a 1% price cut drives a 1.4% rise in volume — which means that as the unit price falls, total dollar spend rises rather than falls. The real-world evidence is blunt: inference got about 1,000× cheaper while demand rose an estimated 10,000×, and enterprise AI spend jumped 320% in 2025 even as per-token prices dropped. This is Jevons' paradox — the observation that making a resource cheaper to use increases, not decreases, total consumption of it. Cheaper tokens don't shrink the bill. They enlarge the market.

Two token economies

Consumption is not evenly distributed, and the geography is itself an allocation decision. Two distinct token economies have emerged, climbing the same demand curve from opposite ends.

Exhibit 3
Two token economies — the same Jevons curve, approached from opposite ends
United States — frontierChina — commodity / open-weight
Model postureClosed frontier labs; per-token API + subscriptionsOpen-weight, deliberate race to zero
Representative price ($/M input)~$5.00 (frontier flagship)~$0.14 — a ~35× gap
Scale of consumptionOpenAI ~15B tokens/min (~650T/mo); MS Foundry >500T in FY2025National >140T tokens/day; ByteDance Doubao alone >120T/day
Growth driverCapability and agentic depthPrice → mass, embedded super-app usage
How it monetisesToken margin + subscriptionsCloud/infra + ecosystem lock-in, not token margin
The crossover: on the neutral OpenRouter platform, Chinese open-weight models overtook US models in weekly token consumption for the first time in February 2026, and by mid-2026 served up to ~46% of US-enterprise token usage. Sources: OpenAI, DeepSeek, Microsoft FY2025, National Data Administration, ByteDance, Dealroom. Figures per different disclosures/dates; treat as directional.

The United States sells fewer, dearer, higher-value tokens: closed frontier labs monetising capability directly. China manufactures vastly more, near-free, commoditised tokens, using open weights as a strategic vehicle and monetising through cloud, infrastructure, and super-app lock-in rather than the token itself. The most under-priced fact in the table is the footnote: on neutral ground, Chinese open-weight models have already overtaken US models by volume, and are now serving nearly half of US-enterprise usage for cost-sensitive work. Frontier margins have not yet met a 35×-cheaper substitute at full force. They will.

The model: a market that grows in dollars as it falls in price

Put the legs together. Price per token is collapsing (Exhibit 1). Tokens per task and tokens overall are exploding (Exhibit 2), with demand elastic enough that lower prices raise total revenue. The product of the two — the actual size of the token economy in dollars — is therefore growing, and the growth curve is young. Exhibit 4 is the scale check.

Exhibit 4
The demand curve is young — provider run-rates and the road to 2030
SignalLevel (2026)Growth
Google — tokens processed / month3.2 quadrillion~330× since Apr 2024 (all surfaces)
OpenAI — API tokens / minute15 billion~2.5× in six months
China — national tokens / day>140 trillion>1,000× since early 2024
Enterprise AI spend+320% (2025)while unit prices fell
Goldman forecast — tokens / month by 2030~120 quadrillion~24× from ~2026
Provider figures use different definitions of a "token" (Google's spans multimodal and internal surfaces; OpenAI's is paid API; China's is a government aggregate) and are not additive. Sources: Google I/O 2026; OpenAI; National Data Administration; Fortune; Goldman Sachs. Directional.

So the first-order conclusion is bullish for the market and bearish for per-unit margin: the token economy compounds in dollars even as each token races toward free. But total market size is not where an investor lives. The question is who keeps the money.

Where the money pools

Deflation does not distribute its gains evenly. It hands them to whoever owns something scarce, and it strands whoever owns something substitutable. Follow the gross margin down the stack.

Exhibit 5
Where the money pools — gross margin down the AI stack, 2026
Chips (Nvidia) — the scarce input~73%
Foundation-model labs (improving)~60%
AI applications — token cost is COGS~50–60%
Classic SaaS (benchmark)~80–90%
Where is electricity? It is not a separate bar — it sits inside the chip and cloud layers' cost of goods, and it is the fastest-rising part of it. That is why “compute scarcity” is turning into “power scarcity,” and why the data-centre energy build-out is itself becoming an investable layer.
Representative gross margins. Nvidia ~72–75% (FY2026 filings); model-lab inference margins improving toward ~70%; AI-application margins from Bessemer / industry surveys; classic SaaS as reference. “Inference is the new COGS.” Directional.

The chip layer keeps roughly three-quarters of every dollar because compute is the genuinely scarce input, and the whole 24× demand curve flows through it. The model labs are improving their margins — and, tellingly, doing so partly by competing with their own customers, absorbing application features into the model itself. The application layer inherits the mirror image: token spend becomes a permanent variable cost of goods sold, dragging gross margins to 50–60% against the 80–90% that classic software enjoyed. "Inference is the new COGS" is not a slogan; it is a structural haircut on every AI business that resells someone else's intelligence.

This is also why pricing is changing shape under everyone's feet. Per-seat software pricing — charge per user, serve them at near-zero marginal cost — breaks when every active user carries a real, rising token bill. The market is migrating to consumption and outcome pricing precisely because the cost base now scales with usage. For an investor, a per-seat AI business is carrying an un-priced margin risk.

The investor's verdict

A model is only useful if it tells you what to do. Here is the whole verdict in one sentence: an AI application that owns no proprietary data, no defensible workflow, and no claim on the outcome is renting intelligence and reselling it at a shrinking margin — and that business does not fly. Everything else is a corollary. Durable margin belongs to whoever owns the scarce input — frontier compute, proprietary data, or the outcome — not to whoever writes the prompt.

Invest in — where you own the scarce input:

  • Frontier compute, and the challengers to its pricing power. The layer that doesn't deflate and that all demand flows through. That includes the custom-silicon threat to the incumbent's ~73% margin, not only the incumbent.
  • Inference infrastructure — the picks and shovels of the demand explosion. Serving, routing, caching, and optimisation monetise elastic volume growth regardless of which model wins.
  • Proprietary-data and workflow-moat applications. Apps whose defensibility is data or embedded workflow a frontier API cannot replicate keep pricing power as token cost falls.
  • Outcome- and consumption-priced businesses. Models whose revenue rises with the token bill instead of being crushed by it.

Avoid — where you rent someone else's input:

  • Thin "wrapper" applications with no data or workflow moat — a pass-through the labs are actively absorbing.
  • AI apps underwritten to classic-SaaS margins. The real, structural number is 50–60%; a plan that assumes 85% is mispriced.
  • Per-seat pricing on AI products — a dying model carrying un-priced repricing risk.
  • Commodity model resellers with no cost or distribution edge; the China price war is their extinction event.

And a call the horse-race commentary misses entirely: treat the US frontier and the China commodity economies as distinct asset classes, underwritten on different logic — frontier exposure for capability-driven value, the open-weight stack for the volume-and-cost-floor play — and never assume frontier margins survive contact with a substitute priced 35× lower.

Coda

For our own field, financial services, the model resolves a debate that consumes too much airtime. The token bill for a bank deploying AI is a rounding error against the value at stake — gen-AI is credibly worth $200–340 billion a year to global banking. The binding constraint is not the price of intelligence; it is that most institutions are stuck at pilots, with only a fraction of adopted use cases actually in production. The stranded value is in deployment and governance, not in shaving cents off a token.

This is the first in a series in which we apply this model — to financial-services adoption, to where value accrues as agents proliferate, and to the businesses being built on both sides of the US–China divide. The through-line is the one principle worth carrying out of it: when something gets radically cheaper, the value does not vanish. It migrates. The builder-investor's job is to own where it lands.