The model labs, the chipmakers, and the investors who back them are all watching the same number fall. The price of a token — the atomic unit an AI model reads and writes — has dropped about a thousandfold between 2023 and 2026. The reflex is to treat that as unambiguous good news: intelligence is becoming free, so anyone building on top of it wins.
The reflex is wrong, and the number that matters is the one going up. In May 2026, Google alone reported processing about 3.2 quadrillion tokens a month — seven times as many as a year earlier (Google I/O 2026). And over the same period that per-token prices collapsed, enterprise spending on generative AI more than tripled in a single year, from $11.5 billion in 2024 to $37 billion in 2025 (Menlo Ventures). Both facts — price down a thousandfold, spending up threefold — are true at once, and reconciling them is the whole game. You cannot value the AI economy, or decide where in it to put capital, on the price of a token, because price is only one leg of the model. This is a demand story wearing a deflation costume.
We build alongside founders and co-invest with capital partners in AI companies, so this is not an academic question for us. It is the question we answer before we write a cheque: when the unit cost of intelligence falls this fast, where does the value actually go? Here is the model we use, and the verdict it produces.
Why the token is the right unit
Start with the unit itself, because it is what makes AI newly legible as an economy. A token is a chunk of text — very roughly three-quarters of a word — and every interaction with a large language model is metered in tokens read (input) and tokens written (output). A sentence is tens of tokens; a document, thousands; an autonomous "agent" grinding through a task, millions.
That means machine intelligence, for the first time, has an atomic unit of consumption with a published price — the way electricity has the kilowatt-hour and data has the gigabyte. And once something is metered, it can be modelled: total spend is simply the price per token multiplied by the number of tokens consumed. Almost all the public commentary obsesses over the first term. The money is in the second.
The supply side: intelligence is getting cheaper — but not everywhere
The deflation is real and historically fast. For a fixed level of capability, the cost of inference has fallen roughly 10× per year for three years running — what a16z labelled "LLMflation." Measured against specific capability benchmarks, Epoch AI puts the fall at anywhere from 9× to 900× per year depending on the task, with a median around 50× and the frontier of hardest problems deflating fastest. The drivers are stacked and durable: better silicon, more efficient algorithms, brutal competition, and a wave of open-weight models that reset the price floor.
But that headline — intelligence roughly 1,000× cheaper — hides the single most important nuance for an investor, and Exhibit 1 makes it visible: the frontier barely deflates; only the commodity tier collapses.
This matters because the workloads that actually generate return — the reasoning and agentic tasks we come back to below — run on the frontier that isn't getting cheaper. The near-free tokens are last year's capability. Any business plan that assumes it will ride commodity pricing into fat margins is quietly betting that its hardest, most valuable work can run on the cheap tier. It usually can't.
The demand side: consumption is exploding
Now the leg everyone underweights. As price falls, the number of tokens consumed is rising far faster — and it is doing so for a structural reason, not a cyclical one. Two forces compound.
The first is a shift in what a task costs. A single-turn chat answer is a few hundred tokens. But the industry has moved decisively toward models that "think" before they answer, and toward agents that loop through a task autonomously. Reasoning tokens crossed 50% of all output tokens in 2025; agentic workloads consume anywhere from 10× to 100× the tokens of a simple chat. A single deep-research-style run can burn 300,000 tokens. Exhibit 2 shows the multiplier.
The second force is price elasticity, and it is the linchpin of the entire model. Demand for tokens is not merely elastic; it is super-elastic. Simulation work puts the elasticity above 1 — roughly a 1% price cut drives a 1.4% rise in volume — which means that as the unit price falls, total dollar spend rises rather than falls. The real-world evidence is blunt: inference got about 1,000× cheaper while demand rose an estimated 10,000×, and enterprise AI spend jumped 320% in 2025 even as per-token prices dropped. This is Jevons' paradox — the observation that making a resource cheaper to use increases, not decreases, total consumption of it. Cheaper tokens don't shrink the bill. They enlarge the market.
Two token economies
Consumption is not evenly distributed, and the geography is itself an allocation decision. Two distinct token economies have emerged, climbing the same demand curve from opposite ends.
| United States — frontier | China — commodity / open-weight | |
|---|---|---|
| Model posture | Closed frontier labs; per-token API + subscriptions | Open-weight, deliberate race to zero |
| Representative price ($/M input) | ~$5.00 (frontier flagship) | ~$0.14 — a ~35× gap |
| Scale of consumption | OpenAI ~15B tokens/min (~650T/mo); MS Foundry >500T in FY2025 | National >140T tokens/day; ByteDance Doubao alone >120T/day |
| Growth driver | Capability and agentic depth | Price → mass, embedded super-app usage |
| How it monetises | Token margin + subscriptions | Cloud/infra + ecosystem lock-in, not token margin |
The United States sells fewer, dearer, higher-value tokens: closed frontier labs monetising capability directly. China manufactures vastly more, near-free, commoditised tokens, using open weights as a strategic vehicle and monetising through cloud, infrastructure, and super-app lock-in rather than the token itself. The most under-priced fact in the table is the footnote: on neutral ground, Chinese open-weight models have already overtaken US models by volume, and are now serving nearly half of US-enterprise usage for cost-sensitive work. Frontier margins have not yet met a 35×-cheaper substitute at full force. They will.
The model: a market that grows in dollars as it falls in price
Put the legs together. Price per token is collapsing (Exhibit 1). Tokens per task and tokens overall are exploding (Exhibit 2), with demand elastic enough that lower prices raise total revenue. The product of the two — the actual size of the token economy in dollars — is therefore growing, and the growth curve is young. Exhibit 4 is the scale check.
| Signal | Level (2026) | Growth |
|---|---|---|
| Google — tokens processed / month | 3.2 quadrillion | ~330× since Apr 2024 (all surfaces) |
| OpenAI — API tokens / minute | 15 billion | ~2.5× in six months |
| China — national tokens / day | >140 trillion | >1,000× since early 2024 |
| Enterprise AI spend | +320% (2025) | while unit prices fell |
| Goldman forecast — tokens / month by 2030 | ~120 quadrillion | ~24× from ~2026 |
So the first-order conclusion is bullish for the market and bearish for per-unit margin: the token economy compounds in dollars even as each token races toward free. But total market size is not where an investor lives. The question is who keeps the money.
Where the money pools
Deflation does not distribute its gains evenly. It hands them to whoever owns something scarce, and it strands whoever owns something substitutable. Follow the gross margin down the stack.
The chip layer keeps roughly three-quarters of every dollar because compute is the genuinely scarce input, and the whole 24× demand curve flows through it. The model labs are improving their margins — and, tellingly, doing so partly by competing with their own customers, absorbing application features into the model itself. The application layer inherits the mirror image: token spend becomes a permanent variable cost of goods sold, dragging gross margins to 50–60% against the 80–90% that classic software enjoyed. "Inference is the new COGS" is not a slogan; it is a structural haircut on every AI business that resells someone else's intelligence.
This is also why pricing is changing shape under everyone's feet. Per-seat software pricing — charge per user, serve them at near-zero marginal cost — breaks when every active user carries a real, rising token bill. The market is migrating to consumption and outcome pricing precisely because the cost base now scales with usage. For an investor, a per-seat AI business is carrying an un-priced margin risk.
The investor's verdict
A model is only useful if it tells you what to do. Here is the whole verdict in one sentence: an AI application that owns no proprietary data, no defensible workflow, and no claim on the outcome is renting intelligence and reselling it at a shrinking margin — and that business does not fly. Everything else is a corollary. Durable margin belongs to whoever owns the scarce input — frontier compute, proprietary data, or the outcome — not to whoever writes the prompt.
Invest in — where you own the scarce input:
- Frontier compute, and the challengers to its pricing power. The layer that doesn't deflate and that all demand flows through. That includes the custom-silicon threat to the incumbent's ~73% margin, not only the incumbent.
- Inference infrastructure — the picks and shovels of the demand explosion. Serving, routing, caching, and optimisation monetise elastic volume growth regardless of which model wins.
- Proprietary-data and workflow-moat applications. Apps whose defensibility is data or embedded workflow a frontier API cannot replicate keep pricing power as token cost falls.
- Outcome- and consumption-priced businesses. Models whose revenue rises with the token bill instead of being crushed by it.
Avoid — where you rent someone else's input:
- Thin "wrapper" applications with no data or workflow moat — a pass-through the labs are actively absorbing.
- AI apps underwritten to classic-SaaS margins. The real, structural number is 50–60%; a plan that assumes 85% is mispriced.
- Per-seat pricing on AI products — a dying model carrying un-priced repricing risk.
- Commodity model resellers with no cost or distribution edge; the China price war is their extinction event.
And a call the horse-race commentary misses entirely: treat the US frontier and the China commodity economies as distinct asset classes, underwritten on different logic — frontier exposure for capability-driven value, the open-weight stack for the volume-and-cost-floor play — and never assume frontier margins survive contact with a substitute priced 35× lower.
Coda
For our own field, financial services, the model resolves a debate that consumes too much airtime. The token bill for a bank deploying AI is a rounding error against the value at stake — gen-AI is credibly worth $200–340 billion a year to global banking. The binding constraint is not the price of intelligence; it is that most institutions are stuck at pilots, with only a fraction of adopted use cases actually in production. The stranded value is in deployment and governance, not in shaving cents off a token.
This is the first in a series in which we apply this model — to financial-services adoption, to where value accrues as agents proliferate, and to the businesses being built on both sides of the US–China divide. The through-line is the one principle worth carrying out of it: when something gets radically cheaper, the value does not vanish. It migrates. The builder-investor's job is to own where it lands.
