AI Token Prices Collapse Nearly 60%, Raising Questions About the $600 Billion AI Boom
AI is getting cheaper at a speed that could reshape the economics of the biggest infrastructure buildout in tech history.
Silicon Data’s LLM Token Expenditure Index fell to $0.9665 per million tokens on August 31, down more than half from its May peak of $2.0651. Seeking Alpha analyst Damir Tokic, citing an earlier peak near $2.20, puts the decline at nearly 60% and argues the drop could be an early warning sign for hyperscaler AI spending.
“The LLM Token Expenditures Index has collapsed by almost 60% in Q2, and this might force a stall or cut in AI capex as profit margins for hyperscalers shrink,” Seeking Alpha wrote.

The plunge comes at an awkward moment for the AI industry. Alphabet, Amazon, Meta and Microsoft are spending hundreds of billions of dollars on GPUs, data centers, networking equipment and energy infrastructure, betting that demand for AI computing will keep climbing for years.
Yet the price companies are effectively paying for AI output is moving sharply in the opposite direction.
That does not mean AI demand has collapsed. Silicon Data’s Token Expenditure Index measures the usage-weighted effective price of one million LLM tokens, rather than total token consumption or aggregate AI spending. A falling index tells us that intelligence is getting cheaper to consume. It does not mean people are consuming less of it.

That distinction may decide whether the AI spending boom ends in a painful reckoning or becomes one of computing’s biggest economic breakthroughs.
Why AI token prices are collapsing
The simplest explanation is competition.
Companies no longer have to send every task to an expensive frontier model from OpenAI, Anthropic or Google. Smaller models can handle classification, summarization, extraction, and many routine business tasks at a fraction of the cost. Open-weight models have improved enough that companies can route simpler workloads away from premium models and reserve costly frontier systems for jobs that truly need them.
That behavior is already showing up in the market. Reuters reported in June that businesses were shifting toward cheaper and smaller models as AI bills climbed, with the share of open-model tokens on one major routing platform rising from 34% in January to 65% in June. Model routers can automatically send each request to the cheapest model capable of doing the job.
Chinese AI developers have intensified that pricing pressure. Models from DeepSeek and other labs have pushed capable inference into price ranges that would have looked extraordinary a few years ago. At the same time, companies can increasingly run open-weight models on their own hardware rather than paying a premium every time an employee, application, or AI agent sends a request to a proprietary API.
Why AI token prices are collapsing
Model routing may be accelerating the price decline. OpenRouter, which Stripe acquired last month for more than $7 billion, gives 8 million users access to more than 400 AI models through a single platform. Part of its appeal is eliminating model lock-in, allowing developers to shift workloads among providers without rebuilding their applications around each model. That makes it easier to reserve expensive frontier models for difficult tasks and send routine requests to cheaper alternatives.
That behavior matters for Silicon Data’s benchmark because the index is usage-weighted. Token prices can fall even without sweeping price cuts from OpenAI, Anthropic or Google if developers simply route a growing share of their workloads to cheaper models. In that sense, falling token prices may reflect a more competitive and efficient AI market rather than weakening demand.
Silicon Data’s methodology makes that shift especially interesting. The benchmark tracks more than 200 open models and 120 proprietary models using pricing and consumption data from more than 20 sources. Prices are weighted by actual market activity. A major AI provider does not have to announce a price cut for the index to fall. The benchmark can drop simply as customers migrate toward cheaper models.
That makes the chart less a story about a sale at OpenAI or Anthropic and more a story about commoditization.
AI models are becoming more capable at the same time that buyers are becoming less willing to pay premium prices for every token. The token-price collapse might not be evidence that AI is failing. It could partly be evidence that the AI model market is finally becoming a real market.
The $600 billion question
That price compression is arriving during an extraordinary spending cycle.
Reuters estimated earlier this year that Alphabet, Amazon, Meta and Microsoft were on track to pour roughly $600 billion into AI in 2026. Another estimate put their total capital spending closer to $650 billion. The money is flowing into data centers, accelerators, servers, networking gear, and the infrastructure required to keep enormous AI workloads running.
Here is where the economics become uncomfortable.
AI infrastructure spending does not translate dollar for dollar into token sales. Data centers support training, cloud computing, storage, databases, and many other services. Still, inference is one of the main ways costly AI infrastructure becomes something customers can buy.
If the effective selling price of that output keeps falling, usage has to make up the difference.
A simplified version of the equation is:
token price × token volume = revenue
The first part of that equation is falling hard. The entire AI investment thesis increasingly depends on what happens to the second.
Signs suggest the spending race is already pressuring cash generation. A Reuters analysis in July found that Microsoft, Alphabet, Amazon, Meta and Oracle are on a trajectory where their combined capital expenditures could exceed their combined free cash flow by 2027. That does not mean the hyperscalers are running out of money. Several remain among the strongest cash-generating companies ever built. It does mean investors are asking harder questions about the return each new dollar put into AI infrastructure generates.
Borrowing costs make that calculation tougher. Global bond yields have moved higher, increasing financing costs at the same time tech companies are leaning more heavily on debt and long-term infrastructure commitments.
The bearish interpretation is easy to see: falling token prices squeeze the revenue side just as rising capital costs squeeze the investment side.
Yet another possibility may be the most important part of this story.
Cheap AI could create so much new demand that falling token prices become fuel for the next phase of the boom rather than the trigger for its collapse.
Why collapsing token prices could be bullish for AI
There is a very different way to read the same chart.
Falling token prices may pressure model providers, yet cheaper AI can make far more applications economically viable. The cost of inference has become one of the main constraints on AI agents, coding systems, research tools and enterprise automation. Cut that cost sharply, and developers can afford to let software think longer, call models more frequently, and take on jobs that were too expensive months earlier.
That matters far more in the agentic era than it did during the chatbot era. A person asking ChatGPT a question might trigger a relatively small amount of inference. An AI coding agent working for hours can read files, generate code, inspect its own output, call tools, correct errors and repeat that cycle many times. The economic unit starts shifting from “one prompt” to an entire machine-executed workflow.
The scale of consumption is already striking. OpenRouter, the AI model marketplace Stripe agreed to acquire in August, is processing more than 10 trillion tokens per day across more than 400 models, according to Reuters. The platform serves more than 10 million developers and companies, giving a glimpse of how large multi-model consumption has become.
That creates the central counterargument to the AI-bubble thesis. A collapsing price per token does not automatically produce collapsing AI revenue or lower infrastructure demand. The result depends on price elasticity, or how much additional consumption appears after costs fall.
Computing history offers plenty of precedent. Semiconductor costs fell and the number of computers exploded. Storage prices crashed and companies began keeping enormous amounts of data. Bandwidth became cheaper and internet traffic surged far beyond what early networks were built to carry.
AI could follow a similar path. Cheap tokens could encourage software developers to embed intelligence into products where using a frontier model once made little financial sense. AI agents could run continuously instead of waiting for a human prompt. Businesses could process every document, customer interaction, software log, or internal workflow through models rather than reserving AI for a narrow set of high-value tasks.
Current infrastructure demand does not yet look like a market preparing for a sudden stop. Anthropic recently committed tens of billions of dollars to additional compute after demand for Claude and Claude Code surged beyond its earlier expectations. Broadcom raised its longer-term AI chip sales forecast this week, citing continued infrastructure commitments from major AI customers.
That does not invalidate the warning from token prices. It changes the question.
The issue is no longer simply whether AI output is becoming cheaper. It is whether consumption can grow faster than prices fall.
Q3 becomes the real test
The next earnings cycle should give investors a much clearer view of that equation.
Silicon Data’s token benchmark fell sharply after its late-May peak. The index stood at $1.06 per million tokens as of August 28, with Silicon Data describing it as a usage-weighted measure of what the active LLM market is effectively paying. The benchmark can fall through direct price cuts, customer migration to cheaper models, or changes in the mix of models being consumed.
That makes Q3 particularly revealing. Investors will get another look at cloud growth, AI revenue, infrastructure utilization, free cash flow and, perhaps most telling, what executives say about returns on AI spending.
Capex guidance may matter more than almost any other number.
A slowdown would give credibility to the bearish argument that falling inference prices are beginning to weaken the economics behind the infrastructure race. Another round of increases would suggest hyperscalers still see enough demand ahead to justify spending through the price compression.
Financing adds another layer of pressure. American hyperscalers have been tapping global bond markets more aggressively to fund data-center construction. Reuters reported this week that U.S. tech giants have issued roughly €40 billion in euro-denominated debt, with the European Central Bank watching the growing presence of Big Tech borrowers in Europe’s credit markets. Estimates cited by Reuters put AI-related capital expenditure at $5 trillion to $7 trillion through 2030.
Investors are effectively watching two races at once.
AI companies must drive token consumption upward faster than falling prices erode revenue per unit. At the same time, the returns generated from that consumption must justify the enormous amount of capital being committed to the infrastructure underneath it.
That is why the collapse in token prices deserves attention without automatically declaring the AI boom over.
Cheap intelligence could become one of the greatest demand catalysts the technology industry has ever seen. It could just as easily expose a brutal economic mismatch between what customers are willing to pay for AI and the cost of building the infrastructure needed to deliver it.
The threat to the $600 billion AI boom isn’t simply that intelligence is getting cheap. It’s that intelligence could get cheap faster than the industry can find profitable ways to consume it.

