08/07 2026
345

▲ All images in this article are sourced from the internet. Please contact us for removal if there is any infringement.
Even Giants Can't Sustain Endless Spending
This article was first published in Shadow Memo by Mo Yingsheng
The story begins with an internal email.
In early August 2026, Microsoft Executive Vice President Jay Parikh sent an internal letter to all employees. The gist was that starting from July 2026, all Microsoft business units would have 'AI Token budget targets,' and employees could view their AI expenditures in real-time on an internal dashboard. The default internal model was also switched from previous options to OpenAI GPT-5.6 for one reason only: it was cheaper.
Parikh wrote a line in the email worth pondering: 'Tokenmaxxing is not our optimization goal.'
What is Tokenmaxxing? This term combines 'token' and 'maxxing' (pushing to the extreme), following the same internet slang logic as gymmaxxing (intense fitness training) and looksmaxxing (extreme appearance modification).
However, it describes a distorted incentive mechanism where AI usage becomes a KPI within companies.
Simply put, employees use AI for the sake of using AI, Crazy brushing (frantically racking up) Tokens.
According to Microsoft's internal data, many engineers spend hundreds to thousands of dollars monthly on Tokens.
One person burning thousands of dollars a month is just the cost side. At the company level, the numbers are even more alarming.
More interestingly, Microsoft CEO Satya Nadella admitted on a podcast in June this year that he himself was a Tokenmaxxer, finding the soaring numbers on the usage dashboard addictive.
But CEOs ultimately have to do the math. He added: The marginal benefit of productivity gains must match the marginal cost of Tokens. Not every problem warrants using the strongest, most expensive cutting-edge models.
Two months later, the CEO's podcast remarks became company policy.
A $3.5 trillion giant that just poured $13 billion into OpenAI is now ordering its employees to use AI sparingly—a scene reeking of dark humor.
Some Microsoft employees complained internally: A company heavily investing in AI and subsidizing massive inference workloads is now guiding its employees to save money? If even companies hosting AI infrastructure worry about affording internal AI products, what should ordinary businesses buying these services do?
That question hits the mark.

Burning Through Annual Budgets in Four Months—Who Can Take It?
Microsoft isn't the first to hit the brakes, nor will it be the last.
In April this year, Uber's CTO sent an internal memo revealing the company burned through its 2026 AI budget in just four months, largely due to widespread adoption of Claude Code internally.
Among 5,000 engineers, monthly usage rates ranged from 84% to 95%, with per-capita monthly bills ranging from $150 to $2,000.
The CTO himself reportedly consumed $1,200 worth of Tokens during a two-hour internal demo. Uber COO Andrew Macdonald was 'literally speechless with shock' upon learning this figure.
Notably, Uber is a company valued at over $200 billion. If even such a company burns through its AI budget in four months, imagine the situation for SMEs?
Meta isn't faring much better. Internal projections show that if current AI usage rates persist, internal AI expenses alone will reach billions of dollars in 2026. Meta sent memos to about 6,000 core employees imposing company-wide Token quotas.
Amazon executives publicly warned employees to 'stop using AI for the sake of using AI,' shifting KPIs from Token consumption to standardized business deliverables.
Amazon also disclosed a more specific case: An author-to-product-page matching system built on Claude Sonnet ultimately cost $1.8 million, far exceeding the original budget. The system was only used to handle simple tasks.
$1.8 million for menial tasks—any CFO would have a heart attack at this ROI.
Another unnamed tech giant forgot to set usage caps on employees' Claude licenses, racking up $500 million in AI spending in just one month. Social media speculation almost unanimously points to Amazon.
From Uber to Meta to Amazon to Microsoft, Silicon Valley giants are collectively slamming on the brakes. This isn't an isolated issue for one or two companies but a systemic contradiction erupting across the industry.

Why Are Tokens Getting More Expensive Despite Falling Unit Prices?
Here's a counterintuitive phenomenon: While Token unit prices are dropping, companies' AI bills are rising.
In 2024, large model vendors burned money training bigger models, flooding the market with free Tokens and rock-bottom prices. Some even said 'selling Tokens is less profitable than selling bottled water,' with per-million-input Tokens costing just cents, or a dollar or two for pricier models. The industry was immersed in (immersed in) narratives of 'continuously declining costs.'
And indeed, mainstream large model Token prices have dropped by up to 98% since early 2023.
But the problem lies in usage volume.
According to OpenRouter, global weekly Token consumption surged from 2.1T to 24.5T over the past year, with year-to-date weekly growth reaching 280% in 2026.
In China, daily Token calls skyrocketed from 100 billion in early 2024 to 140 trillion by March 2026—a 1,000-fold increase in two years.
Unit prices dropped 98%, but usage surged 1,000-fold. Simple math shows total spending can only rise, not fall. Even minor unit costs, when multiplied by massive usage scales, result in astronomical total costs.
More critically, the explosion of AI agent applications has caused Token consumption per task to surge dozens of times over.
Previously, having AI write code consumed controllable amounts of Tokens. Now, when an AI agent autonomously executes a task, it may call models dozens or hundreds of times, resulting in exponential Token growth.
The supply side is equally strained. Global Blackwell chip computational power grows about 3.4x annually, while global Token demand grows about 10x yearly.
A 3.4x vs. 10x gap widens yearly. HBM high-bandwidth memory is monopolized by Samsung, SK Hynix, and Micron, with expansion cycles lasting 24-36 months.
The computational leasing market has tightened accordingly. Renting NVIDIA's cutting-edge B200 chips has doubled to nearly $6 per hour.
So you see—with demand surging and supply constrained, how can Tokens not become expensive?
If expensive Tokens delivered high output, it might be justifiable. But where's the output?
Developer productivity platform Entelligence.AI aggregated data from 2,444 enterprises, revealing alarming figures: For every $1 spent on AI Tokens, only 18 cents generates actual user-facing value.
44 cents go to fixing AI-introduced bugs, 27 cents to rework, and 11 cents to review friction.
In other words, throwing in a dollar yields less than 20 cents in value. The remaining 80+ cents go to 'cleaning up after AI.'
Uber's COO put it more bluntly: The correlation between Token consumption growth and actual product improvements 'doesn't exist yet.'
That's the core issue.
Companies initially pushed for enterprise-wide AI adoption to replace repetitive human work, cut labor costs, and seize AI transformation dividends.
But reality shows companies haven't reduced labor costs—instead, they've acquired unpredictable massive computational bills. Whether heavy Token investments translate to business growth remains unproven.
More awkwardly, current large models still can't deeply adapt to niche business scenarios, often requiring employees to double-check and correct outputs. Employees merely shifted from 'producers' to 'reviewers,' without truly liberating human resources.
Most companies also use AI only for superficial copywriting and coding assistance, failing to build complete AI business loops.
Microsoft, which invested $13 billion; Uber, valued at over $200 billion; and Amazon, with about $200 billion in annual capex—all these giants are hitting the brakes when costs spiral out of control. This isn't a technical issue but a commercial one.

More Users, Bigger Losses
Speaking of Microsoft, we can't ignore Copilot.
Microsoft originally charged fixed fees for Copilot users: $10 monthly for Pro, $39 for Pro+ with top-tier models.
But after AI agents emerged, data throughput and inference costs soared. Those Dozens of dollars (tens of dollars) in monthly fees couldn't cover costs. More users meant bigger losses for Microsoft.
So in early June 2026, Microsoft changed its policy to usage-based billing after certain thresholds.
The global developer community howled: Many engineers heavily reliant on top-tier AI models burned through basic quotas in days, with AI usage fees skyrocketing from $39 to hundreds or even thousands of dollars.
GitHub Copilot also shifted entirely to usage-based billing. An official discussion post received nearly 900 downvotes, with users calculating that a typical agent programming session consumed $30-40 in Tokens. This meant a $10 monthly plan could be exhausted in a single use.
Even Microsoft itself was shocked by employees' Token bills, forcing quotas. What about enterprise clients paying for Copilot? Their bills look even worse.
This creates a deadlock: AI companies dare not charge high prices for fear of scaring clients, but low prices can't cover costs—more users mean bigger losses.
Selling Tokens isn't a guaranteed profit. When computational power is fixed but revenue variable, every order gambles on whether clients will max out usage.
Wall Street has already shifted gears. Surging Token costs pose the biggest valuation threat to star AI firms planning IPOs by late 2026.
Public market investors scrutinize gross margins and profit paths mercilessly. When core input costs climb 40% every six months, who can smile?
You might think: If Tokens are too expensive, just lower prices.
Indeed, price cuts are happening. On July 30, 2026, OpenAI suddenly slashed GPT-5.6 Luna prices by 80%, reducing output costs from $6 to $1.2 per million Tokens.
Terra cut prices by 20%. DeepSeek was more aggressive, reducing V4 series Token prices to one-fourth of original tags. Competition and falling costs have triggered a wave of Token price cuts.
But can price cuts solve the problem?
Short-term, they ease cost anxiety for enterprise clients. Long-term, as long as Token consumption grows exponentially and AI agents keep driving per-task Token use skyward, price-cut dividends will quickly be devoured by usage volume.
More critically, price cuts squeeze AI vendors' profit margins. OpenAI disclosed 2026 computational expenses would reach $50 billion. Google, Microsoft, Amazon, and Meta are projected to spend $725 billion on AI in 2026—a 77% YoY surge.
With massive investments on one side and continuously falling revenue on the other, profit margins are crushed from both ends.
Currently, Anthropic is the only leading AI firm expected to achieve quarterly operating profits. Others generally operate at a loss.
This is the Token business's dilemma: Charge high prices and clients can't afford it; charge low and you can't sustain it.
Tokens are sliding toward zero margins along IaaS's old path. General-purpose inference Tokens are commoditizing, with profits thinning.

Final Thoughts
In May 2026, Microsoft revoked internal Claude Code licenses for most employees. Nearly 100,000 engineers were ordered to migrate from Claude Code to Microsoft's own GitHub Copilot CLI by June 30. The reason was blunt: Bills were too high.
Note the timeline: Claude Code revocation in May, Token budget targets in August. Within three months, Microsoft struck twice at AI cost control. What does this signal? That the problem didn't erupt suddenly but accumulated to a breaking point.
Looking deeper, Microsoft's Claude Code cancellation had an unspoken reason: Claude Code was too popular internally, overshadowing Microsoft's own GitHub Copilot CLI.
Before granting Claude Code access, 91% of Microsoft's engineering teams used GitHub Copilot. After Claude Code became available, this usage rate was 'severely eroded.'
On one hand, their own products are being outperformed by competitors, and on the other hand, competitors charge by Token, requiring payment to Anthropic for every use.
Microsoft has invested $13 billion in OpenAI and has built most of the computing power infrastructure for Anthropic on Azure. However, when their own engineers extensively use Claude Code, it's equivalent to financially supporting their most direct competitor.
After reviewing the bills, this tech giant with a market capitalization of $3.5 trillion decided it would no longer play the 'fool.'
This issue goes beyond cost; it has risen to a strategic level.
However, even for GitHub Copilot, which operates under 'internal cost accounting,' the marginal costs, no matter how low, cannot withstand massive consumption. Hence, the Token budget target was set in August.
All of this points to a harsh conclusion: Even tech giants can't make the Token business work.
It's not due to inadequate technology or poor products; it's because the economics don't add up. For every dollar invested, only 18 cents generate real value; the more users there are, the greater the losses; unit prices have dropped by 98%, but total bills have multiplied several times over.
These three facts combined create the most awkward situation for the Token economy.
IDC predicts that by 2026, Token consumption in China's MaaS market will reach 40,000 trillion, growing approximately 20-fold from 2025.
Usage is still increasing, and rapidly. But the question is, who will foot the bill? Tech giants are already hitting the brakes—what about small and medium-sized enterprises? What about individual developers?
Token, as a core resource in the AI era, is equivalent to internet traffic in the digital age and energy in the industrial age.
These production factors must be affordable and accessible to all; they should not constitute a major operational cost for businesses. However, the reality is that Tokens are not yet that cheap.
Perhaps in a few years, with increased chip production capacity, improved model efficiency, and enhanced infrastructure, Tokens will truly become so inexpensive that they can be overlooked. But for now, even Microsoft, which sells AI, can't bear the burden.
This is not a failure of AI; it's the commercial laws at work. Any technology transitioning from the laboratory to large-scale commercialization must undergo a process of realigning costs and values, and AI is currently experiencing this growing pain.
Token limits and slow AI progress are not signs of AI's failure; they are lessons in economic laws for everyone.