08/07 2026
377

▲ All images in this article are sourced from the internet. Please contact us for removal if any infringement occurs.
Even the Giants Can't Sustain Endless Spending
This article was first published in Shadow Memo by Mo Yingsheng.
The story starts with an internal email.
In early August 2026, Jay Parikh, Microsoft's Executive Vice President, circulated an internal memo to all staff. The essence: starting from July 2026, every business unit at Microsoft must adhere to "AI token budget targets," with employees able to track their AI expenditures in real-time via an internal dashboard. The default internal model was also switched from previous options to OpenAI's GPT-5.6 for one primary reason: cost efficiency.
Parikh included a thought-provoking line: "Tokenmaxxing is not our optimization goal."
What is Tokenmaxxing? The term combines "token" and "maxxing" (pushing to extremes), following the same internet slang logic as gymmaxxing (obsessive fitness training) or looksmaxxing (extreme physical transformation).
However, it describes a perverse incentive structure where AI usage becomes a Key Performance Indicator (KPI) within companies.
In simpler terms, employees use AI for the sake of using AI, rapidly depleting tokens.
According to Microsoft's internal data, many engineers spend hundreds to thousands of dollars monthly on tokens.
For one person, burning through thousands each month is just the cost side. At the corporate level, the numbers become staggering.
More amusingly, Microsoft CEO Satya Nadella admitted on a podcast in June 2026 that he himself was a Tokenmaxxer, finding the rising numbers on the usage dashboard addictive.
But CEOs ultimately focus on numbers. He added: "The marginal benefit of productivity gains must match the marginal cost of tokens. Not every problem warrants deploying the strongest, most expensive cutting-edge models."
Two months later, the CEO's podcast remarks translated into company policy.
Here we have a $3.5 trillion behemoth that just invested $13 billion in OpenAI, now instructing employees to ration AI usage—a scene tinged with dark humor.
Some Microsoft employees vented internally: "A company heavily invested in AI, subsidizing massive inference workloads, is now coaching its own staff to save money? If even AI infrastructure hosts worry about affording internal AI products, what hope do regular enterprises buying these services have?"
That question hits the mark.

Burning Through Annual Budgets in Four Months—Who Can Keep Up?
Microsoft isn't the first to hit the brakes, nor will it be the last.
In April 2026, Uber's CTO circulated an internal memo revealing the company had exhausted its entire 2026 AI budget in just four months, largely due to the widespread adoption of Claude Code internally.
Among 5,000 engineers, monthly usage rates ranged from 84% to 95%, with per-capita bills ranging from $150 to $2,000.
The CTO's own two-hour internal demo reportedly consumed $1,200 worth of tokens. Uber COO Andrew Macdonald was "literally speechless" upon learning this figure.
Notably, Uber is a company valued at over $200 billion. If even such a behemoth exhausts its AI budget in four months, imagine the plight of SMEs.
Meta faces similar woes. Internal projections show that if current AI usage rates persist, Meta's internal AI expenses alone could reach billions in 2026. Meta circulated a memo to roughly 6,000 core employees imposing company-wide token quotas.
Amazon executives publicly urged staff to "avoid using AI for its own sake," shifting KPIs from token consumption to standardized business deliverables.
Amazon also disclosed a specific case: a Claude Sonnet-based system matching authors with product pages cost $1.8 million, far exceeding budgets, only to handle trivial tasks.
Any CFO would balk at $1.8 million for menial work.
Another unnamed tech giant forgot to cap employees' Claude licenses, racking up $500 million in AI spending in one month. Social media speculation nearly unanimously points to Amazon.
From Uber to Meta to Amazon to Microsoft, Silicon Valley giants are collectively slamming on the brakes. This isn't an isolated issue but a systemic industry crisis erupting en masse.

Why Are Tokens Getting More Expensive Despite Falling Unit Prices?
Here lies a counterintuitive phenomenon: token unit prices drop, yet corporate AI bills soar.
In 2024, foundational model vendors burned cash to train larger models, flooding markets with free tokens and rock-bottom prices to seize market share. Some even quipped, "Selling tokens is less profitable than selling bottled water," with input tokens priced at cents per million, or $1-2 for pricier models. The industry wallowed in narratives of "ever-declining costs."
This held true initially. Token prices for mainstream models plummeted up to 98% from early 2023 levels.
But usage exploded.
OpenRouter statistics show global weekly token consumption skyrocketed from 2.1TB to 24.5TB over the past year, with a 280% YoY surge since 2026.
In China, daily token calls leaped from 100 billion in early 2024 to 140 trillion by March 2026—a 1,000-fold increase in two years.
Unit prices dropped 98%; usage surged 1,000-fold. Elementary math shows total spending can only rise. Even minor unit costs, multiplied by massive scale, yield astronomical totals.
More critically, the rise of AI agent applications has multiplied per-task token consumption by tens of times.
Previously, having AI write code consumed predictable token volumes. Now, an AI agent executing a task autonomously might invoke models dozens or hundreds of times, spiking token use exponentially.
Supply-side constraints are equally dire. Global Blackwell chip computing power grows ~3.4x annually, while token demand grows ~10x yearly.
The 3.4x vs. 10x gap widens annually. HBM memory, monopolized by Samsung, SK Hynix, and Micron, requires 24-36 months for capacity expansion.
The computing rental market tightened accordingly. Renting NVIDIA's cutting-edge B200 chips doubled to nearly $6/hour.
With demand surging and supply choked, how could tokens not become expensive?
Expensive tokens might be acceptable if output justified costs. But where's the output?
Developer productivity platform Entelligence.AI aggregated data from 2,444 enterprises, revealing a sobering statistic: for every $1 spent on AI tokens, only 18 cents generates tangible user-facing value.
44 cents fix AI-induced bugs, 27 cents go to rework, and 11 cents are lost to review friction.
In other words, less than 20 cents of every dollar creates value. The remaining 80+ cents go toward "cleaning up AI's mess."
Uber's COO put it bluntly: "The correlation between token consumption growth and actual product improvements—that line doesn't exist yet."
This is the crux.
Enterprises initially pushed AI adoption to replace repetitive human labor, cut personnel costs, and seize AI transformation dividends.
Instead, they've incurred unpredictable computing bills without verifying whether token investments drive business growth.
Worse, current foundational models poorly align with niche business scenarios, often requiring employees to double-check outputs. Staff transition from "producers" to "reviewers," failing to truly liberate human resources.
Most enterprises use AI superficially—for copywriting or coding assistance—without building closed-loop AI business systems.
Even Microsoft ($13B invested), Uber ($200B+ valuation), and Amazon (~$200B annual capex) are slamming on the brakes when numbers don't add up. This isn't a technical issue—it's commercial.

More Users, Deeper Losses
Speaking of Microsoft, we must mention Copilot.
Microsoft initially charged fixed fees for Copilot: $10/month for Pro (using top models), $39/month for Pro+.
But AI agents increased data throughput and inference costs exponentially. Those fixed fees couldn't cover costs. More users meant deeper losses for Microsoft.
In early June 2026, Microsoft shifted to usage-based pricing beyond certain thresholds.
The global developer community howled: engineers heavily reliant on top-tier AI models exhausted basic quotas in days, with AI costs spiking from $39 to hundreds or even thousands monthly.
GitHub Copilot also switched to pay-per-use. An official discussion post drew nearly 900 downvotes, with users calculating that a single agent programming session typically consumed $30-40 in tokens. A $10/month plan could be exhausted in one session.
Even Microsoft balked at its employees' token bills, imposing quotas. What about enterprise clients paying for Copilot? Their bills look even worse.
This creates a vicious cycle: AI firms dare not charge high prices for fear of losing clients, but low prices can't cover costs. More users mean deeper losses.
Selling tokens isn't a guaranteed profit machine. When computing costs are fixed but revenue variable, every order gambles on whether clients will max out usage.
Wall Street is already pivoting. Soaring token costs pose the ultimate valuation killer for star AI firms eyeing IPOs by late 2026.
Public market investors scrutinize gross margins and profit paths mercilessly. When core input costs climb 40% semi-annually, who can smile?
You might think: "If tokens are too expensive, just cut prices."
Indeed, price cuts are happening. On July 30, 2026, OpenAI suddenly slashed GPT-5.6 Luna's prices by 80%, reducing output costs from $6 to $1.2 per million tokens.
Terra cut prices by 20%. DeepSeek was more aggressive, reducing V4 series token prices to one-fourth of original tags. Increased competition and falling costs triggered a wave of token price cuts.
But can price cuts solve the problem?
Short-term, they alleviate corporate clients' cost anxiety. Long-term, as long as token consumption grows exponentially and AI agents keep multiplying per-task token use, price-cut dividends will quickly be devoured by usage.
More critically, price cuts squeeze AI vendors' profit margins. OpenAI disclosed 2026 computing costs would hit $50 billion. Google, Microsoft, Amazon, and Meta are projected to spend $725 billion on AI in 2026, up ~77% YoY.
With massive investments on one side and relentlessly falling revenues on the other, profit margins are crushed from both ends.
Currently, Anthropic is the only leading AI firm likely to achieve quarterly operating profits. Most others operate at a loss.
This is the token economy's dilemma: charge high prices and clients can't afford it; charge low prices and you can't sustain it.
Tokens are sliding toward zero margins along the IaaS path. General-purpose inference tokens are rapidly commoditizing, with profits thinning.

Final Thoughts
In May 2026, Microsoft revoked Claude Code internal licenses for most employees. Nearly 100,000 engineers were ordered to migrate from Claude Code to Microsoft's own GitHub Copilot CLI by June 30. The reason was blunt: bills were too high.
Note the timeline: Claude Code revoked in May, token budget targets imposed in August. Within three months, Microsoft struck twice at AI cost control. What does this signal? That problems didn't erupt suddenly but accumulated to a breaking point.
Digging deeper, Microsoft's move had an unspoken motive: Claude Code had become too popular internally, overshadowing Microsoft's own GitHub Copilot CLI.
Before Claude Code access, 91% of Microsoft's engineering teams used GitHub Copilot. After Claude Code opened, that share was "severely eroded."
On one hand, their own products are being outperformed by competitors, and on the other hand, competitors charge per Token, requiring payment to Anthropic for every use.
Microsoft has invested $13 billion in OpenAI and built most of the computing power infrastructure for Anthropic on Azure. Yet, when their own engineers extensively use Claude Code, it's akin to financially supporting their most direct competitor.
After reviewing the bills, this tech giant with a market capitalization of $3.5 trillion decided it would no longer play the 'fool.'
This goes beyond cost issues—it has risen to a strategic level.
However, even for GitHub Copilot, which operates on 'internal cost accounting,' no matter how low its marginal costs are, it cannot withstand massive consumption. Hence, the Token budget target set in August.
All of this points to a harsh conclusion: Even tech giants can't make the Token business work.
It's not due to inadequate technology or poor products; it's that the economics simply don't add up. For every dollar invested, only 18 cents generate real value; the more users there are, the greater the losses; unit prices have dropped by 98%, but total bills have multiplied several times over.
These three facts together create the most awkward situation for the Token economy.
IDC predicts that by 2026, Token consumption in China's MaaS market will reach 40,000 trillion, growing about 20-fold from 2025.
Usage is still rising—and rapidly. But the question is, who will foot the bill? Giants are already hitting the brakes. What about small and medium-sized enterprises? Individual developers?
As a core resource in the AI era, Tokens are equivalent to internet traffic in the digital age or energy in the industrial age.
These essential production factors must be affordable and accessible to all, not constituting a major operational cost for businesses. But the reality is that Tokens are not yet that cheap.
Perhaps in a few years, with increased chip production capacity, improved model efficiency, and enhanced infrastructure, Tokens will truly become negligible in cost. But for now, even Microsoft, which sells AI, can't bear it anymore.
This is by no means a failure of AI; rather, it is simply commercial logic playing out its role. Any technology, as it makes the leap from the laboratory to widespread commercialization, is bound to go through a phase of rebalancing costs and value, and AI is presently navigating through this challenging phase.
The presence of token limits and the seemingly sluggish progress of AI do not signal its failure; instead, they serve as valuable lessons in economic reality, imparting wisdom to all.