09/22 2026
492


The AI Price War is Entering Its Next Phase
Image source | Internet (Please contact us for deletion if infringement occurs). Partially generated by AI
If you are a developer or business decision-maker keeping an eye on the AI industry, you may have noticed a subtle shift: Over the past year, those enticing "rock-bottom prices," "free access," and "unlimited calls" that once excited you are quietly fading away.
Large model vendors are no longer competing to be the cheapest. Instead, they are engaging in a more "mature" practice—tiered pricing.
On the usage side, the boldness of companies encouraging employees to "use AI without restraint" is being replaced by cost-benefit analyses in Excel spreadsheets.

The "Commodity Price" Era of AI
Let's rewind the clock to just over a year ago.
From 2024 to the first half of 2025, the domestic large model market experienced an unprecedented wave of price competition. Alibaba Cloud slashed the input price of its Tongyi Qianwen flagship model from 0.02 yuan to 0.0005 yuan, a 97% reduction; Baidu made its two main Wenxin large models fully free; Tencent reduced prices for its Hunyuan large model by up to 87.5%.
DeepSeek entered the scene as a "cost killer," driving API prices to astonishingly low levels.
The logic at the time was simple: Capture developers, secure entry points, and cultivate user habits.
Leading vendors actively lowered API pricing to preempt (seize) developer traffic entry points, using low prices to build their own ecological barriers.
Meanwhile, a farcical scene unfolded on the usage side. Within Silicon Valley tech giants, management rewarded employees heavily reliant on AI, even using token call volume as a key performance metric.
Weekly calls of 210 billion tokens, monthly spending of $150,000...
These extreme cases reflected programmers' "token anxiety," with many brushing up call volumes uncontrollably to climb internal rankings.
This frenzy also swept through China. Many internet companies proudly announced their All in AI strategies, granting employees unlimited call privileges.
Rapidly expanding demand drove token prices higher. According to the China Academy of Information and Communications Technology, the average price of large model token services on domestic public clouds rose by 51% in the first half of 2026 compared to the second half of 2025.
But the party came at a cost. During reviews, companies suddenly found that token consumption accounted for an increasingly large share of operational costs, yet the actual business conversion and efficiency gains fell far short of expectations.
Survey data from Entelligence AI was staggering: Of every $1 spent on tokens, only $0.18 directly contributed to output, while $0.44 went to fixing bugs, $0.27 to rewriting AI-generated code, and $0.11 was consumed in review processes.
In other words, over 80% of token costs were "idling."


The Turning Point: Price Hikes and Cuts
In the spring of 2026, the global AI industry reached a clear inflection point.
Doubao experimented with paid services, DeepSeek completed its first round of financing, and ByteDance raised its AI infrastructure spending to 200 billion yuan...
A series of signals indicated that the large model industry was shifting from "burning money for traffic" to a rational maturity phase focused on "who can generate profits first."
Those most feeling the pressure were independent model vendors without cloud business cross-subsidies.
In February 2026, Zhipu took the lead in raising prices, increasing the Coding Plan price alongside the release of its GLM-5 overseas version.
Three subsequent rounds of price hikes followed closely: a 20% API price increase with the March release of GLM-5-Turbo, another 10% rise for GLM-5.1 in April.
In June, MiniMax doubled API prices across the board while releasing its third-generation flagship model M3.
In July, Yuezhi Anmian released Kimi K3, raising input prices to 20 yuan per million tokens and output prices to 100 yuan per million tokens, more than tripling previous rates.
DeepSeek's adjustments were most representative: New prices introduced peak and off-peak tiers, along with differentiated pricing for cache hits and misses, with peak cache hit prices rising to 12 times the original.
Morgan Stanley statistics showed that the average Chinese large model API input price rose to 4.9 yuan per million tokens in Q2 2026, with output prices reaching 21.9 yuan, up about 48% and 80% respectively from Q1 2025's 3.3 yuan and 12.2 yuan.
Why couldn't they hold on? The reasons lie in two transmission chains.
The first is compute supply. Long-context, deep reasoning, and agent tasks significantly increase inference loads, while GPU compute supply remains tight during peak hours.
When call volumes grow exponentially, the true compute cost per million tokens cannot be proportionally diluted through economies of scale.
The second is business models. Alibaba, ByteDance, and Tencent have proprietary compute clusters and cloud businesses, making model APIs Drainage (traffic entry points) for cloud services rather than profit centers.
Independent vendors lack cloud ecosystem support, with API revenue as their core or even sole income source. Cost inversions forced them to raise prices.
Interestingly, price hikes didn't scare away users. According to Liang Wei, a partner at Grant Thornton, Zhipu's API pricing rose 83% in Q1 compared to late 2025, yet call volumes surged 400%.
This suggests that enterprise demand has lower price elasticity than expected, with customers willing to pay for certainty.


The "New World" of Tiered Pricing
If you thought price hikes were the whole story, you'd be mistaken.
While independent vendors were forced to raise prices, industry giants chose another path: tiered pricing.
Bank of America analysts noted in a report: "China's large language model price war is evolving from across-the-board cuts to tiered pricing."
Specifically, general-purpose basic APIs follow an affordable pricing route to drive traffic, while ultra-long text, multimodal complex reasoning, and flagship exclusive models maintain high premiums, forming a diversified pricing system where "basic models drive traffic and premium models generate profits."
Consider the data: Current domestic large model API prices show clear "tiered differentiation." Economy models' input prices generally fall to around 1 yuan per million tokens, while price gaps among flagship models widen significantly. Kimi K3's input and output prices reach 20 yuan and 100 yuan per million tokens respectively, whereas MiniMax M3 costs just 2.1 yuan and 8.4 yuan per million tokens after a "permanent 50% discount."
Calculations show that for a typical task, cost differences between models can approach 60-fold.
On the consumer side, tiering is also accelerating. Mainstream models increasingly offer "basic free + advanced paid" tiered systems, with monthly fees ranging from 49 yuan to 700 yuan.
Kimi offers four tiers (49-699 yuan), Doubao three (up to 500 yuan), and video model Hailuo AI even has six tiers, with its premium version reaching 2,299 yuan per month.
Giants' tiering strategies are more subtle. In April, Alibaba Cloud first raised prices for its Bailian large model units by 2%-5%, then canceled unlimited free calls for DataWorks APIs, switching to pay-as-you-go.
Tongyi Qianwen offers three independent billing plans by scenario: pay-as-you-go, Token Plan monthly subscriptions, and Coding Plan exclusive programming subscriptions.
This shows that giants haven't abandoned price competition but have changed tools. They've replaced "one-size-fits-all" price cuts with tiered pricing and simple low-price acquisition with scenario-based charging.
Pricing reforms on the supply side are just one side of the coin. On the consumption side, a quiet revolution is also underway.
On July 24, 2026, a Wall Street Journal report marked a watershed: U.S. companies were collectively stopping money-burning on large model calls.
Over the past two years, firms took pride in "token maximization," with the highest API spenders seen as digital transformation pioneers. Now, this narrative has completely flipped.
Cursor conducted an experiment: Building a web browser from scratch using GPT-5.5 cost just over $10,000, while using Cursor's own Composer model with Anthropic's Opus cost only $1,339—an order of magnitude difference.
This "consumption downgrade" reflects a key cognitive shift: Most tasks simply don't require the most expensive, cutting-edge models.
Uber exhausted its 2026 AI budget by April and set monthly spending caps of $1,500 per employee. Meta, Amazon, Walmart, and Tesla also introduced spending limits or guided employees toward cheaper models.
Domestic companies' attitudes are also subtly shifting. Some programmers note that excessive token calls are now seen as "too costly and unproductive," as these expenses count toward departmental costs, requiring extra approvals for overages.
More notably, companies are developing clear "task-model" matching awareness: Simple text processing uses extremely low-cost open-source or local small models, while only high-difficulty logical reasoning tasks call expensive frontier models.
This tiered routing strategy significantly optimizes token spending structures without reducing actual business value.
UBS Securities analyst Xiong Wei's judgment confirms this trend: "Model performance, model monetization, and token ROI will become three key observation lines for China's AI industry."


"Pay-for-Performance": A New Business Logic
When companies stop paying for token consumption and instead pay for business results, a new pricing model is emerging.
Intercom's Fin AI Agent in the U.S. no longer charges by call volume but by outcomes: 0.99 USD per user problem solved, 9.99 USD per qualified lead screened.
Salesforce's Agentforce Help Agent adopts a "pay-for-resolution" pricing model: Billing only occurs when the agent autonomously solves a problem from start to finish.
Domestic practices are also accelerating. 100Credit adopts a RaaS (Results-as-a-Service) model charging by outcomes, directly converting processing volume and deployment scale into revenue.
Yunxi Technology confidentially submitted a listing application to the Hong Kong Stock Exchange, with its core selling point being the rewriting of enterprise service industry billing rules from "selling software" to "pay-for-performance."
"Companies used to buy information transmission capabilities (traffic packages), now they buy intelligent generation capabilities (token packages)," noted the director of China Telecom AI's Xingchen General AI Lab, describing this as a leap from "bit economy" to "intelligence economy."
The deeper change is that when pricing shifts from "by token" to "by results," the entire AI application workflow needs redesigning.
Research shows that for every dollar spent on tokens, about 80% is implicitly wasted on bug fixes, code rewrites, and review delays. Vendors are shifting from token-based to outcome-based pricing, with rational returns forcing workflow reconstruction.
Looking back from the second half of 2026, the evolution of the AI price war is clear:
Phase 1 (2024-H1 2025): Across-the-board price cuts to capture market share, with companies encouraging unlimited token use and the industry stuck in "burning money for growth" inertia.
Phase 2 (2026-present): Independent vendors are forced to raise prices to correct cost inversions, while giants shift to tiered pricing and scenario-based charging. Companies move from "unlimited use" to "pay-for-performance consumption," with pricing logic shifting from "cost-determined" to "value-determined."
The underlying driver of this shift is the AI industry's inevitable transition from a "technological inflection point" to an "industrial inflection point."
As UBS analysis points out, Chinese large models demonstrate systemic cost advantages in both training and inference. Some leading models' training costs may be just one-tenth of overseas peers', with inference-side API pricing at 10%-20% of comparable overseas models while still achieving 20%-40% gross margins.
But cost advantage does not equate to business success. In the long run, AI chips, semiconductor equipment, memory chips, and large-scale cloud computing platforms are likely to maintain the strongest profitability, while standalone AI model developers may remain at a relative disadvantage.
What truly warrants attention are the corporate practices that are redefining the 'AI economics'.
Wuxi's 'Token Supermarket' enables enterprises to call upon token services on-demand, much like purchasing 'rations', resulting in a nearly 30% reduction in AI R&D costs and a compression of the R&D cycle from two weeks to one.
The monthly token costs of Servyou Group have decreased by 90% year-on-year, with each yuan spent on tokens generating 460 yuan in output. Despite the cost reduction, token consumption has doubled compared to the same period last year.
Uber's engineers run over 30,000 AI Agent tasks daily, and the number of people using AI tools has more than quadrupled, yet token costs have decreased.
These cases all point to the same conclusion: The true commercial value of AI can only be unlocked when the unit cost of tokens continues to decline while token consumption becomes more precise and efficient.

The AI price war is not over. It has simply shifted from one dimension to another.
In the past, competition revolved around the 'price per million tokens'; now, the key metric is shifting from the price per token to the cost per task completed.
This means that what determines a company's competitiveness is no longer just model benchmark scores but a comprehensive operational capability that includes compute resource scheduling, inference optimization, and workflow design.
For enterprises and developers, an important reminder is this: Stop treating token consumption as a KPI for AI transformation.
True competitiveness lies in delivering the best business outcomes with the most suitable models, the most precise workflows, and the lowest overall costs.
The ultimate outcome of the price war is not about being the cheapest, but about being the most worthwhile.