09/28 2026
539
In September, AI vendors worldwide collectively slashed prices—is a new round of price wars upon us?
Domestically, DeepSeek reduced pricing for its Flash series, with input prices for cache hits dropping by up to 60%. Alibaba unveiled five new voice models at its Cloud Town Conference, announcing across-the-board price cuts of up to 95%.
In the U.S., Microsoft was reported to be offering discounts of up to 50% on Copilot enterprise subscriptions, expected to be implemented in October. On the same day, OpenAI and Anthropic released new models and reduced prices. These two companies, previously vocal advocates for "slowing down the pace of frontier model development," shifted their narrative to emphasize "cost savings" following media reports of algorithmic breakthroughs.
Capital markets reacted with panic. Zhipu and MiniMax were caught in the crossfire, with their valuations retracting by 46% and 20%, respectively, since the start of the month.
Is the market interpreting all price reductions as a price war, and all price wars as inevitable profit declines?
Morgan Stanley offers a contrasting perspective. Its report argues that the "premium Pro + budget Flash" dual-track approach is becoming the standard product strategy for large models. While competition in the Flash tier will intensify, a disorderly price war is unlikely to emerge.
A telling detail supports this: even after DeepSeek's price cuts, its pricing remains higher than comparable tiers from Alibaba and Zhipu.
The key to evaluating this round of price reductions lies in whether API call volumes increase and cost curves shift downward post-reduction.
If these metrics respond positively, this round of price cuts may diverge from the zero-sum logic of internet-era price wars, instead acting as an accelerator for AI industrialization and incremental growth.
A Second Wave of Price Cuts: Cost Curves Provide Confidence
Capital markets resist price reductions due to the lingering scars of recent internet-sector price wars.
The first round of AI price cuts in China followed the same zero-sum logic.
In May 2024, DeepSeek V2 slashed API prices to 1 yuan per million tokens, forcing ByteDance, Alibaba, and Baidu to follow suit. At the time, model capabilities were indistinguishable, leaving price cuts as the sole strategy to acquire customers at a loss. By year-end, no winner emerged from this industry-wide price war.
However, the latest round of price reductions differs markedly in logic.
First, reducers now operate with cost buffers.
While 2024's price cuts breached cost baselines, this round is supported by declining cost curves.
The semiconductor industry's Wright's Law states that unit costs decline by a fixed percentage with each doubling of cumulative production. Large-model inference is now following this trajectory.
Zhipu's GLM-5.3-Flash, released in August, reduced attention computation by 3x and KV Cache usage by 4.4x using sparse attention technology. An anonymous test model, "Niulai," ran entirely on domestic chips, achieving single-token costs on par with NVIDIA GPUs after optimization.
Architectural innovations and domestic computing power are jointly compressing costs, with price reductions reflecting technological dividends rather than eroded profit margins.
Second, pricing mechanisms have evolved.
DeepSeek's API employs peak-valley pricing, charging double during peak hours compared to off-peak. Input prices differ by 50x between cache hits and misses.
Microsoft's discounts tier by seat count: 30% for 1,000 seats, 50% for 10,000+.
Alibaba's cuts coincided with the release of five new models under Qwen-Audio-3.1, devaluing legacy capabilities while repricing new ones.
Economically, this is price discrimination: segmenting demand by willingness to pay, usage timing, and procurement scale to maximize compute utilization.
2024 saw prices slashed to the bone; 2026 will use tiered pricing to filter demand. The former chased market share, while the latter optimizes demand structure and compute efficiency.
Budget models will proliferate faster, but the race for cost-effectiveness will persist. More powerful models will occupy lower price tiers, driving sustained declines in unit intelligence costs—a trend inevitable with AI industrialization and technological adoption.
Morgan Stanley thus frames this round of price cuts as competitive defense rather than destructive competition. The critical question is whether reductions can maintain profitability.
The true lesson of 2024's price war is that reductions alone cannot retain customers. This round's dare to cut prices rests on a single premise: call volumes will rebound exponentially after price drops.
Flash for Volume, Pro for Profit: Call Volume Elasticity Reveals the Truth
A 160-year-old economic principle perfectly explains AI vendors' pricing logic.
In 1865, British economist William Stanley Jevons observed: as steam engines became more efficient, total coal consumption rose. Lower usage costs unlocked new demand, outpacing efficiency gains.
This is the Jevons Paradox.
Large models are replicating this dynamic: declining inference costs drive exponential growth in API call volumes, amplifying total compute demand.
Third-party platform OpenRouter's global data shows 127 trillion tokens processed in the second week of September, a 10.4% weekly increase. Chinese models have outpaced U.S. counterparts in call volumes for 20 consecutive weeks, holding eight of the top ten positions. (PS: This explains why U.S. AI giants advocate for slowing development—they aim to lower prices of existing models first to capture market share.)
Chinese vendors strategically pace their growth, cutting prices while ramping up compute investments. Alibaba Cloud's AI-related revenue has grown triple-digit for 12 consecutive quarters, even as it expands AI data center investments. Price reductions and infrastructure upgrades represent two ends of the same cost curve.
DeepSeek and QianWen's volume-focused pricing strategies are expected to further stimulate market demand.
Rising call volumes will trigger two shifts.
Commercially, large-model revenue simplifies to call volume multiplied by unit price. With unit prices trending downward, success hinges on gaining call volume share while outpacing cost reductions.
The dual-track system proves ingenious here.
Flash tiers serve as volume entry points, using low prices to drive penetration. Pro tiers generate profits, with capabilities justifying premium pricing. After DeepSeek's cuts, its model's weekly call volume on coding tool OpenCode surged 30%, validating the entry-point strategy.
Technically, call volumes now generate intrinsic value: each API call accumulates real-world data, which refines models and attracts more calls.
Once this flywheel spins, competition shifts from models to scenarios. Vendors embedding models into more real-world businesses capture both market share and data.
However, for this volume-for-price exchange to work, call volume growth must consistently outpace unit price declines. If penetration falters, only gross margin pressure remains.
Only quarterly financial reports can validate this.
Yet, certain actions prove that constructing new "commercial-technical" flywheels has become imperative.
Industrialization Continues: Dual Engines Propel AI's Future
In mid-September, a faction within the U.S. AI industry called for slowing frontier model development.
Anthropic and OpenAI publicly urged caution, triggering a global sell-off in AI and semiconductor stocks.
BOC International debunked this rhetoric in a subsequent report.
It argued that "slowing down" masks an attempt by frontier labs to convert "safety" concerns into regulatory gatekeeping, assessment rights, and standard-setting powers—codifying their lead.
More damningly, these companies acted incongruously.
While OpenAI and Anthropic democratized access to powerful models by slashing prices, they simultaneously channeled their most expensive compute resources into next-gen frontier models.
On September 22, both companies released even stronger models at lower price points, without reducing R&D spending.
This underscores the prisoner's dilemma of frontier competition. Training next-gen models costs billions, but stopping cedes the intelligence ceiling to rivals. Far from braking, firms are accelerating.
Historically, this is inevitable.
Post-Industrial Revolution, steam and electricity commoditized power, making it cheaper and stronger. Post-Information Revolution, storage and bandwidth followed suit over 40 years. In the AI era, cognition becomes the commodity. Commercial instincts dictate continuous improvement—as long as thinking can be priced, firms will push cognitive capabilities upward.
Breaking cost floors and raising intelligence ceilings will persist, as these define the industry's product evolution.
Thus, this round of price cuts demands reevaluation for its industrial significance.
Profits now derive from two sources: Flash-tier reductions drive application adoption, amortizing compute costs through volume and feeding models with scenario data. Meanwhile, frontier models push intelligence ceilings, justifying premium pricing for advanced capabilities.
Looking ahead, simultaneous breakthroughs in cost floors and intelligence ceilings will form "dual engines" sustaining AI valuations.
Gartner predicts global inference spending will reach $20.6 billion by 2026, surpassing training expenditures for the first time. The industry's revenue model is shifting from training burn to inference monetization—a trend the dual-engine strategy aligns with.
For capital markets, this correction offers repricing opportunities.
September's sell-off stemmed from three pressures: panic over price cuts (sentiment), Zhipu's HK$39.3 billion funding announcement and MiniMax's lockup expiration (supply), and regulatory uncertainties around open-source weight restrictions and anti-distillation measures (compliance).
Note that all three pressures are temporary, leaving industry revenue curves intact.
Zhipu's H1 revenue surged 399.7% YoY, MiniMax grew 283%, and Alibaba Cloud's AI revenue maintained triple-digit growth for 12 quarters.
J.P. Morgan's September 24 report maintained a constructive view on China's AI value chain, favoring Alibaba, Tencent, and Zhipu.
Ultimately, long-term valuations depend on financials.
The market is transitioning from narratives to metrics like ARR, call volumes, and gross margins. As long as Chinese AI firms sustain commercial efficiency, absorb sentiment-driven corrections, and convert call volumes into revenues—and revenues into profits—they will ascend to new heights.
Source: HK Stock Research Society