US Large Models Also Start a Price War

09/28 2026 464

The Wind of Large Model Price War Blows Across the Pacific

Image Source | Internet (Please contact for removal if infringing) Partially AI-generated

After competing on parameters, large models are finally competing on price. This time, the companies leading the price cuts are the two with no shortage of customers.

On September 22 (US time), Anthropic released Claude Opus 5.5. About an hour and a half later, OpenAI launched GPT-6 Sol and GPT-6 Luna.

Both prominently featured "affordability" on their posters. GPT-6 Sol reduced prices to $2 per million tokens for input and $10 for output, half of the previous generation's prices. Luna slashed input prices to $0.10, a maximum reduction of 58%. Opus 5.5 reduced input and output prices to $4 and $20, respectively.

An AI company with a trillion-dollar valuation and the most users, and a model vendor with the highest paid adoption rate among US enterprises, both voluntarily cut prices on the same day. What's behind this?

Two Different Price-Cutting Strategies

Although both companies cut prices on the same day, their approaches are vastly different.

OpenAI adopted a "tiered pricing, volume-driven" strategy. GPT-6 Sol is positioned for complex coding and agent workflows, while Luna targets cost-sensitive, high-volume tasks.

In essence, different model tiers are assigned specific roles: high-difficulty tasks go to the Astra flagship, daily coding and batch processing use Sol, and high-frequency, lightweight conversations go to Luna.

Critically, OpenAI explicitly stated that the new prices are long-term, not promotional or launch discounts, but permanent pricing.

This means OpenAI is betting on a judgment: by lowering prices sufficiently, daily usage will grow exponentially, and economies of scale will eventually reduce unit costs to sustainable levels.

Anthropic's approach is entirely different. While Opus 5.5's unit price reduction is only 20%, Anthropic emphasizes "overall cost reduction" by introducing a new dynamic caching mechanism that slashes cache read prices from $0.50 to $0.20 per million tokens, a 60% reduction.

At the same time, the total tokens required to complete tasks with Opus 5.5 are effectively halved, with output speed increasing by over 30%, reducing overall operational costs for typical workloads by about 40% compared to the previous generation.

To illustrate, OpenAI is like a supermarket offering a 50% discount on price tags to attract buyers, while Anthropic is like a restaurant optimizing its supply chain—menu prices remain largely unchanged, but each dish uses fewer ingredients and is served faster, so you get more for the same money. One reduces unit prices, the other reduces total costs.

Why did both companies choose to cut prices at this time? On the surface, it's competition, but deeper down, three structural forces are at play.

First Force: Technological breakthroughs in architecture provide confidence for price cuts.

This round of price cuts is not subsidized at a loss. Anthropic introduced a new dynamic caching mechanism in Opus 5.5, reducing the storage and retrieval costs of intermediate state data to a quarter of the original.

With this underlying improvement, when enterprises call the same core business data, the actual token consumption is effectively halved.

OpenAI's significant concessions on the Sol series are also based on a reconstruction of basic computational logic, while Luna further strips away non-essential general-purpose computing modules to focus on high-frequency, short-duration conversational Q&A scenarios.

When the underlying computational paradigm undergoes a qualitative change, "affordability" is no longer a reluctant compromise but a natural byproduct of new technological architectures.

Second Force: Market data reveals a harsh reality—expensive models are no longer selling.

Data from enterprise spending management platform Ramp shows that a month after Anthropic's most expensive Fable series launched, it accounted for only 6% of tokens purchased by enterprises from Anthropic.

This figure illustrates a simple fact: most tasks don't require paying for the top tier. If a task can be done well with Opus, enterprises have no reason to pay extra for Fable.

Meanwhile, Ramp's tracking data shows that the effective price per million tokens has dropped from $1.15 in March this year to $0.68, a 41% decline in six months.

Enterprises are becoming more cost-conscious with AI spending, no longer blindly paying for the "strongest model" but instead calculating: Can this task be done with a cheaper model?

Third Force: Low-priced offensives from Chinese models are forcing a response.

Domestic models like DeepSeek have lowered the global price threshold for large models by an order of magnitude. DeepSeek V4 Flash costs just 1 RMB per million tokens for input and 2 RMB for output, approximately $0.14 and $0.28 respectively.

Even after OpenAI slashed Luna's prices, its output price is still more than four times that of DeepSeek V4 Flash.

A Morgan Stanley research report notes that while the average price of Chinese large model APIs has risen significantly over the past year, prices of closed-source US models have continued to decline, narrowing the price gap between Chinese and US large models.

This "you raise, I lower" price divergence indicates that US vendors are feeling the pressure, as price-sensitive customers will without hesitation switch to Chinese models if prices don't drop.

Different Price War Tactics Across the Ocean

Shifting focus back to China, this price war script actually played out two years earlier.

In May 2024, ByteDance's Doubao take the lead [original Chinese pun not translatable] set the price for its Pro model at 0.0008 RMB per thousand tokens, 99.3% lower than the industry average.

Subsequently, Alibaba Cloud's Tongyi Qianwen main models reduced prices by 97%, Baidu's Wenxin large models went fully free, and Tencent's Hunyuan large models saw price reductions of up to 87.5%. For a time, the entire industry descended into a frenzy of "selling tokens at a loss."

But two years on, the domestic market has seen a reversal. Calculations by Guolian Minsheng Securities show that China's overall daily token consumption surged from 100 billion in early 2024 to 180 trillion in February 2026.

Amid explosive demand, companies like Zhipu AI and Tencent Cloud have issued price hike notices, with some products rising by over 400%.

In Q2 2026, the average API input price for Chinese large models had risen to 4.9 RMB per million tokens, with output prices at 21.9 RMB, increases of about 48% and 80% respectively compared to Q1 2025.

The price war could no longer continue. A ByteDance executive stated internally: "In the next 18 months, only players controlling the computing power supply chain will survive."

This judgment applies equally to the US market, but the underlying logic of the price wars differs fundamentally between the two countries:

The domestic price war was about "burning money to seize market share." ByteDance used its pre-stocked computing power advantage and the industry's lowest marginal costs to crush competitors, forcing Alibaba and Baidu to follow suit. For example, Alibaba compressed 397 billion parameters to 17 billion using MoE architecture, reducing memory usage by 60% and increasing inference throughput by 19 times—a case of "being forced to cut costs."

The problem with this approach is that when the market shifts from growth to consolidation, burning money becomes waste, not investment.

The US price cuts this time resemble "technological dividend release." OpenAI and Anthropic's price reductions are clearly supported by architectural optimizations, not sacrificing margins for market share but recalculating costs under new technological paradigms.

A Bank of America report also notes that Chinese AI companies are abandoning simple price wars and increasingly pricing based on task completion costs. To some extent, pricing strategies in China and the US are converging.

So what will this price war ultimately hinge on?

First, technological scale. The confidence to cut prices comes not from funding amounts but from the engineering capabilities required to reduce inference costs by each percentage point.

Anthropic's cache read mechanism optimizations and OpenAI's inference efficiency reconstructions represent costly technological barriers. Xia Lixue, CEO of Infinite Chip Technology, provided data showing that through collaborative optimization between models, system software, and hardware, inference costs have dropped 90% over the past year, with potential for another 90% reduction.

This means the price war's outcome won't be determined by who collapses from losses first, but by whose technological cost reductions are faster. The faster cost reducer can maintain profit margins at lower prices, while the slower one will be forced out through losses.

Second, valuations. OpenAI is currently valued at $852 billion, with Anthropic's latest valuation reaching $965 billion, both approaching the $1 trillion mark. This valuation means they can sustain astronomical computing power investments.

ByteDance's capital expenditures exceeded 150 billion RMB in 2025, with about 90 billion going to AI computing power. In 2026, it plans to invest 160 billion RMB, with 85 billion for AI chip procurement.

Only companies of this scale can continue ramping up R&D while cutting prices—a game small players simply can't afford.

More critically, the price war itself will accelerate industry differentiation. Global startups have raised about $510 billion in funding, with OpenAI and Anthropic alone absorbing about $217 billion, or 43% of global venture capital.

With funds highly concentrated at the top, mid-sized and small model vendors lack both the technological capability to reduce costs and the capital to burn money, not to mention the confidence to sustain computing power investments. When the price war ends, not all players will remain at the table.

In 2013, Didi and Kuaidi's subsidy war burned billions before merging.

In 2020, the community grocery price war lasted over a year, leaving only two or three survivors.

The large model price war will follow the same logic—it's an elimination round about "who can provide the best capabilities at the lowest prices."

OpenAI and Anthropic's price cuts may appear to be a price war, but essentially, they're using technological advantages and capital scale to redefine industry cost benchmarks.

With per-million-token prices dropping below $1, large models are shedding their luxury image and becoming infrastructure like utilities.

Only those with sufficiently robust technology, massive scale, and valuations capable of sustaining long-term investments can continue operating and innovating at this price level.

The price war's outcome has never been about who is cheaper, but who can thrive while being cheap.

Technological scale and valuations are the thresholds for survival.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.