The AI Price War Reignites! Unveiling the Dynamics Behind Stronger and More Affordable Models

09/28 2026 556

Accelerating the Pace of AI Adoption

Are Large AI Models Poised for Another Round of Price Wars?

Lei Technology (ID: leitech) has observed a notable trend in recent days: OpenAI, Anthropic, Xiaomi, and SpaceXAI have nearly simultaneously unveiled new models. GPT-6 Sol/Luna, Claude Opus 5.5, MiMo-V2.6, and Grok 4.7—each boasting unique names and market positions—are collectively making AI more accessible and cost-effective.

OpenAI has taken a direct approach, nearly halving the price of GPT-6 Sol compared to its predecessor, with Luna pushing the envelope even further by offering even lower prices. Meanwhile, Claude Opus 5.5 has reduced its API unit price by 20%, with Anthropic asserting that the actual costs for typical tasks have plummeted by approximately 40%.

However, it's not all about price slashing. Xiaomi's MiMo-V2.6 and SpaceXAI's Grok 4.7 have maintained their pricing while significantly enhancing performance. Yet, for developers, the distinction is somewhat nuanced:

With the same budget, they now access more machine intelligence at an unprecedented speed.

Image Source: Lei Technology@GPT

This trend is not entirely novel. Over the past few years, the prices of large model APIs have been on a downward trajectory, particularly due to intense competition among domestic open-source large model providers, which has substantially reduced the cost per million tokens. However, this latest round of changes is occurring at a particularly critical juncture.

On one hand, OpenAI and Anthropic have begun openly discussing the need to decelerate the development pace of cutting-edge models, citing concerns over safety and alignment risks associated with rapid AI capability growth. On the other hand, several model providers are simultaneously intensifying their efforts to make existing capabilities more affordable and efficient.

Simultaneously, Agents are transitioning from chat windows to programming, research, and office tasks. A single task now may consume not just a few thousand tokens but hundreds of thousands or even millions, necessitating hours of continuous model interaction.

Stronger and Cheaper: Deciphering the Source of 'Savings'

Why can large models become both more powerful and more affordable? According to explanations from several providers, the primary reason lies in improvements in model training and inference efficiency.

Gone are the days when simply expanding pre-training scale was the norm. Model providers now have a plethora of areas to 'optimize efficiency.'

OpenAI attributes the price reductions of GPT-6 Sol/Luna directly to advancements in caching and inference efficiency. While both models inherit some training methods and capabilities from GPT-6 Astra, they further slash costs on the inference side, enabling OpenAI to reflect these savings directly in API pricing.

Anthropic has taken it a step further. Although Claude Opus 5.5's API unit price has only decreased by 20% compared to Opus 5, Anthropic claims that the actual operational cost for typical tasks has dropped by 40%. In addition to individual tokens becoming cheaper, the new model requires fewer tokens to complete the same task, with caching read prices plummeting by 60%.

Image Source: Anthropic

As models grow smarter, they inherently act as a form of 'price reduction.'

Xiaomi has focused more on the training side. MiMo-V2.6 hasn't reduced API prices but has further expanded the scale of Agentic reinforcement learning compared to V2.5.

During less than six days of live reinforcement learning training, Flash and Pro generated approximately 750,000 trajectories. Xiaomi asserts that by utilizing larger batches, more complex task environments, and finer reward mechanisms, the model learned to complete tasks with shorter paths and fewer tokens.

Ultimately, V2.6 Flash has fully surpassed the previous-generation V2.5 Pro while maintaining the same price.

SpaceXAI has adopted a somewhat different strategy. Grok 4.7 hasn't shrunk its model but instead employs a larger base model, along with longer and more challenging reinforcement learning tasks, while bolstering the model's self-verification, long-context, and Agent capabilities.

Finally, while maintaining the same API price as Grok 4.6, it has further enhanced performance in coding and Agent tasks.

The four technical approaches may differ, but the outcomes converge. Today, model providers are pursuing efficiency across the entire chain, from post-training reinforcement learning to inference optimization, prompt caching, and reducing invalid tokens and tool calls.

This forms the most direct technical backdrop for this round of 'price wars.'

Large models are undoubtedly still evolving, but more and more capability enhancements no longer necessitate simply trading higher inference costs. For developers, this translates to a simple reality: for the same task, models can now accomplish it with less money.

Why Decelerate in One Area While Accelerating 'Price Cuts' in Another?

However, technical efficiency improvements alone do not fully explain this round of 'price wars.'

Just a month ago, OpenAI announced a temporary slowdown in model scaling, including pausing reinforcement learning training for the latest deployed models for two weeks and postponing planned large-scale frontier reinforcement learning training. The direct reasons were twofold: internal Agent safety incidents and GPT-6 Astra reaching the threshold for 'critical cybersecurity capabilities.'

The core issue is that cutting-edge models' capabilities are growing too rapidly, and monitoring, alignment, and safety measures need time to catch up.

Anthropic CEO Dario Amodei took it a step further, publicly proposing 'Pace the Frontier,' calling for the entire industry to decelerate the capability growth rate of frontier AI models. Anthropic even deliberately reiterated this point in the recently released Claude Opus 5.5.

Image Source: Wikimedia

At first glance, this seems somewhat contradictory to the current 'price war.'

However, one easily overlooked aspect is that OpenAI and Anthropic are not advocating for a slowdown in technical progress across the entire AI industry but rather in the continued expansion of the most cutting-edge capability boundaries.

According to Amodei, 'Pacing' does not mean halting model training or technical progress. Instead, he believes today's models themselves are already a vast 'research goldmine' worth investing more time to understand, align, and utilize.

This provides another lens through which to view the current price war.

OpenAI offers a very intuitive data point. After imposing stricter safety restrictions on Astra in August, GPU allocation for Astra-class models once dropped by 59.2%, but GPU allocation for other model categories immediately increased by 17.2%, offsetting about 85% of the decline. Computational power was not idle; it simply flowed to other models and experiments.

Of course, this does not prove that the price reductions for GPT-6 Sol/Luna are a direct result of slowing down Astra's development. However, it at least demonstrates that competition among top model providers does not have only one path: 'training stronger models.'

The same computational power and R&D investment can be used to push the next capability frontier or to optimize existing models, making their existing capabilities cheaper and more reliable through reinforcement learning, inference optimization, caching, distillation, and engineering improvements.

The rise of Agents has further amplified the value of this latter approach.

In the chatbot era, a single response might consume only a few thousand tokens. But in the Agent era, models work continuously for hours, repeatedly reading context, calling tools, self-checking, and even orchestrating other Agents. Every slight reduction in model cost can directly determine whether an Agent product can scale.

But Agents have another dimension. When Amodei explained why he began supporting a slowdown in frontier models, he highlighted a core change: RSI (Recursive Self-Improvement), or in simpler terms:

AI is increasingly involved in developing the next generation of AI.

Anthropic disclosed in August that Claude already plays a 'leading' role in 26% of internal AI R&D tasks, with over 90% of related work involving human-AI collaboration. OpenAI has observed a similar trend, with Coding Agents increasingly participating in writing code, setting up experiments, analyzing results, and advancing model research.

Cheaper and more usable models mean more Agents can run. More Agents can participate in model R&D, improving researchers' efficiency and accelerating the training and iteration of next-generation models.

The cheaper intelligence becomes, the easier it is to produce more intelligence. And this rapidly accelerating diffusion of intelligence may be the most noteworthy aspect of this round of 'price wars.'

In Closing

Over the past few years, large model price reductions have almost become a given. With more models and continuous improvements in computational and algorithmic efficiency, prices have naturally trended downward.

However, as OpenAI and Anthropic begin seriously discussing how to decelerate the growth of the most cutting-edge capabilities, another competition has become clearer: top-tier machine intelligence is being optimized, decentralized, and diffused at an even faster pace.

As William Gibson wrote in Neuromancer, 'The future is already here—it's just not evenly distributed.' But now, this 'unevenness' in machine intelligence is rapidly disappearing.

Source: Lei Technology

Images in this article come from: 123RF Licensed Image Library       Source: Lei Technology

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.