The Scripts of Chinese and U.S. AI Have Completely Swapped Overnight

08/17 2026 448

Author|Tang Fei

Editor|Li Xiaotian

From late July to early August 2026, the global AI industry witnessed a rare instance of “reverse operations.”

On one side, OpenAI suddenly announced an 80% price cut for its lightweight model Luna, reducing input costs from $1 to $0.20 per million Tokens and output costs from $6 to $1.20. Its mid-range Terra model saw a simultaneous 20% price reduction, while only the flagship Sol model maintained its original pricing.

On the other side, DeepSeek, a Chinese company dubbed the “Token Price Butcher” by developers, issued an announcement: It planned to raise API service pricing across the board in the near future, “with a significant expected increase.”

One lowered prices, the other raised them. One was “rolling downward,” the other “moving upward.”

This was no coincidence. The two events, occurring within less than ten days of each other, reflected starkly different strategic paths chosen by Chinese and U.S. companies in AI—American firms were using tiered pricing to build a defensive system, while Chinese firms were leveraging cost-effectiveness advantages to compete for global pricing power.

A Set of Data Reveals the “Scissor Gap” in Pricing Between China and the U.S.

On one hand, leading U.S. large model companies initiated a wave of price cuts.

Consider OpenAI’s price reductions. In July, OpenAI completed a systemic adjustment of its products and pricing around GPT-5.6, establishing for the first time a three-tier product matrix—Sol, Terra, and Luna—within the same generation. Just three weeks after launch, it swiftly implemented structural price adjustments.

Specifically, Sol (flagship) maintained its original pricing (input: $5.00/output: $30.00 per million Tokens), Terra (mid-range) saw a 20% price reduction (adjusted to input: $2.00/output: $12.00), and Luna (lightweight) experienced a massive 80% price cut (adjusted to input: $0.20/output: $1.20). Simultaneously, a Fast acceleration mode was introduced for Sol at double the price (processing speed increased to 2.5 times that of the standard mode). The overall adjustment pace was significantly faster than previous product cycles, revealing a clear intent to respond to competition.

This three-tier architecture signified that OpenAI was no longer using a single flagship model to dominate the entire market but was instead building a tiered supply system more closely aligned with enterprise budget management. Sol was positioned as the flagship for cutting-edge capabilities, targeting high-failure-cost scenarios such as complex reasoning, scientific research, and long-term agents. The core purpose of maintaining its original pricing was to preserve OpenAI’s capability premium, profit anchor, and brand ceiling. The concurrently launched Fast mode disaggregated latency from service attributes into a separately sellable commodity, extending the flagship layer’s commercial logic from a single capability premium to a dual premium of capability and latency. Terra was positioned as the mainstream production layer, targeting daily production workflows. It intercepted medium-complexity requests originally destined for Sol at a lower cost layer, reducing overall client bills while enhancing routing efficiency within the product family—making it the core hub of the entire operating system. Luna was positioned as the high-concurrency volume layer, targeting high-throughput, price-sensitive tasks such as classification, summarization, extraction, and lightweight agents. After the significant price reduction, it directly entered the ultra-low-price market, competing head-on with cost-effective models from Chinese vendors.

In addition to GPT-5.6’s price cuts, on August 10, Anthropic canceled a planned 50% price hike for Claude Sonnet 5, originally set to take effect in September.

According to reports, when Sonnet 5 was released in late June, its API debut price was $2 for input and $10 for output per million Tokens, with plans to revert to $3 and $15, respectively, starting September 1. The latest update revealed that this price hike had been canceled, with $2/$10 continuing as the long-term pricing.

Google’s newly released Gemini 3.7 Flash also demonstrated significant pricing Sincerity (goodwill). According to its official pricing strategy, Gemini 3.7 Flash charged just $0.75 for input and $3.75 for output per million Tokens. This represented a direct 50% discount off the previous standard pricing of Gemini 3.6 Flash.

Moreover, this preferential pricing was not a short-term measure but would persist until January 1, 2027, meaning Google had left the market with a multi-month low-price window.

On the other hand, Chinese large model companies were one after another (collectively) initiating price hikes.

First was DeepSeek. On August 11, DeepSeek previewed an imminent across-the-board price increase on its official API pricing page, issuing a warning of a “significant expected increase.”

On the evening of August 12, DeepSeek V4 Pro’s official version was launched. In terms of pricing, calculated per million Tokens, V4 Pro’s input was 0.025 yuan (cache hit) and 3 yuan (cache miss), while output was 6 yuan. In other words, both the output Tokens and input (cache miss) of V4 Pro were triple those of the previous version, V4 Flash.

This was not an isolated case. On August 11, Alibaba’s QianWen App quietly launched paid services. Its Office Assistant Pro membership came in three tiers—200 yuan for continuous annual subscription, 568 yuan, and 1,499 yuan. AI video generation was charged by quota, with 500 quotas costing 968 yuan.

Looking back, in June, ByteDance’s Doubao introduced three membership tiers, priced between 688 and 5,088 yuan annually. Zhipu raised prices three times within the year, accumulating an 83% increase, yet its usage volume surged by 400%. Yuezhi’s Kimi K3 increased prices by 3 to 4 times and added a clause for “up to 30% revenue sharing.”

A Morgan Stanley report noted that in the first quarter of 2025, vendors such as ByteDance, Alibaba, Baidu, Tencent, MiniMax, Zhipu, Yuezhi, and DeepSeek had an average input price of approximately 3.3 yuan and an output price of about 12.2 yuan per million Tokens. By the second quarter of 2026, input prices had risen to 4.9 yuan, and output prices to 21.9 yuan, representing overall increases of 48% and 80%, respectively.

OpenAI’s “Tiered Defense” Strategy

After presenting the data, we must examine the reasons behind the price adjustments.

Why did OpenAI lower its prices? Let’s first consider OpenAI’s official explanation: “Improvements in GPU cores, inference systems, and production environment efficiency have been achieved, and the company is passing some of the benefits from these efficiencies onto users.”

These optimizations are real. However, the issue is that this explanation does not account for Luna’s staggering 80% price reduction while Sol’s price remained unchanged.

The real pressure comes from competition.

Chinese models are surpassing their U.S. counterparts across multiple dimensions: According to a CNBC investigation, Chinese models like DeepSeek and Kimi have captured 46% of Token usage by U.S. enterprises on the OpenRouter platform, even surpassing U.S. models during certain periods. DeepSeek V4 Pro is priced at $0.435/$0.87, superposition (with) a 75% long-term discount. If OpenAI does not compete head-on in the lightweight model segment, corporate usage could continue to decline.

Meanwhile, Alibaba’s Qwen has surpassed 3 billion global downloads in the past six months, exceeding Meta and Google to become the world’s most downloaded open-source AI model. Alibaba’s emails reveal that the Qwen family has open-sourced over 460 models, spawning more than 300,000 derivative models. According to a Hugging Face report on August 14, Google and Meta’s open-source model downloads in 2026 were just 418 million and 227 million, respectively.

After Luna’s price reduction, its input price fell below DeepSeek’s. According to Artificial Analysis evaluations, Luna’s performance surpassed Google’s Gemini 3.6 Flash and even the older Gemini 3.1 Pro. Users could now obtain better performance for less money, potentially triggering a new wave of user migration.

A deeper change lies in strategy.

Luna’s significant price reduction was essentially OpenAI’s proactive move to defend the mid-to-low-price market. It did not need to achieve the absolute lowest price but only to narrow the price gap sufficiently so that it did not justify the costs of customer migration, compliance, and ecosystem switching, thereby retaining API traffic and developer access. Overall, this tiered price adjustment marked OpenAI’s revenue model officially shifting from “flagship premium driving single-user revenue” to “tiered routing driving total usage volume and customer retention.”

Another factor may be the influence of Wall Street investors.

OpenAI has secretly filed an IPO application with the U.S. Securities and Exchange Commission, with CEO Sam Altman stating in an internal Slack message plans to “go public within a year.”

A clear IPO timeline means OpenAI needs to deliver impressively robust financial results within the next year. The core metrics that capital markets value are undoubtedly total user count, profit margins, and whether user growth can offset margin compression. Regardless, a stable user base is essential for maximizing future profitability.

The differing magnitudes of price adjustments reveal OpenAI’s strategic intent: defend high-end pricing, consolidate mid-range positioning, and capture low-end volume. By actively widening the price gradients between tiers, it forms a complete product matrix covering diverse budgets and scenarios.

However, defense remains defense. OpenAI has not relinquished pricing power but is redefining it—shifting from past per-Token premiums on a single flagship model to budget share across the entire product family at different task levels. This signifies that the era of relying solely on generational upgrades to sustain price hikes has ended.

DeepSeek’s “Commercialization Repair”

While OpenAI was defending downward, DeepSeek was attacking upward.

In fact, DeepSeek’s pricing adjustment route had been foreshadowed earlier. In April, it lowered input prices for cache hits and introduced limited-time discounts. In July, it introduced peak-valley time-based pricing. By August, it previewed an overall price hike.

Why did it dare to raise prices?

First, its capabilities had caught up.

From a capability standpoint, DeepSeek V4 Flash maintained its foundational architecture of 284B total parameters, 13B activated parameters, and a 1M context window, achieving significant capability improvements solely through post-training. According to Artificial Analysis data, V4 Flash 0731 scored 50 on the Intelligence Index, just 1 point below GPT-5.6 Luna’s 51, indicating extremely close comprehensive intelligence levels.

Second, its cost advantages stemmed from genuine engineering efficiency, not subsidies.

V4 Flash employed a CSA+HCA hybrid attention mechanism capable of highly compressing KV caches while retaining only the most relevant Tokens through sparse filtering, significantly reducing invalid computations. DeepSeek’s technical report revealed that under a 1M context setting, V4-Pro required just 27% of the single-Token inference FLOPs and 10% of the KV Cache compared to DeepSeek-V3.2. V4-Flash further reduced single-Token computation and cache usage through additional sparse filtering, with engineering estimates showing another significant decline from V4-Pro.

Most aggressively, its cache pricing set cache hit input prices at just 2% of miss prices (the industry average was a 90% discount). In agent workflows, characterized by multi-round calls and repetitive contexts, 80-90% of inputs are redundant. Prompt caching temporarily stores fixed content computations, skipping Transformer layer operations upon cache hits and consuming negligible computational resources. Thus, DeepSeek could offer nearly free cache pricing.

According to Artificial Analysis data, DeepSeek-V4-Flash’s average testing cost was just $0.03, compared to $1.86 for OpenAI’s GPT-5.6 Sol and $3.15 for Anthropic’s Claude Fable 5—a difference of tens to hundreds of times.

So even after the price increase, DeepSeek's pricing remains significantly lower than that of overseas models in the same tier, and its cost-effectiveness advantage still exists.

Professor Jiang Zhenhui from the School of Business at the University of Hong Kong, specializing in Innovation and Information Management, analyzed that DeepSeek's ability to offer such competitive pricing stems from both architectural optimizations like its Mixture of Experts (MoE) design and efficient training strategies such as multi-teacher knowledge distillation. However, the cost savings from these technological advantages ultimately cannot offset the reality of serving massive traffic volumes at near-loss levels.

Therefore, the price increase appears more like a rebalancing of DeepSeek's business model after scaling up.

Morgan Stanley analysts Gary Yu and Lydia Lin interpreted DeepSeek's price adjustment as a positive signal for improved pricing discipline in the industry. They pointed out that the sector is shifting from price competition to commercialization driven by intelligent capabilities, with three structural changes occurring simultaneously: pricing becoming more rational, open-source licensing tightening, and model parameter scales leaping forward.

The "Rise of the East, Decline of the West" has just begun

When viewed alongside GPT-5.6's tiered pricing adjustments and DeepSeek's pricing shift, the power structure of the global AI industry is being rewritten.

First, pricing power is shifting.

In the past, OpenAI's pricing served as the anchor for the entire industry. When this anchor begins to move downward, it rewrites not just the price of a single product but the pricing logic of the entire sector.

Now, Chinese models are becoming the new price reference points. DeepSeek's price increase is not a sign of "collapsing under pressure" but rather "securing victory"—first capturing the market with extreme cost-effectiveness, then reclaiming pricing power.

In the future, the global large model market may enter a clearly stratified competitive phase, where the market is no longer dominated by a single strongest model. Different tiers will have entirely distinct competitive logics and pricing rules.

For example, the top tier consists of high-end markets with high failure costs, covering scenarios like complex code repair, long-duration knowledge work, long-chain agents involving multi-tool coordination, and high-risk professional decision-making. In these scenarios, the business losses from a single task failure often far exceed the token cost differences between models. Customers prioritize task success rates, reliability, and stability, willing to pay significant premiums for stronger capabilities. Participants in this tier mainly include advanced versions of Claude, GPT-5.6 Sol, Kimi K3, etc., which maintain relatively high pricing, possess strong pricing power, and compete primarily on sustained capability leadership and scenario deepening, with price not being the primary decision factor.

The second tier is the mass automation market, covering high-frequency scenarios like conventional text generation, simple coding, information extraction, classification summarization, lightweight agents, and batch document processing. These scenarios share characteristics of high task standardization, frequent invocation, and low per-task value, with modest requirements for extreme reasoning capabilities. Once model capabilities reach a usable threshold, price becomes the core decision variable for customer selection. This tier represents the most fiercely competitive area currently, with numerous products like DeepSeek V4 Flash, GPT-5.6 Luna, GLM series, and Gemini Flash competing in this space.

Second, the AI gap between China and the US is narrowing, while the price gap is widening.

According to OpenRouter platform statistics from January-June 2026, DeepSeek's token invocation share doubled from 9% to 18%. Based on Ramp's billing data from over 50,000 US enterprises in June 2026, DeepSeek topped the monthly software trend list, with American companies like Coinbase and Lindy directly purchasing rather than merely deploying open-source weights locally. According to Stanford's "2026 AI Index Report" data from March 2026, the performance gap between top Chinese and US models was just 2.7%. Combined with significant price advantages, US enterprises switching to DeepSeek could reduce inference costs by 30%-95%.

Multidimensional data indicates that this pattern of "similar capabilities, vastly different prices" is driving global developers to vote with their feet.

A research report by Sinolink Securities suggests that on the demand side, domestic models' capabilities are crossing the usability threshold, with token invocation volumes sustained growth (continuously growing). High-bandwidth communication domains have shifted from engineering optimization items to model deployment requirements. On the supply side, improvements in domestic GPU and CPU supply, coupled with sustained increases in capital expenditures by major firms, are expected to propel domestic computing power from chip introduction to system-level volume expansion.

The institution further analyzed that the competitive barrier in large models lies in intelligence level rather than price. Top players can introduce low-cost lightweight versions downward, but mid-tier players face much greater difficulty breaking upward into top-tier models. This asymmetry determines that pure price wars are not sustainable strategies.

The "Rise of the East, Decline of the West" is not just a slogan but an industrial reality unfolding before our eyes.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.