DeepSeek Hikes Prices as 'Price Butcher' Liang Wenfeng 'Retreats'

08/19 2026 438

Produced by Leida Finance. Authored by Zhou Hui. Edited by Meng Shuai.

On August 17, DeepSeek's latest price increase strategy officially came into effect.

It is reported that DeepSeek's API pricing adjustment this time incorporates a peak-valley pricing system. Peak hours are defined from 9:00 AM to 12:00 PM and 2:00 PM to 6:00 PM Beijing time, with the remaining hours classified as off-peak.

Taking DeepSeek-V4-Pro as an example, during off-peak hours, the input price per million tokens is 0.15 yuan (with cache hit) or 4.5 yuan (with cache miss), and the output price is 13.5 yuan. During peak hours, these prices double.

Prior to this price hike, DeepSeek had been dubbed the 'price butcher' in the large model industry.

At the end of 2024, DeepSeek launched the V3 series model, achieving performance on par with top-tier models at a significantly lower training cost. The following year, DeepSeek-R1 was introduced, with its output API price being merely 3% of that of OpenAI's o1.

In May of this year, DeepSeek announced that after the conclusion of a 25% discount promotion on May 31, the API price for the DeepSeek-V4-Pro model would be adjusted to one-fourth of its original price.

Some analysts suggest that DeepSeek's unconventional price increase this time may be linked to the escalating strain on computing resources. DeepSeek's pricing shift also signals a return of domestic large models from price competition to value competition.

DeepSeek Officially Announces Price Increase, Implements Peak-Valley Pricing

At midnight on August 17, DeepSeek, a leading domestic large model, officially initiated a new round of API price adjustments.

In contrast to the previous uniform pricing model across the industry, DeepSeek's latest price adjustment introduces an innovative peak-valley pricing mechanism.

According to DeepSeek's pricing announcement, peak hours are set from 9:00 AM to 12:00 PM and 2:00 PM to 6:00 PM Beijing time, with the remaining hours considered off-peak.

During off-peak hours, the input price (with cache hit) for DeepSeek-V4-Flash is 0.05 yuan per million tokens, the input price (with cache miss) is 1.5 yuan per million tokens, and the output price is 4.5 yuan per million tokens.

Simultaneously, the input price (with cache hit) for DeepSeek-V4-Pro is 0.15 yuan per million tokens, the input price (with cache miss) is 4.5 yuan per million tokens, and the output price is 13.5 yuan per million tokens.

During peak hours, these pricing standards double. For instance, for DeepSeek-V4-Pro from 9:00 AM to 12:00 PM and 2:00 PM to 6:00 PM Beijing time, the input price (with cache hit) rises to 0.3 yuan per million tokens, the input price (with cache miss) rises to 9 yuan per million tokens, and the output price reaches 27 yuan per million tokens.

Compared to the previous prices of 0.025 yuan for input (with cache hit), 3 yuan for input (with cache miss), and 6 yuan for output per million tokens for DeepSeek-V4-Pro, the highest increase during peak hours reaches 1100%.

DeepSeek stated that this API price adjustment aims to allocate resources more judiciously and encourage users to adjust their task schedules based on actual usage.

It is worth noting that DeepSeek's price increase is not an isolated event. Across the domestic large model sector, the industry is experiencing a wave of collective price hikes in 2026.

Among them, Zhipu, hailed as the 'first large model stock listed on the Hong Kong Stock Exchange,' has raised prices three times this year. Meanwhile, Moonshot AI, with the release of Kimi K3, has significantly increased its API pricing, with input prices rising more than threefold and output prices nearly quadrupling.

According to a research report released by Morgan Stanley on August 9 and cited by National Business Daily, in the second quarter of 2026, the average API input price for large models in China rose to 4.9 yuan per million tokens, and the output price rose to 21.9 yuan, representing increases of approximately 48% and 80%, respectively, compared to 3.3 yuan and 12.2 yuan in the first quarter of 2025.

However, despite these substantial price adjustments, DeepSeek still maintains a significant price competitive advantage over its overseas counterparts. According to calculations by Guosheng Securities, the output price of DeepSeek-V4-Pro during peak hours is only about 1/13 of that of Claude's flagship model, and only about 1/25 during off-peak hours.

Unleashing the Ultimate Pricing Strategy, Weekly Call Volume Soars

Information from Tianyancha reveals that DeepSeek's affiliated company, Hangzhou DeepSearch Artificial Intelligence Foundation Technology Research Co., Ltd. (referred to as 'DeepSearch'), was registered and established in 2023.

At the end of 2024, DeepSeek-V3 was released, achieving performance comparable to top-tier models like GPT-4o at an extremely low training cost.

Shortly after, DeepSeek swiftly launched a new model, DeepSeek-R1, in January of the following year.

In terms of pricing, DeepSeek-R1 continued its consistent high cost-effectiveness advantage, with API service pricing set at 1 yuan (with cache hit) or 4 yuan (with cache miss) per million input tokens, and 16 yuan per million output tokens, making its output API price only 3% of that of OpenAI's o1.

With its exceptional cost-effectiveness, DeepSeek emerged as a dark horse among large models. In April of this year, DeepSeek released a preview version of the V4 series model, including DeepSeek-V4-Pro with 1.6T parameters and DeepSeek-V4-Flash with 284B parameters.

On April 26, DeepSeek officially announced an API price adjustment, reducing the input cache hit price for all models in the V4 series to one-tenth of the initial launch price, with an additional limited-time 25% discount for V4-Pro, bringing the input cache hit price for a million tokens down to as low as 0.025 yuan, setting a new global low for large model prices.

According to DeepSeek's official API pricing page, this price reduction applies to all models in the V4 series, with core adjustments focused on input cache hit scenarios. Among them, the input cache hit price for DeepSeek-V4-Flash dropped from 0.2 yuan per million tokens to 0.02 yuan per million tokens.

The following month, DeepSeek once again adjusted its pricing, announcing that after the limited-time 25% discount ended on May 31, the API price for the DeepSeek-V4-Pro model would be formally adjusted to one-fourth of the original price.

After the price adjustment, the input price (with cache hit) per million tokens for DeepSeek-V4-Pro dropped from 0.1 yuan to 0.025 yuan, the input price (with cache miss) dropped from 12 yuan to 3 yuan, and the output price dropped from 24 yuan to 6 yuan.

The low-price strategy also, to a certain extent, propelled DeepSeek's API call volume to rise rapidly. According to statistics from OpenRouter, a globally renowned model distribution platform, from July 27 to August 2, DeepSeek-V4-Flash ranked first globally with a weekly call volume of 7.22 trillion tokens, while DeepSeek-V4-Pro ranked fourth.

On July 31, the official API of DeepSeek-V4-Flash's formal version was launched for public testing. In the following two weeks (August 3 to August 9, August 10 to August 16), the formal version of DeepSeek-V4-Flash consecutively topped the charts with weekly call volumes of 8.83 trillion tokens and 11.2 trillion tokens, respectively.

However, a low price does not necessarily equate to 'poor quality.' According to official evaluation data released by DeepSeek, compared to Opus-4.8, the formal version of DeepSeek-V4-Pro demonstrates significant advantages in multiple test projects, including Terminal Bench2.1, Cybergym, DeepSWE, and AutomationBench (Public).

Even compared to Fable5, the formal version of V4-Pro can also slightly outperform in some projects.

According to Jiangnan Metropolitan Daily, with the popularity of DeepSeek-V4, a new term, 'DeepSeek kill line,' has emerged on social media platforms both domestically and internationally.

It is reported that this term is related to a scatter plot of model cost-effectiveness drawn by Artificial Analysis, with the horizontal axis representing the cost per task and the vertical axis representing the intelligence index score.

The point where DeepSeek-V4-Flash is located has become a dividing line: models with similar performance but significantly higher pricing, or models with inferior performance but even higher prices, are all classified into the 'kill zone,' facing the risk of elimination.

Why Does the 'Price Butcher' Suddenly 'Retreat'?

Although DeepSeek has attracted a massive number of users with its exceptional cost-effectiveness advantage, the surge in call volume has also placed considerable pressure on its servers.

On August 4, OpenCode, an overseas open-source AI Coding Agent platform, stated that DeepSeek-V4-Flash was experiencing capacity issues due to unprecedented traffic, which might cause errors and was undergoing emergency repairs.

According to Yicai, at that time, some users reported that the official API for DeepSeek V4 Flash was almost unusable that morning. DeepSeek's official subsequent statement indicated that the issue had been resolved and service had been restored.

Behind the pressure on DeepSeek's servers lies the severe reality of strained computing resources. According to the 2026 annual meeting of the China Development Forum, in March of this year, China's daily average token call volume exceeded 140 trillion, representing a more than thousandfold increase compared to 100 billion at the beginning of 2024.

However, while demand has surged, supply-side capacity has failed to keep pace. According to Stock Star, current supplies of high-end GPUs and HBM memory remain tight, with NVIDIA's Blackwell and Hopper series facing strong demand.

Data from the China Academy of Information and Communications Technology shows that in the first quarter of 2026, domestic AI computing demand increased by 417% year-on-year, while effective supply growth was only 128%, widening the supply-demand gap.

Constrained by the shortage of computing resources, Moonshot AI's Kimi previously announced the suspension of new C-end user subscriptions after the popularity of Kimi K3, allocating all existing computing resources to serve existing subscribed users.

In addition, some analysts stated that DeepSeek's pricing shift also marks a return of domestic large models from price competition to value competition.

Morgan Stanley pointed out in the aforementioned research report that China's large model industry is establishing a healthier commercial environment, with the focus of competition shifting from price to model capabilities.

Morgan Stanley believes that the key determinant of long-term competitive position for large models lies in model intelligence, not price. Enhancing the capabilities of cutting-edge models requires substantial training and computing investment. Relying solely on low prices to compete for the market will suppress product gross margins, thereby limiting the R&D funds needed for the next generation of models.

Zheshang Securities also pointed out that DeepSeek's lead in raising prices has broken the pessimistic expectation of 'infinite price competition among models,' lifting the valuation anchor for domestic large models.

Radar Finance will continue to monitor DeepSeek's subsequent developments.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.