There's No Such Thing as a Free DeepSeek

08/20 2026 479

Qingping Chuiguo | Author

Star All Empty | Editor

Accidentally Orange Juice | Visuals

Liu Chao | Supervisor

Several years down the line, workers will find themselves staring at exorbitant API renewal bills from DeepSeek, reminiscing about the days when a cache hit cost a mere 0.025 yuan.

DeepSeek, once lauded for slashing large model API prices to rock-bottom levels, now seems to have 'forgotten its roots.' It has officially unveiled a new pricing strategy, introducing 'tiered pricing based on computing power peaks and valleys' among domestic large models.

During peak hours, prices double, while off-peak rates are halved. The cost of a cache hit input has soared from 0.025 yuan to 0.3 yuan, marking a staggering 1,100% increase.

This news has sent ripples through the developer community, causing quite a stir.

Some developers have pulled all-nighters to reschedule batch tasks for the early morning hours, while others have embarked on meticulous comparisons of off-peak price lists across various APIs.

However, a more profound realization is dawning on many: If even DeepSeek, the so-called 'price butcher,' can no longer maintain its low-cost stance, is the AI utopia of unlimited, free access truly nearing its 'end credits'?

Such concerns are not unfounded, as DeepSeek is not the first large model to hike prices.

As early as May this year, Doubao introduced paid memberships for consumer-end users. Tencent Cloud soon followed suit, increasing model input prices severalfold. Kimi, Zhipu, and MiniMax were not far behind, leading to a widespread surge in B-end API prices.

The era of free AI is waning, and the commercial turning point where large models transition from 'competing on price' to 'competing on value' has arrived.

The 'freebies' we once took for granted are now becoming 'unaffordable luxuries.'

Around this time last year, virtually all basic functions of domestic AI applications were available free of charge. Writing copy, creating spreadsheets, organizing meeting minutes, generating images—all these tasks could be accomplished without spending a dime.

An unspoken 'rule' prevailed: AI was meant to be free.

In May this year, Doubao shattered this tacit understanding by introducing three tiers of paid subscriptions: Standard at 68 yuan/month, Enhanced at 200 yuan/month, and Professional at 500 yuan/month.

While free users could still access the platform, long-document processing, batch generation tasks, and priority computing power allocation during peak hours were reserved exclusively for paying subscribers.

Kimi soon followed suit.

After releasing K3 in July, Kimi mandated subscriptions for users wanting access to the latest model. Four membership tiers were offered, ranging from 39 yuan to 559 yuan per month.

Dramatically, within 48 hours of launch, Yuezhi Anmian (the company behind Kimi) had to urgently suspend new consumer subscriptions due to overwhelming demand. Request volumes surged far beyond expectations, putting a strain on GPU clusters.

Even the desire to charge was constrained by computing power supply, highlighting the cost pressures faced by manufacturers.

These changes were felt keenly by ordinary users.

However, the true determinants of industry trends lie not in consumer-end membership fees but in B-end API pricing systems, which directly impact AI application development costs and ecosystem prosperity.

So, let's delve into the B-end price hikes:

In March, Tencent Cloud ended free public testing for third-party models like GLM 5 and MiniMax 2.5, while raising input prices for its self-developed Hunyuan HY2.0 Instruct from 0.0008 yuan/1,000 tokens to 0.004505 yuan, a 463% increase.

In April, Tencent Cloud uniformly raised AI computing power product prices by 5%.

These two 'strikes' in one month covered everything from models to infrastructure.

Zhipu AI also raised prices five times within the year. API pricing in Q1 increased by about 83% compared to late 2025, yet usage volume surged 400%.

Both volume and prices soared, revealing that while users grumbled about costs, their actions spoke louder than words.

Kimi K3 was no slouch in the B-end market, with output prices skyrocketing to 100 yuan/million tokens (3.5 times higher than the previous generation) and input prices to 20 yuan/million tokens (3 times higher), stepping into the price range of overseas closed-source models.

Meanwhile, MiniMax launched the M3 model on June 1st while doubling API input prices from 2.1 yuan to 4.2 yuan and output prices from 8.4 yuan to 16.8 yuan. The cheapest package jumped from 29 yuan to 49 yuan, with no advance notice.

Finally, DeepSeek implemented its 'peak-valley electricity pricing' scheme a few days ago.

The day was divided into two periods: Peak (9:00-12:00, 14:00-18:00) and off-peak (remaining hours). Off-peak prices were set at half of peak prices. V4-Pro output prices during peak hours surged to 27 yuan/million tokens (4.5 times the old price), while cache hit prices skyrocketed 12-fold.

From being the 'price butcher' to adopting 'peak-valley pricing,' this transformation took just over a month.

The era of 'free AI' has truly come to an end, with price hikes sweeping across the industry—from user-end to developer-end, from model calls to underlying computing power.

Why the sudden price hikes?

When large models raise prices, many's first reaction is to accuse them of 'cashing in.'

However, a closer look at AI costs reveals that this round of hikes was inevitable.

We're accustomed to internet products like Douyin and WeChat, where marginal costs approach zero, leading us to subconsciously equate traffic with money. AI, on the other hand, is a different beast—every dialogue, long-document analysis, and image generation consumes GPU power, electricity, and money.

Electricity bills, equipment depreciation, and computing power consumption are all hard costs that cannot be ignored.

Doubao now boasts 382 million monthly active users, far ahead of competitors, with average token usage exceeding 180 trillion.

Impressive, right? But industry estimates suggest Doubao burns tens of millions daily, with daily revenue below 1 million yuan as of May this year. ByteDance's AI infrastructure spending could soar to 200 billion yuan in 2026.

Tencent Cloud also stated in its price adjustment announcement: 'Global AI computing power demand is surging, and core hardware supply chain costs have risen sharply.'

Yuezhi Anmian bluntly admitted when releasing K3: 'Smarter models performing more complex tasks consume vast resources.'

Thus, large models don't refuse free access out of choice—the physical costs of 'free' have hit the hard ceiling of computing power. If every call burns GPUs, electricity, and money, low pricing becomes an unsustainable loss-making proposition.

Not to mention the explosion of AI agents, which has multiplied computing power consumption. Token requirements for completing an agent task can be 10 times or even 100 times those of answering a simple question.

The National Data Bureau estimates that China's daily token calls reached 140 trillion in March this year, growing over 1,000 times since early 2024.

More critically, low pricing not only amplifies losses but also attracts massive low-value, high-concurrency invalid requests (e.g., bots, test scripts), severely straining computing resources for high-value enterprise clients.

DeepSeek's 'peak-valley pricing' essentially borrows from the power industry's load-shifting strategy. It's not about gouging customers but using price levers to redirect peak-hour requests to off-peak hours, avoiding the enormous capital expenditures of continuously scaling GPU clusters to handle peak loads.

Viewed from this perspective, price hikes are inevitable. If major players persist with below-cost pricing, losses will snowball, and capital markets will no longer tolerate 'money-burning' narratives.

The days of relying solely on financing are over. Large model firms must now prioritize revenue and gross margins in their KPIs.

Price hikes aren't the goal—commercialization is

Among this wave of price hikes, ByteDance's approach stands out.

Doubao's model differs from others. It doesn't just sell memberships—when users book hotels or hail rides through Doubao and complete transactions, Doubao earns commissions on payment channel fees (currently in gray-scale testing).

ByteDance CEO Liang Rubo stated at an August 6th all-hands meeting: 'We hope Doubao will become a 'trunk' business, like Douyin, branching into e-commerce, local services, and more.'

Thus, an AI assistant is evolving into a 'transaction platform.'

ByteDance's strategy has always been to acquire users at any cost, build scale, and then monetize. Douyin and TikTok followed this path.

Doubao's current moves mirror Douyin's foray into local services—once user scale is achieved, it's time to reel them in.

But the question remains: Why should users stay within ByteDance's 'net'?

Baidu's combo is 'Wenxin + Cloud,' but its enterprise application product matrix lags behind Feishu. Alibaba's combo is 'Tongyi + Cloud,' with DingTalk actively deploying AI, but its consumer-end user base still trails Doubao.

In contrast, ByteDance is among the few domestic players with a super consumer-end gateway (Doubao), an enterprise collaboration base (Feishu), and cloud infrastructure (Volcano Engine). Their integration creates vast commercialization potential.

Beyond Doubao, other firms are exploring their own paths.

Zhipu has achieved 'volume and price growth' through rapid technological iteration and repeated price adjustments. Kimi focuses on high-value programming scenarios, using the Mooncake architecture to reduce actual usage costs. Tencent Cloud opted for across-the-board price hikes, covering model APIs to underlying computing power.

Clearly, despite differing approaches, the end goal is the same: finding sustainable business models, not just raising prices.

Epilogue

Recall earlier this year, when the AI industry was luring consumer-end users with cash incentives ('treating the nation to milk tea'), only to see the same players collectively raise prices within six months—from handing out freebies to charging fees, faster than flipping a book.

Though swift, this shift isn't surprising. All eras of 'free lunches' eventually end.

Reviewing internet business history, the removal of free offerings follows a familiar pattern: Cultivate user habits and market scale through free access, then monetize through tiered pricing.

In 2002, 263.net, China's largest free email provider, canceled free services, requiring all users to pay. In the 2010s, after a 'free war' among cloud storage providers, the industry shifted to pay-per-use acceleration and membership subscriptions. This May, Meta, the world's largest 'free social empire,' launched paid subscriptions for Instagram, Facebook, and WhatsApp.

Now, AI is following this path.

The retreat of free offerings reflects the industry's transition from wild growth to standardized development. When manufacturers stop burning money to acquire users and instead retain them through genuine capabilities, true competition begins.

As free resources dwindle, AI doesn't lose value—quite the opposite. When a service starts charging, it signals real utility.

Future rules will clarify:

Basic consultations will remain free; productivity services that save significant time or generate tangible benefits will require payment. Computing power, intelligent creation, and efficient information processing are inherently valuable intellectual services worth paying for.

So at this turning point, we don't need to worry about price increases, but we need to think clearly: Which AI capabilities are truly creating value for me,

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.