Nine Days After DeepSeek Raised Prices, Zhipu Enters the Market with One-Tenth the Cost

08/27 2026 453

China's First Validated Large Model Business and Its Initial Price War

Author|Huawen

Editor|Xiaobai

Illustrations|AI-Generated

Produced by|Qiangdiao Next On the evening of August 26, MiniMax founder Yan Junjie reported a figure at an earnings call: token consumption in July was 20 times that of January this year. This company, which went public on the Hong Kong Stock Exchange in January, reported revenue of $117 million in the first half of the year, a 283% year-on-year increase, with revenue from its enterprise-oriented open platform surging sevenfold. A few hours earlier, The Information cited two insiders to disclose DeepSeek's most complete set of financial figures to date.

In the first seven months of this year, DeepSeek's revenue was approximately 475 million yuan ($65.6 million), ten times its full-year revenue for 2025, with a net loss of 715 million yuan, nearing the 935 million yuan loss for the entire previous year. The gross profit margin for its API business was 82.9%. On the same day, Zhipu launched and open-sourced GLM-5.3-Flash. This model, with 320 billion total parameters, is priced at one-tenth of its flagship GLM-5.3 and lower than DeepSeek V4-Flash's off-peak price.

It is powered online by a cluster of over 100,000 domestically produced chips. Behind these three sets of data lies the same business: "lightweight flagship" models with two to three hundred billion total parameters, activating only one to two hundred billion parameters per call. This represents the first fully validated business in China's large model industry.

There is significant demand, suppliers have gross margins, and manufacturers have gained pricing power for the first time. However, Zhipu's entry method reveals another side of the story. Just nine days after DeepSeek raised prices for its V4 series on August 17, a model with higher evaluation scores entered the same market at a lower price. Validation and dilution followed in quick succession.

───

1  Demand Confirmed Three Times in One Day ■

The hardest evidence for judging genuine demand is not user numbers but price increases. At midnight on August 17, DeepSeek's new API pricing for its V4 series took effect, introducing peak and off-peak pricing for the first time: 9 AM-12 PM and 2 PM-6 PM are peak hours, with off-peak prices half those of peak hours. The output price for V4-Flash rose from 2 yuan per million tokens to 9 yuan during peak hours, a 350% increase.

Input prices for cache hits increased by 400%. The higher-tiered V4-Pro has a peak output price of 27 yuan. Even after these increases, it remains one of the cheapest frontier models globally. V4-Pro's peak output price of 27 yuan (approximately $3.96) was previously 0.87 dollars. Goldman Sachs interpreted this round of price adjustments as indicative of "sustained strong demand and tightening computing resources." Zheshang Securities put it more bluntly: this enhances the overall pricing power of domestic frontier open-source models. DeepSeek is not alone in raising prices.

Zhipu canceled first-purchase discounts for its Coding Plan in February and increased API prices for GLM-5-Turbo by 20% in March; Yuezhi's Kimi K3 set its output price at 100 yuan per million tokens, more than triple that of its predecessor. According to BlockBeats, seven cloud and model providers have raised prices since the beginning of the year.

Zhipu's API call pricing had increased by a cumulative 83% by the first quarter, yet call volume surged by 400%. This simultaneous rise in price and volume, with supply lagging behind demand, forms the common backdrop for this round of collective price hikes. However, price increases can also be interpreted as seeking subsidies for ongoing infrastructure investments and presenting a more attractive revenue story for ongoing financing and IPO efforts. These two motivations are not mutually exclusive, as subsequent data in this article will demonstrate.

MiniMax's interim report provides a more rigorously audited perspective. First-half revenue from its open platform reached $73.929 million, a 703% year-on-year increase, accounting for 63.4% of total revenue (up from 30.3%), with second-quarter revenue ringing 82% higher quarter-on-quarter. Vice President Xue Zizhao attributed this to "optimizing inference efficiency while providing frontier model capabilities, achieving attractive cost-effectiveness."

Yan Junjie provided a more intuitive slope: token consumption in July was 20 times that of January. Zhipu confirmed demand most directly. Before the official release of GLM-5.3-Flash, Zhipu listed it anonymously as "Ox-Alpha" on OpenRouter and OpenCode, handling 50T token calls in five days and setting new traffic records for both platforms. The Chinese community named it "Niulai" (Ox Comes) before the model itself gained fame, with all this traffic running on domestic chips. Then came pricing.

Regular pricing for GLM-5.3-Flash's input and output is 53% and 62% of DeepSeek V4-Flash's off-peak prices, respectively; during the half-price period, these ratios drop to 27% and 31%. According to Artificial Analysis's Composite Intelligence Index, GLM-5.3-Flash scored 57, equal to Claude Opus 4.8 and higher than DeepSeek V4-Pro's official score of 53.

However, the opposite is true for cache hits. GLM-5.3-Flash's cache hit input price is 0.23 yuan, higher than V4-Flash's off-peak price of 0.05 yuan. Programming agents have high context repetition rates, with cache hit rates generally exceeding 90%. For these heavy users, DeepSeek remains cheaper. Both models have their strengths, but the pricing framework of "a lightweight flagship costing an order of magnitude less than its own flagship and half that of competitors" has been established.

Demand's sensitivity to prices is also real. According to LatePost, after DeepSeek's price hike, V4-Flash's calls on OpenCode halved, with this demand flowing to other domestic models.

───

2  44.6% vs. 17.9%: Profits Come from Computing Costs, Not Pricing ■

The most striking figure from The Information's disclosure is DeepSeek's API business segment gross margin of 82.9%, with an overall gross margin of 44.6%. MiniMax's interim report shows an overall gross margin of 17.9%, up from 12.1% year-on-year. The 26.7 percentage point difference between 44.6% and 17.9% is significant. International references at the overall level: OpenAI had a gross margin of 39% in Q1; Anthropic expects to raise its from 40% to 63% this year.

Both overall figures are mixed with their respective revenue structures. DeepSeek's overall 44.6% is 38 percentage points lower than its API segment's 82.9%, likely dragged down by limited monetization of C-end apps and other businesses; MiniMax's 17.9% includes 36.6% C-end revenue (Talkie, Hailuo AI), with Hailuo focusing on video generation, generally recognized (Note: " generally recognized " is translated as "widely recognized" in context, but kept as is here; in final text should be "widely recognized") to have far higher inference costs than text.

Structure explains part of the gap, but not all. Pricing does not explain it either, as MiniMax's models are not more expensive than DeepSeek's. The remaining variable lies primarily in costs: computing power consumed per token. The technical foundation for this calculation is the Mixture of Experts (MoE) architecture. V4-Flash has 284 billion total parameters, activating 13 billion per token; GLM-5.3-Flash has 320 billion total parameters, activating 18 billion.

Compared to the GLM-4.5 series, GLM-5.3-Flash reduces activated parameters from 32 billion to 18 billion and layers from 92 to 45; compared to GLM-5.3, its attention computation and KV cache are reduced to roughly one-third and one-fourth, respectively.

The commercial implication of these parameters is singular: single-token computing cost. MiniMax attributes its gross margin improvement to "enhanced infrastructure efficiency," referring to the same factor. DeepSeek's other figures further illustrate the point: revenue of 475 million yuan in the first seven months, with AI infrastructure investment of approximately 11 billion yuan, 23 times revenue. In 2025, this ratio was 25 times: revenue of approximately 47.5 million yuan, with investment of 1.2 billion yuan.

Revenue increased tenfold, yet the ratio barely moved. It remains an unprofitable company, with every yuan of revenue and financing ultimately converted into chip time.

Even discounting these leaked figures, the direction remains unaffected. The profit structure of this business is now clear: revenue is determined by token volume, gross profit by single-token computing cost, and pricing itself offers little defensive moat. Zhipu used a new architecture to reduce activated parameters to 18 billion, then employed a domestic cluster to lower costs further. This move struck at DeepSeek's cost structure, with the price list being merely the surface.

───

3  100,000 Domestic Chips Handle Real Traffic for the First Time ■

This release contained more information beyond the model: GLM-5.3-Flash's online inference runs on a cluster of over 100,000 domestically produced chips. According to LatePost, computing power suppliers may include Huawei, Moore Threads, and Hygon, though Zhipu officially declined to comment. Zhipu's blog describes it as "tens of thousands of domestic accelerators" connected by a self-developed high-bandwidth interconnection network. Zhipu did not first announce the release and then deploy; instead, it listed the model anonymously on OpenRouter, using real traffic from global developers for five days (50T tokens), before revealing that all these requests were handled by domestic chips.

Previously, domestic chips running large models were mostly seen in demo presentations at launches. This time, Zhipu first handled real billing before showing its hand. Zhipu's disclosed engineering details are candid about the chips' shortcomings. The bottleneck lies in memory capacity and bandwidth, especially challenging for supporting 1 million token contexts. Zhipu's solution was to build its own inference engine based on SGLang: W8A8 quantization, INT8/FP8/BF16 mixed cache quantization, and a three-stage separate scheduling for encoding-prefilling-decoding, boosting end-to-end performance of the same hardware by threefold. Zhipu concludes that single-token hardware efficiency and costs now match those of mainstream NVIDIA GPUs.

If this conclusion holds, Zhipu's low pricing is not burning money for market share but represents sustainable pricing. This would also explain its confidence in setting prices at one-tenth of its flagship model. However, only inference has been validated thus far; training remains unaddressed industry-wide. The specifics of the 100,000 chips—which models and their proportions—remain undisclosed. Compensating for hardware shortcomings with aggressive software optimization means redoing engineering work with each chip change. However, for this business, achieving domestic computing power for inference alone is enough to alter the landscape, setting a new cost baseline for price wars.

───

4  Disagreements Over a Trillion-Yuan Valuation: What Is Capital Pricing? ■

The three companies' revenue scales are nearly identical: MiniMax at 786 million yuan in the first half, DeepSeek at 475 million yuan in the first seven months, and Zhipu at 724 million yuan for all of last year (accelerating this year). Yet their market caps/valuations differ by orders of magnitude. Zhipu's current valuation is approximately 440 billion yuan, nearly five times MiniMax's; unlisted DeepSeek has a private market valuation of 450-500 billion yuan.

Divergence stems from strategic narratives. Zhipu has transformed from "China's OpenAI" to "China's Anthropic" by focusing on coding and B-end APIs. DeepSeek is a global symbol of open-source and cost-effectiveness. MiniMax derives over 60% of revenue from overseas, with C-end AI-native products still accounting for more than one-third, though open platform revenue has rapidly increased from 30.3% to 63.4%. MiniMax's revenue structure is actually shifting toward B-end, but capital market revaluation has not kept pace with this transition, leaving it without a position in the current B-end narrative. Even at the most optimistic estimates, these figures seem exorbitant. Anthropic's valuation reached $965 billion after its May funding round, roughly 21 times annualized revenue.

Goldman Sachs rated Zhipu "neutral" in July, citing a ~1600% stock price increase in six months and a "largely reasonable" valuation. JPMorgan maintained an "overweight" rating with a target price of HK$1,800. Capital stories are converging toward the same structure: go public to secure funding, convert every yuan into computing power, and wait for revenue to catch up.

MiniMax has over $3 billion in cash on hand, including HK$16 billion from its July placement; Zhipu raised HK$31.4 billion in July and plans an additional 15 billion yuan on the STAR Market; DeepSeek completed a 50 billion yuan initial funding round in June and is targeting another 50 billion yuan in its second round, having already spent 11 billion yuan in the first seven months. All three are using capital market funds to purchase the same ticket: chip time. Returning to August 26.

MiniMax's interim report proves rising demand, DeepSeek's data confirms gross margins, and Zhipu's new model shows accelerating supply entry. The lightweight flagship business model is validated, but it has been a buyer's market from day one. Demand is strong enough to support collective price hikes, yet supply is abundant enough that any price increase creates space for competitors.

DeepSeek spends 11 billion yuan on chip time, Zhipu uses 100,000 domestic chips to lower bills, and MiniMax trades inference efficiency for cost-effectiveness. The final outcome will depend not on a 57 vs. 53 score in evaluation rankings but on who first reduces single-token costs by an order of magnitude. The price war has just begun.

Note: Data in this article comes from public sources and media reports, with DeepSeek's financial data being leaked by a single source and not confirmed by the company. This article does not constitute investment advice.

- END -

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.