Zhipu’s Dual-Track Strategy: Balancing Growth and Profitability

09/10 2026 474

Source | Bohu Finance (bohuFN)

Following MiniMax, Zhipu has ramped up its efforts.

Recently, Zhipu—often dubbed the “first global large-scale model stock”—released its inaugural interim results since going public. Revenue reached 954 million yuan in the first half of the year, marking a 399.7% year-on-year surge and surpassing the total for all of 2025. Revenue from open platforms and APIs skyrocketed 27-fold, now accounting for 86.5% of total revenue. Meanwhile, API gross margin improved from -0.4% in the same period last year to 24.6%.

During the earnings call, management revealed that as of late August, Zhipu’s ARR (Annual Recurring Revenue) had climbed to $1.6 billion, up 60% from $1 billion in early July.

Overall, Zhipu has delivered a strong performance—shifting its revenue mix, achieving explosive growth, and further enhancing API profitability. Zhipu is no longer merely a large-scale model company surviving on project-based work but is steadily positioning itself as the “Chinese counterpart to Anthropic” in market narratives.

However, the market’s response has been somewhat muted. In June, Zhipu’s market capitalization briefly exceeded one trillion Hong Kong dollars but has since retreated to around 520 billion Hong Kong dollars. The company’s stock price opened roughly 4% lower following the earnings release.

Despite Zhipu’s explosive revenue growth, it has fallen short of market expectations. While earnings have risen, so have losses.

To a certain extent, Zhipu and MiniMax face similar challenges.

They have demonstrated that companies can generate revenue based on model capabilities. However, to truly “thrive,” they must retain profits—a feat that requires balancing intellectual premium with computing costs.

01 ARR Surges, Zhipu Accelerates

The most notable shift in Zhipu’s interim report lies in its revenue mix.

According to the report, in the first half of 2026, revenue from open platforms and API businesses soared to 825 million yuan, a 2735.7% year-on-year increase. Its share of total revenue rose from 15.2% in the same period last year to 86.5%. Meanwhile, revenue from localized deployments fell from 162 million yuan to 129 million yuan, with its proportion dropping from 84.8% to 13.5%.

This clearly signals that Zhipu is breaking through its business ceiling.

Historically, Zhipu primarily provided model localization services to institutional clients in government, finance, and energy sectors. While stable, this business faced challenges such as scalability issues and a clear growth ceiling.

Now, open platforms and API services have emerged as the company’s new revenue engines.

By late August 2026, Zhipu’s ARR had reached $1.6 billion. In August alone, it hit $133 million—roughly equivalent to the revenue generated in the first half of the year.

Behind this ARR surge lies simultaneous growth in user volume and willingness to pay. As of August, Zhipu’s MaaS platform had over 7.4 million registered users, a 144% increase from the beginning of the year. Paid daily active users grew by 603% over the same period.

During this timeframe, Token usage on Zhipu’s MaaS platform increased more than 40-fold from the start of the year, with Coding Plan usage growing over 23-fold. The average selling price of APIs rose by approximately 101%.

These two sets of data are crucial—they indicate that Zhipu’s revenue growth is not the result of “trading price for volume.” Instead, users are willing to maintain high-frequency usage at higher prices, reflecting Zhipu’s pricing power over its models to some extent.

Liu Debing, Chairman of Zhipu, stated during the earnings call, “If last year’s keyword for Zhipu was the ‘upper limit of intelligence,’ this year’s keyword is ‘capability delivery.’”

This can be interpreted as follows: When a company’s model capabilities are insufficient, customers essentially purchase delivery services, with models serving merely as tools to help integrate, customize, and deploy models into data centers. However, sustainable and explosive revenue growth is difficult to achieve under this model.

But when a company’s model capabilities can independently drive entire engineering projects, what customers purchase becomes a capability applicable in scenarios such as coding and collaboration. Users call upon models based on tasks, generating sustainable revenue.

Zhipu’s management summarizes this evolution as “selling models → selling calls → selling subscriptions → selling end-to-end task results.” Once models are fully operational, the revenue flywheel naturally begins to turn.

Moreover, the smarter the model, the more willing customers are to pay higher prices for higher-value task processing capabilities—further boosting revenue. Zhipu’s management believes, “The upper limit of intelligence determines pricing power, while the scale of Token consumption determines value.”

While everything appears to be moving in a positive direction, overseas investment institutions remain cautious.

Despite Zhipu’s nearly fourfold revenue growth in the past six months, it still falls about 30% short of analyst expectations—primarily due to delivery challenges.

Overseas investment institutions base their revenue expectations on Zhipu being labeled the “Chinese counterpart to Anthropic.” However, compared to Anthropic’s revenue growth curve (ARR surging from $9 billion to $65 billion in seven months), Zhipu has yet to achieve “absolute leadership” domestically.

Additionally, while ARR reflects the sustainability of future revenue, whether users’ subscription habits can be maintained and whether all this revenue can be converted into operating income will depend on factors such as model capabilities, API pricing, and competitor actions.

In response, Xiao Lei, Secretary of Zhipu’s Board of Directors, acknowledged that Zhipu’s ARR curve is not smooth but shows “pulsed” growth. He also compared Zhipu to Anthropic, stating that Zhipu is about 17 months behind—but that the gap is narrowing.

02 High Earnings, but Also High Losses

However, maintaining Zhipu’s leading position in model capabilities comes at a cost—as reflected in its less-than-impressive profit figures.

In the first half of the year, the company incurred a loss of 2.072 billion yuan, a 12.1% year-on-year narrowing. However, after excluding items such as stock-based compensation and changes in the book value of financial instruments, the adjusted net loss widened from 1.75 billion yuan to 1.96 billion yuan. Operating losses also expanded to 2.15 billion yuan.

The reasons for these losses are closely tied to the nature of large-scale models. In the first half of the year, Zhipu’s research and development expenses reached 2.13 billion yuan—2.2 times its revenue for the same period and more than eight times its total gross profit. In other words, the money earned was insufficient to cover Zhipu’s new round of investments.

The direct result? A near-halving of Zhipu’s gross margin.

During the reporting period, Zhipu’s overall gross margin fell from 50% to 26.4%. The company attributed this decline to the shift in revenue mix and the rapid expansion of cloud-based deployment services—whose gross margins are still climbing—structurally diluting the overall level.

This is a collective challenge faced by model companies. While the MaaS model—which charges based on usage volume—can bring scale, as user numbers swell, computing costs climb simultaneously.

Take Zhipu as an example: In addition to training costs, its computing investments are also visibly increasing.

On the one hand, as model capabilities continue to improve, inference itself demands higher computing power.

Xiao Lei revealed during the earnings call that Zhipu’s coding subscription product, Coding Plan, was effectively on sale restriction in the first half of the year. It was only fully reopened in August after significant improvements in the company’s computing capacity.

The reason for these computing capacity improvements? Zhipu “went all in.”

According to the interim report, Zhipu’s prepaid computing service fees and other expenses surged from 116 million yuan to 1.244 billion yuan. Computing service fees payable rose from 727 million yuan to 1.427 billion yuan—likely representing prepaid inference and training fees to cloud providers or chip clusters.

Interestingly, the prepaid and payable computing service fees are almost equal in volume—indicating sustained tightness in computing supply within the large-scale model industry. To the point where production capacity must be locked in with advance payments. To earn “tomorrow’s” money, large-scale model companies must first bear “today’s” pressure.

On the other hand, there is the issue of intelligent infrastructure.

Zhipu stated that it fully implemented an All-in-Infra strategy in 2026 and has now completed the construction of a 1GW-scale domestic AI computing data center. Additionally, through its ecosystem fund, Starlink Capital, it has invested in AI Infra companies such as Jiliu Technology, Infinite Chip Space, and Silicon-Based Flow—covering areas such as computing clusters and inference optimization. It also acquired Zhongke Jiahe to further enhance its underlying model engineering capabilities.

Rising costs continue to squeeze profits, creating an awkward situation where, despite increases in Zhipu’s revenue and gross profit, most of the profits merely pass through the books before ultimately flowing to computing suppliers.

Barclays’ AI industry research also highlights this issue, noting that for every $100 in revenue generated by model companies, about $35–40 flows to the three major cloud platforms—Amazon AWS, Microsoft Azure, and Google Cloud Platform—in the form of inference computing fees.

The pressure on input and output for large-scale model companies is forming a “kill zone”:

Either use stronger model capabilities to increase users’ willingness to pay, or provide “good enough” models at more cost-effective prices—forcing large-scale model companies to find a balance between SOTA (State-of-the-Art) status and cost-effective models.

But in the business world, “both high quality and low cost” is almost a false proposition. Why does Zhipu think it can “have both”?

03 The ‘Ox-Alpha Moment’ for Domestic Large-Scale Models

Five days before the interim report’s release, Zhipu officially launched and open-sourced the GLM-5.3-Flash model. Previously anonymous as “Ox-Alpha,” this model went viral on overseas platforms and was dubbed “Ox-Alpha” by the Chinese developer community.

Coincidentally, the film Ox-Alpha, released this summer, perfectly illustrates the concept of “having both”—as long as a product addresses user pain points, low costs can still yield high returns.

Take GLM-5.3-Flash as an example: First, it packs a punch. As Zhipu’s first multimodal model since focusing on coding strategy, it features native multimodal capabilities—supporting image and video input.

During its anonymous release, GLM-5.3-Flash topped the OpenRouter usage charts within just 12 hours of launch. On the OpenCode platform, it ended DeepSeek’s 56-day streak at the top.

But a capable model doesn’t have to be expensive.

GLM-5.3 Flash has 300 billion parameters—roughly 40% of GLM-5.3’s parameter count. By reducing the number of parameters invoked per inference, it leaves more room to lower inference costs and improve generation speed.

At the same time, GLM-5.3 Flash is the first flagship model to fully utilize domestic chip clusters. Through a self-developed inference engine and a separated architecture, its unit Token inference cost has dropped by 80% since the beginning of the year.

Additionally, Zhipu views post-training as a critical link in enhancing model capabilities at this stage.

Management mentioned that, while keeping total and active model parameters unchanged, by increasing investments in post-training, reinforcement learning, and long-term task feedback, GLM-5.3’s end-to-end task completion rate improved by over 50% compared to GLM-5.2 with the same architecture and scale.

This is precisely the “first path” Zhipu aims to take:

In the fiercely competitive large-scale model industry, price wars have become an unignorable factor.

At this point, product tiering becomes essential—using lightweight, cost-effective models to attract price-sensitive customers, acquiring them at low cost, and then converting these users into highly engaged customers as usage volume grows and scenarios deepen.

Recently, Zhipu also opened an official flagship store on Tmall, selling subscription packages for the GLM Coding Plan (individual and team versions)—packaging AI services into standardized products that are easier for consumers to understand and reaching a broader consumer base.

Zhipu’s “second path,” meanwhile, lies in next-generation SOTA models.

Tang Jie, founder of Zhipu, stated during the earnings call that his understanding of SOTA is not “how long it leads” but rather “how strong its ability to sustain leadership is.”

Thus, the technical route for the next-generation GLM-6.0 large-scale model is Full Self-Training—aiming to serve users at scale from day one of release rather than creating models with massive parameters that cannot be economically deployed.

Zhipu judges that next-generation SOTA models will be cutting-edge models capable of opening up new task boundaries and significantly improving task success rates. As long as they can handle high-value tasks, they will still command a premium.

On one front, Zhipu aims to acquire customers via cost-efficient models to expedite market expansion. On the other, it seeks to uphold pricing power through highly intelligent models, thereby boosting the gross margins per unit of Token and safeguarding profits—this dual strategy epitomizes Zhipu's pursuit of 'the best of both worlds.'

However, for these two strategies to be successful, Zhipu needs to demonstrate one crucial point: even when its models become more cost-effective, their state-of-the-art (SOTA) status remains unassailable.

The success of Zhipu's endeavors will be evident in its upcoming earnings report:

Can the Annual Recurring Revenue (ARR) maintain its growth trajectory and translate into reported revenue? Can SOTA models continue to elevate model prices? Can the gross margins of the API business keep improving, resulting in a sustained reduction in overall company losses?

Until then, Zhipu must strive to extend its lead as much as possible, affording itself more time to provide answers and await its 'Ox-Alpha Moment.'
The cover image and illustrations are the property of their respective copyright holders. If the copyright owners deem their works inappropriate for public viewing or believe they should not be used without compensation, please notify us promptly, and our platform will make immediate corrections.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.