09/11 2026
536

Cover | Illustrated by Qianbidao
In just six months, computing power services worth nearly 900 million yuan were sold, with net profits soaring sevenfold. This was the 'harvest' Parallel Technology reaped in the first half of 2026. Similarly, several AI giants reported 'barely any idle GPUs'.
Contrastingly, a thousand-card intelligent computing center in a western city had an occupancy rate of less than 50%, with the actual utilization of deployed servers below 30%. Yet, its annual operating costs exceeded 30 million yuan. Industry insiders estimate that the national utilization rate of intelligent computing centers hovers around 20%–30%.
On one hand, there's a scramble for computing resources; on the other, many resources remain idle.
Leading AI companies are in dire need of high-performance, large-scale, and stable computing clusters, which remain in short supply. Meanwhile, some regions have preemptively built computing power infrastructure but have yet to find suitable clients or applications.
Peng Yuanyang, the founder of Lingkong Intelligence, told Qianbidao that many companies are not averse to using AI but are uncertain about which models to employ, how much computing power to purchase, how to process data, deploy systems, or whether their investment will ultimately yield revenue or cost savings. There's still a missing 'link' between computing power and corporate needs.
- 01 - Who is Profiting First?
Peng Yuanyang observed that some small computing power rental companies in Shenzhen are thriving. They may not possess ten-thousand-card clusters but are willing to accept small orders for a few or dozens of cards, staying close to clients and better understanding their model requirements and usage durations.
Parallel Technology has taken this business model even further. In the first half of 2026, the company's computing power service revenue reached 895 million yuan, accounting for 94.62% of its total revenue. Enterprise client computing power service revenue surged to 705 million yuan, up 168.07% year-on-year.
Its core strategy doesn't rely solely on selling cards from a single data center but on integrating computing power resources from different regions and clusters through a computing network, followed by efficient scheduling and delivery.
This is a business model focused on improving turnover rates.
The recently launched 'Computing Power Supermarket' in Jiangning, Nanjing, provides another case in point. During its trial operation, only 20 companies participated, yet they collectively invoked models 1.308 million times, consuming 51 billion Tokens.
China Electric Power Transformer Co., Ltd. previously wanted to use large models for bid document reviews and code reviews but was deterred by computing power prices and model selection. After accessing the 'Computing Power Supermarket,' the company could directly select GPUs and mainstream large models on demand.
It's not that companies lack AI demand; rather, the traditional way of purchasing computing power was too cumbersome.
In the past, a short-term project might require renting an entire server and signing a long-term contract. Now, more and more platforms charge by card-hours, core-hours, or even Tokens. Companies only pay for actual usage, while computing power companies can allocate a batch of GPUs to more orders.
This is precisely what 'Computing Power Supermarkets,' 'Computing Power Banks,' GPU rentals, and computing power scheduling platforms are doing: repackaging scattered, heterogeneous, and idle computing power into purchasable standard services.
However, matching supply and demand is not highly profitable.
Parallel Technology's gross margin for computing power services in the first half of the year was 16.97%. This indicates that simply selling cards from A to B doesn't offer as much profit as imagined. What truly determines profitability is whether the same batch of GPUs can be utilized by more clients for longer periods.
- 02 - A Tale of Two Markets: Scarcity and Idleness
Why are many intelligent computing centers 'sleeping' despite the high demand for computing power?
Many localities build intelligent computing centers based on infrastructure logic: first determine the scale, purchase equipment, build the center, and then wait for demand to arrive.
But computing power doesn't automatically attract users like a highway does after construction.
Investigations by Yicao, a publication under Xinhua News Agency, revealed that a central province planned three intelligent computing centers with a total capacity of 1,200 PFLOPS. Yet, local digital economy companies numbered fewer than 100, with actual demand below 200 PFLOPS. Two districts in an eastern city built intelligent computing centers for biomedical and industrial internet applications, respectively, but fewer than 10 companies actually moved in.
According to data released by the National Data Administration, as of June 2026, China had over 500 operational intelligent computing centers and nearly 400 projects under construction or planning.
If the industry average utilization rate for these 500 centers is only 20%–30%, and an intelligent computing center typically needs to reach about 40% occupancy to break even, then a rough estimate suggests that 250–350 centers may still be operating at a loss or struggling to cover full costs.
Large model training requires high-performance chips, high-speed interconnects, storage, cluster stability, and a mature software ecosystem. Industrial, financial, and other scenarios may impose different requirements on latency, data security, and privatization. Just because a locality has idle GPUs doesn't mean it can directly handle training tasks for leading model companies.
This explains why 'giants lack computing power' and 'local computing power remains idle' occur simultaneously.
Moreover, most ordinary companies don't know how to find suitable computing power.
A factory might only need visual quality inspection, a retail company might only want customer service integration, and another company might only require a private knowledge base. Project scales may involve just a few or dozens of cards, necessitating someone to handle model selection, data, deployment, security, and costs.
Peng Yuanyang said there's a lot of 'dirty work' involved. Large companies prefer to build standard platforms and rarely enter fragmented scenarios one by one. Small companies are willing to do so, but the service chain is not yet mature.
What computing power centers often lack is not the next batch of GPUs but the next batch of orders.
- 03 - The True Value Lies in Tokens
Peng Yuanyang predicts that for most ordinary companies, what they will purchase long-term is inference computing power—directly invoking existing models to integrate AI into customer service, marketing, code development, industrial inspection, and knowledge bases.
Training has project cycles, but inference occurs continuously with business operations.
This is transforming the financial models of computing power companies.
QingCloud Technology's AI computing power cloud service revenue was 43.36 million yuan in 2025, still declining year-on-year, but its gross margin rose from 18.99% to 31.62%. One reason the company cited was improved utilization of computing power resources and optimized operating costs.
Another company demonstrates the challenges of a capital-intensive business.
Runjian Co., Ltd.'s computing power network business revenue was 519 million yuan in the first half of 2026, up 50.31% year-on-year. It claimed that the 'Token Factory' business model was largely closed-loop; however, the business's gross margin was only 6.97%. The company explained that one main reason was high depreciation costs in the early stages of fixed asset investment.
These two sets of figures point to the same fact: acquiring GPUs is just the beginning. As long as cards remain idle, depreciation continues. Once utilization improves, each additional task can help dilute fixed costs.
Therefore, the next wave of opportunities may not belong to those who build another intelligent computing center but to those who can generate more Tokens from the same batch of cards.
GPU scheduling, elastic inference, inference acceleration, multi-model routing, and cost optimization represent the first layer of opportunities. Beneath that lie model deployment, privatization, enterprise knowledge bases, Agents, and vertical industry applications.
The latter is closer to where clients' money is.
When AI remains at the Demo stage, computing power consumption is one-time. Once it enters production, with factories conducting daily quality inspections, customer service handling daily orders, and programmers writing code daily, computing power becomes a recurring order.
Peng Yuanyang believes that intelligent computing centers should focus on providing stable, low-cost infrastructure rather than trying to handle all applications themselves. Companies that truly understand industries should then integrate models and computing power into corporate operations.
Over 500 intelligent computing centers already exist nationwide. In the next phase, the industry will not compete solely on 'who has more cards.'
More importantly, it will be about who can keep a GPU idle for one minute less and generate one more billable Token; who can turn these Tokens into business services that enterprises are willing to renew long-term.
This article does not constitute any investment advice. Relevant reporting from Saige Avenue and Yicao was also referenced in the writing of this article, and thanks are extended accordingly.
