Farewell to 'Card Stacking': 2026 Marks China's Computing Power Shift to 'Token Value Output' Era

08/19 2026 513

Behind this vigorous intelligent computing 'city-building movement' lies an extremely rational 'low-cost Token production defense.' Only by thoroughly reducing the comprehensive production cost per Token to be as cheap as 'tap water' and 'electricity' can the massive demand for Agent-era calls transform from a heavy user's carnival into a sustainable long-term business.

Author | Dou Dou

Produced by | Industrial Home

In 2026, data center hubs in northwest China are witnessing a near-frenzied 'city-building movement.'

In June, DeepSeek bets big on a 1GW-scale mega intelligent computing base, landing directly in Ulanqab, Inner Mongolia. In July, Shuguang 8000 announces the first fully domestic 100,000-card supercluster officially connecting to the National Supercomputing Internet. In August, Alibaba Cloud officially declares reducing AIDC delivery cycles to just 100 days, aggressively expanding along the northwest green energy belt.

Notably, unlike previous focuses on hardware card counts or theoretical peak computing power displays, the emphasis has shifted to system engineering. Some ongoing changes indicate that the true benchmark for computing infrastructure seems to have transitioned from 'how many cards are owned' to 'whether 100,000 cards can stably, continuously, and bottleneck-free collaborate on a real large model task.'

Questions arise: computing infrastructure is moving from simple hardware stacking to a brutal system engineering test. What is the fundamental driving force behind this shift? Why, as card counts soar, does the entire industry collectively fall into anxiety over 'effective output?'

I. AI Coding Era Sees Call Volume Surpass 140 Trillion

The industry's crazy 'money-burning city-building' must stem from explosive downstream demand outstripping supply.

In fact, AI technology implementations have continuously evolved over the past few years, with growth engines shifting from Chat AI to Agents capable of sustained model calls. Consequently, model call volumes have skyrocketed. The National Data Bureau disclosed that by March 2026, China's daily average Token call volume exceeded 140 trillion, growing over 1,000-fold from 100 billion in early 2024.

This surge in call demand is reshaping the global model usage landscape. OpenRouter platform data shows that from February 9-15, 2026, Chinese model call volume reached 4.12 trillion Tokens, surpassing the 2.94 trillion Tokens for US models during the same period.

Why is this happening?

The answer is straightforward. Chat AI typically involves question-and-answer exchanges, while Agents must autonomously plan tasks, call tools, read returned results, and continue execution and verification based on those results. A single task often requires multiple rounds—even dozens—of continuous model operation, resulting in far higher Token consumption than ordinary dialogues. In short, what drives this call volume surge is that individual tasks are becoming increasingly 'heavy.'

Among these, AI Coding represents the most typical Token-consuming scenario.

A study using eight cutting-edge models and OpenHands Agent on SWE-bench Verified reveals that Agentic Coding consumes approximately 3,500 times more Tokens per round than single-round code inference and 1,200 times more than code chatting, with input Tokens far exceeding output Tokens at an average input/output ratio of 154:1.

As these high-consumption tasks move from limited experiments to large-scale applications, the Token structure at the platform level also changes. OpenRouter and a16z's '2025 AI Usage Report,' based on over 100 trillion Token anonymous metadata, shows that programming tasks' Token share rose from 11% in early 2025 to over 50%, becoming the platform's largest single usage category. Meanwhile, the most prominent general-purpose Agent framework in March 2026, OpenClaw, contributed only about a quarter of the platform's weekly consumption.

From the current perspective, it is not all general-purpose Agents but AI Coding—with its clear workflows and willingness to pay—that truly converts high-frequency Agent calls into stable Token consumption.

This judgment is directly validated in service providers' financials.

Zhipu is among China's early model vendors concentrating resources on Coding. Its GLM Coding Plan has surpassed 242,000 global paying developers, with Token call volume surging 15-fold in six months. More notably, in Q1 2026, GLM API prices increased cumulatively by about 83%, yet call volume still grew roughly 400% during the same period. Its model API's ARR reached 1.7 billion yuan in March, soaring 60-fold year-on-year, and hit 1 billion USD in overall ARR by July. From 100 million to 1 billion USD, Anthropic took 15 months; Zhipu achieved it in just 5.

Moonshot AI's growth curve is even steeper. Its long-context programming-focused K3 hit user request limits for existing clusters just 48 hours after release. MiniMax M2.5 surpassed 3.07 trillion Token calls within seven days of launch.

While model vendors see demand explode for individual products, cloud vendors witness rapid expansion of the entire Token market.

Calculated by MaaS call volume, Volcano Engine's Doubao large model saw daily average Token calls grow from 2 trillion at the end of 2024 to 63 trillion by the end of 2025, reaching 120 trillion in March 2026 and 180 trillion by June—a growth exceeding 1,500-fold in two years. IDC data shows Volcano Engine capturing 9.5% of China's public cloud MaaS market by call volume. Internally, one of its fastest-growing businesses is the code tool Trae, with daily Token consumption rising from an initial 8 billion to the 300-400 billion range.

Alibaba Cloud presents a different trajectory.

In Q4 FY2026, its AI-related revenue reached 8.971 billion yuan, accounting for over 30% of external commercial revenue for the first time. Customer numbers on its Bailian platform grew eightfold year-on-year, with daily average Token revenue growing about 15-fold over the past five months. Alibaba Cloud's financial report (financial reports) and business disclosures repeatedly cite growth sources including Bailian model calls and AI Coding products like Qoder.

Overall, Agents have unlocked far greater Token consumption than traditional chatting, while AI Coding has led the way in converting this demand into stable revenue. This has created insatiable demand for computing power, bursting supply and ushering in a golden age of 'demand explosion—price hikes—scale expansion' for the entire AI industry.

II. The Computing Black Hole of the Agent Era Amid Heavy Users' Carnival

However, a call volume explosion does not equate to a viable business model.

Under normal circumstances, stronger demand should amplify economies of scale, boosting both revenue and profits for vendors. Yet after entering the Agent and AI Coding phase, the industry faces a paradox: the more popular the product, the more frequently vendors impose purchase limits, price hikes, or even suspend new users.

Moonshot AI exemplifies this. Just 48 hours after Kimi launched K3, it had to suspend new C-end subscription entries due to overwhelmed computing power. If new demand could steadily convert into new profits, vendors would have no reason to voluntarily 'close the door' during peak demand.

Bo Wenxi, Vice Chairman of the China Enterprise Capital Union, argued that 'the core issue is that K3 passed the demand test but failed the supply and profit test.' Other experts attribute this to computing shortages and service guarantees, reflecting a shift among independent large model companies from user acquisition to focusing on revenue and efficiency.

This contradiction becomes even more pronounced in model vendors' pricing strategies.

Alibaba Cloud's Bailian Coding Plan Lite package halted new purchases from March 20, 2026, and discontinued renewals and upgrades from April 13, eliminating the previous 40-yuan low-price tier. Zhipu's GLM partially switched to fixed-time limited sales, effectively making certain packages unattainable—its Max tier opened only about 20% of daily quotas. Additionally, Alibaba Cloud, Tencent Cloud, MiniMax, and Xiaomi have all introduced or renamed packages to Token Plans, replacing per-request billing with fine-grained Credits/Token metering.

Previously, Chat AI's Token consumption was relatively smooth and predictable, but this logic breaks down in the Agent era—especially for AI Coding. As mentioned earlier, certain Agentic Coding tasks can reach input/output Token ratios of 154:1.

This rapidly widens cost disparities among users. An ordinary user might call Tokens dozens of times daily, while a heavy user integrating Coding Agents into development workflows could keep models running continuously for hours or even all day. Token and computing costs per user can thus differ by tens or even hundreds of times.

Problems emerge. Monthly subscriptions essentially use fixed revenue to cover uncertain computing consumption, but Agents amplify this uncertainty to extremes. The stronger the product's capabilities, the more professional developers and high-frequency users it attracts. The more these users engage, the higher the vendor's marginal computing costs become. Once capacity limits are exceeded, vendors must ultimately regain control through quotas, limits, split benefits, or switching to Token-based billing.

If Token gross margins were extremely high, vendors might subsidize heavy users with light users, but the opposite is true. For model vendors, profit margins are already thin. Sun Yuanhao, CEO of Xinghuan Technology, noted that 'the root cause of fierce competition in the Token factory track (track) lies in the tiny price gap between Token prices and computing procurement costs, squeezing profitability.'

Coupled with persistently high computing hardware procurement costs and multiple rounds of price wars in the frontend MaaS market, absolute gross margins per Token remain extremely low. When heavy users' Agent tasks trigger long-context retrieval, already thin profits vanish instantly, turning into net losses.

Beyond frontend business model distortions, extreme waste in backend computing supply further inflates per-Token amortized costs.

Data from the China Mobile Research Institute shows that GPUs in traditionally modeled intelligent computing centers typically average below 30% utilization. Computing power physically remains inefficiently absorbed by the market, with many data centers running idle tasks or stalling due to network packet loss or memory fragmentation. These idle and loss-related costs ultimately factor into every Token generated, exacerbating already narrow price gaps.

This creates 2026's most paradoxical AI industry scene: on one side, exploding usage desire and soaring call volumes; on the other, vendors lamenting razor-thin margins and forced price hikes. The root cause is that traditional SaaS monthly subscription logic completely collapses before the computing black hole of the Agent era. Demand is real, but at current Token production costs, no vendor can sustain heavy users' carnival with existing business models.

III. 'Cut Electricity Bills, Open Pathways, Smart Division of Labor, Peak Shaving': Industrialized Token Factories Take Shape

AI industry evaluation criteria have shifted.

Technologically, competition has moved from merely stacking TFLOPS peak computing power to 'effective computing power' metrics like model Flops utilization, linearity, long-term stable operation duration, and checkpoint resumption time. Commercially, model vendors no longer pursue user growth alone but focus on high-quality Token output per watt, per card, and per yuan—and whether these Tokens can sell above cost.

The problem is that demand-side price hikes, traffic throttling, and reduced low-price packages are merely short-term fixes. The real battleground must shift to reducing underlying Token production costs.

To win this cost defense war, suppliers have launched system engineering reforms across four layers: 'cutting electricity bills, opening pathways, smart division of labor, and peak shaving.'

The first step is relocating computing power to cheaper electricity zones.

A supply-side 'city-building movement' unfolds across northwest China.

Eastern coastal data centers typically face industrial electricity prices of 0.6–0.8 yuan/kWh, while northwest regions like Ulanqab, Inner Mongolia, and Qingyang, Gansu, boast abundant wind and solar green energy, with some areas' comprehensive electricity prices dropping to 0.25–0.3 yuan/kWh. For intelligent computing centers reaching hundreds of megawatts or even 1GW scale, this slashes Token energy costs at the source.

But cheap electricity doesn't automatically maximize card utilization. In large-scale clusters, communication bottlenecks, network congestion, data packet loss, and node failures leave GPUs idle. The larger the scale, the more any minor bottleneck amplifies, resulting in 'many cards but few actually working.'

Thus, the first fully domestic 100,000-card supercluster and Huawei's Atlas 950 SuperPoD hypernode architectures primarily solve 'path widening.' By implementing higher-speed interconnection networks and tighter node collaboration, they link scattered servers into a larger computing system, reducing communication wait times and ensuring more GPU cycles go toward actual computation.

If hypernodes address 'how cards collaborate,' software layers tackle 'what different cards should do.'

Facing Agent scenarios' extreme 154:1 input/output ratio, assigning the same GPU to both data ingestion and output causes severe mismatch—'overfed then starved.' Runtime systems like SenseTime's Large Device and Zhongke Jiahe's software and compilation layers introduce 'assembly line division of labor.' They assign high-computing-power, fast-reading GPUs to ingest massive historical contexts and high-memory, high-bandwidth GPUs to rapidly output code.

This assembly line approach of specialized cards for specific tasks completely eliminates memory fragmentation and computational waiting bubbles, significantly enhancing the actual Flops utilization across the entire cluster.

Even if electricity is cheap and machines connect quickly, if the data center is overwhelmed during the day and largely idle at night when everyone is sleeping, the depreciation costs for the machines remain high. As a result, compute scheduling has evolved from simply "sending tasks to wherever there are available cards" to a more sophisticated approach of determining "which tasks, at what times, should be sent where."

Conversations requiring millisecond-level responses can remain on eastern nodes closer to users; Agent tasks that need to run for minutes or even hours can be scheduled to large-scale western clusters; non-real-time tasks can leverage off-peak nighttime hours when electricity is abundant. Some platforms are also beginning to guide users to actively shift their usage through nighttime discounts and flexible Token pricing. This ensures that high-value real-time tasks occupy the most expensive resources, while delay-tolerant tasks consume the cheapest and most idle compute power.

Source: Implementation Opinions of the National Data Bureau on Deepening the "East Data West Calculation" Project

From low-cost green power in the northwest, to super-nodes solving communication bottlenecks, and then to PD decoupling and cross-regional, cross-temporal scheduling, this combination of "energy + hardware + software + scheduling" transforms previously scattered and inefficient GPU stacks into high-throughput, high-utilization "industrialized Token factories."

Overall, behind this vigorous movement to build intelligent computing "cities" lies an extremely rational "battle to produce Tokens at low cost." Only by thoroughly reducing the comprehensive production cost per Token to be as cheap as "tap water" and "electricity" can the massive demand for calls in the Agent era transform from a celebration for heavy users into a sustainable long-term business.

IV. New Survival Rules for Five Industrial Roles Under the New Compute Paradigm

When the ultimate goal of compute competition shifts from "stacking GPU quantities" to "reducing the comprehensive output cost per Token," this fundamental change in underlying logic is no longer just an issue for data centers and electricity bills. Instead, it rapidly propagates along the industrial supply chain, triggering a realignment of positions across the entire industry.

The fact is that every player, from chip giants, cloud providers, and large model unicorns to software Infra service providers and end-user enterprises, will be drawn into a survival elimination race centered around "effective compute power."

First are chip and hardware vendors.

Vendors such as Huawei, Sugon, Hygon, and Moore Threads are evolving from "selling individual cards/competing on theoretical TFLOPS peaks" to "selling software-hardware integrated super-nodes and hyper-intelligent fusion systems." In the future, the performance gains from individual cards will peak, and network interconnection and ecosystem adaptability will determine survival. Hardware form factors will accelerate toward cabinet-style, pooled super-nodes and deeply integrate with upper-layer compute scheduling and heterogeneous compilation frameworks.

Cloud providers like Alibaba Cloud, Tencent Cloud, and local AIDC operators, along with intelligent computing center operators, will transition from "selling compute power" to "low-cost production and refined sales of Tokens." Cloud providers such as Alibaba Cloud and Tencent Cloud will compete on ultimate delivery efficiency and low-cost green power layouts in northwest China while dismantling monthly subscription models at the front end and fully introducing Token Plans. Intelligent computing center operators like Paratera Technology will achieve over 90% utilization through precise scheduling, widening the gap with traditional local data centers that operate at less than 30% utilization.

Future differentiation will intensify. Small and medium-sized intelligent computing centers unable to access unified scheduling networks and operating at low utilization will face elimination or be absorbed into larger networks. Leading cloud providers will evolve into high-throughput Token production networks integrating "green power + super networks + dynamic scheduling."

Large model vendors are no longer content to be mere purchasers of compute power.

Independent large model vendors like DeepSeek, Zhipu AI, and Yuezhi's Dark Side are gradually evolving into builders and operators of full-stack Infra. This trend is already evident, as seen in DeepSeek's plan to build a 1GW self-owned intelligent computing base in Ulanqab and Zhipu AI's launch of a 1GW domestic intelligent computing center, along with its direct acquisition of the compilation team "CAS Jiahe" from the Chinese Academy of Sciences. These moves are fundamentally aimed at reducing reliance on external suppliers and bringing the cost of converting GPUs to Tokens under their control.

As a result, full-stack vertical integration will become the threshold for top-tier Frontier AI enterprises in the future. Only model companies capable of sourcing low-cost electricity themselves, building modular factories, and performing heterogeneous compilation optimization will survive. Small and medium-sized model vendors lacking low-cost Token production capabilities will be reduced to contract manufacturers or forced to shut down.

As model vendors move deeper into infrastructure, the value of Infra and compute software service providers like CAS Jiahe, SenseTime's Large-Scale AI Infrastructure, DataCanvas, and iSoftStone is being re-amplified. These vendors are shifting from simply forwarding API interfaces to deep cultivation compilers, Runtime, heterogeneous co-usage, and PD decoupling engines.

For example, SenseTime's Large-Scale AI Infrastructure has improved the MFU of mainstream domestic chips by 85%~152% through heterogeneous hybrid inference, doubling Token output for the same cost. iSoftStone and DataCanvas are reconstructing compute power into end-to-end "token factories," squeezing out every ounce of performance from each card through software engineering.

In the future, the value of "software-defined compute" will become prominent. Heterogeneous chip adaptability and distributed compilation capabilities may become extremely scarce assets, triggering a wave of industry mergers and acquisitions targeting underlying Infra teams within the next 1~2 years.

This cost restructuring will ultimately transmit to the very end of the industrial chain. Enterprise and developer customers will be forced to accept fine-grained billing by volume, Token, or Credits. Enterprises will become extremely focused on Token procurement costs, favoring cost-effective domestic Tokens or building their own private fine-tuning models.

Under this trend, the procurement logic of downstream enterprises and developers is also changing. "Pay-for-performance" and multi-model hybrid scheduling will become the norm. Enterprises will establish refined "Token cost control systems," using low-cost Tokens for simple rule-based tasks and reserving high-priced top-tier models only for extremely complex reconstruction and debugging tasks.

By 2026, the core of China's compute competition will shift toward ultimate exploitation of "effective output" per watt of electricity, per card, and per second of depreciation. Enterprises capable of reducing unit Token costs to be as cheap as tap water and electricity through systems engineering will truly bridge the profitability gap and secure their tickets to the Agent era. Those still trapped in low-utilization, high-cost card-stacking traps will be eliminated by the wave of industrialized Tokens.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.