08/03 2026
428
The Coexistence of Large and Small Nodes
Written by: Chen Dengxin
Edited by: Li Ji
Typeset by: Annalee
The significance of computing power has once again come to the fore.
Recently, the releases of GLM 5.2 and Kimi K3.0, coupled with the leak of DeepSeek investor meeting minutes, have underscored the pressing issue of computing power scarcity.
This indicates that for domestic large-scale models to progress further, they must surmount the computing power barrier.
Controversies surrounding computing power had temporarily ebbed, so why are they resurfacing now? How much progress has domestic computing power made? Can super nodes serve as the key to unlocking further advancements?
The Computing Power Barrier Cannot Be Overlooked
There was a period when amassing computing power was a global trend, humorously dubbed in the industry as 'brute force'.
For various reasons, this path has been fraught with challenges in China, with computing power scarcity long hindering the industry. It wasn't until DeepSeek emerged, mitigating reliance on computing power through algorithmic optimizations, that the previous brute-force approach was disrupted.
Since then, leveraging technical means to reduce computing power demands has become an industry norm: flexibly adjusting the parameter scale of the Mixture of Experts (MoE) architecture can alleviate the computational load per inference; optimizing the attention mechanism can ease memory pressure from extended contexts; utilizing open-source models to construct vertical models can curtail model training volume...

Image Source: Turing Editorial Department
After all, achieving more with less is a widely embraced strategy in China.
Against this backdrop, the narrative of perpetual computing power scarcity gradually faded in China, while the notion of eliminating computing power premiums gained momentum, naturally allaying related concerns.
Unexpectedly, domestic large-scale models have caught up and joined the international elite, leading to relative computing power shortages. Measures such as purchase restrictions and price hikes have been implemented to curb the escalating demand.
Thus, the strategy of bypassing the computing power barrier through alternative means has crumbled.
In hindsight, optimizing single-inference efficiency does not equate to satisfying the total computing power demand. Once demand spikes, pushing models to their limits, a computing power gap may emerge.
Take Kimi K3 as an example, with model parameters reaching 2.8 trillion. To enhance efficiency, only 104 billion parameters are activated per token, usually sufficient. However, after Kimi K3 gained popularity, token consumption surged, increasing the total token throughput and, consequently, the demand for computing power.

Image Source: Moonshot Official WeChat
'Bohu Finance' remarked: 'Kimi K3 excels in long-term programming and Agent tasks, where a single task involves not just one round of Q&A but the model repeatedly reading code, calling tools, and self-correcting. What users perceive as one request may entail dozens of inferences on the backend.'
This phenomenon closely aligns with the Jevons Paradox in economics.
According to Baidu Baike, in 1865, economist William Stanley Jevons posited a theory: when technological progress enhances resource utilization efficiency, it may lead to an increase, rather than a decrease, in total resource consumption.
In simpler terms, while the inference cost of large-scale models has decreased, their capabilities continue to strengthen, enabling more and complex application scenarios and attracting an increasing number of users. Once boundaries are fully opened, the user base expands, driving up demand and necessitating computing power expansion.
In fact, infrastructure construction for computing power has not waned.
Public data reveals that the domestic intelligent computing power scale reached 1590 ExaFLOPS in 2025 and surged to 2185 EFLOPS in the first half of 2026, a 177% year-on-year increase, indicating explosive growth.
Super Nodes: The Ideal Partner for Large-Scale Models
Behind the explosive growth in intelligent computing power lies the ascent of domestic AI chips to the forefront.
For instance, Kunlunxin unveiled its fourth-generation AI chip, the M100, optimized for large-scale inference scenarios (especially MOE models), positioning itself as a cost-effective alternative to NVIDIA H20 with a focus on high-performance inference applications.
J.P. Morgan data projects that Baidu Kunlun chip revenue will soar from approximately 1.3 billion yuan in 2025 to 8.3 billion yuan in 2026, a more than sixfold increase.
Another example is T-Head Semiconductor, which introduced the new-generation AI chip, Zhenwu M890, with performance three times that of its predecessor, meeting the demands of larger large-scale model training and inference.

Image Source: T-Head Semiconductor
Official data indicates that Zhenwu AI chips have shipped a cumulative 560,000 units, serving over 400 clients across more than 20 industries, with over 130,000 cards deployed in the intelligent driving sector, including clients like Xpeng, Li Auto, and NIO.
In essence, the rise of domestic AI chips has provided reassurance for domestic computing power.
Notably, super nodes have emerged from laboratories to become the next battleground for domestic AI chips, offering greater potential for domestic computing power.
A super node integrates multiple, tens, or even hundreds of AI chips into a supercomputing system with computer-like attributes, significantly boosting performance, narrowing the gap with foreign high-end chips, and substantially reducing token costs.
Huang Xiangjun, Vice President of Research and Development at MetaX, stated: 'As model scales escalate, especially for trillion-parameter large-scale models, larger-scale GPUs connected via high bandwidth are objectively required. Additionally, communication among MOE model experts demands minimal latency during both training and inference, making the super node architecture more suitable than ordinary servers.'
It should be emphasized that super nodes are not merely about amassing chips.
The core competitiveness of super nodes lies in harnessing technical advantages such as ultra-high bandwidth, ultra-low latency, and software-hardware synergy to maximize 'high efficiency, low cost' objectives, transforming nominal computing power into usable computing power.
Luo Guozhao, Director of CHIP Lab, said: 'A 2 trillion-parameter MoE model often necessitates simultaneous activation of multiple expert networks, requiring frequent data exchanges among different GPUs for each inference step. If interconnect bandwidth is insufficient, many GPUs must halt for data synchronization, rendering theoretical computing power inadequate. This is commonly referred to in the industry as 'GPUs idling rather than computing.'
Clearly, super nodes represent the optimal solution to the current computing power scarcity.
Nevertheless, as a novel concept, super nodes face technical disagreements, primarily regarding the optimal scale, with different companies offering varied solutions.
One perspective advocates for larger nodes.
As large-scale models enter the trillion-parameter era, with context lengths exceeding millions and concurrent inferences sustaining growth, the transition to hundred-card, thousand-card, or even ten-thousand-card super nodes is inevitable.
Only large nodes can meet the full-scenario demands of large-scale models.
A typical representative of this technical path is Ascend, which has sold over 750 sets of its 384-card super node, now iterated to a 1024-card super node, achieving the industry's largest 256TB unified memory addressing space across physical nodes, and planning an 8192-card super node to further explore the technical limits of full-scale super nodes.

Image Source: Tencent Technology
Another perspective favors more economical nodes.
Large-scale models require both high bandwidth and high performance, as well as low latency and strong stability, necessitating a balance. More importantly, configuring fewer AI chips reduces costs, aligning with the practical needs of small and medium-sized enterprises and avoiding computing power waste.
Thus, small nodes have also gained traction.
A typical representative of this technical path is Inspur Information, which initially launched a 64-card super node, aligning with the mainstream small node approach, and recently reduced it to 32 cards.
Zhao Shuai, Vice President of Inspur Information, said: 'Internally, we utilize trillion-parameter large-scale models for AI Coding, which is undoubtedly more efficient than using hundred-billion-parameter models. That's the value of trillion-parameter large-scale models. Therefore, we reintroduced the 32-card product to make it truly affordable for more enterprises.'
In summary, super nodes were initially overlooked but, with enhanced delivery capabilities and visible computing power scarcity, have emerged as the ideal partner for large-scale models. Companies of all sizes, including Baidu, Alibaba, Inspur Information, MetaX, and Moore Threads, have entered the fray, injecting new vitality into the AI infrastructure ecosystem.
Thus, the time is ripe for a breakthrough in domestic computing power.