"Wudang" Unveiled: Arm China's Next-Gen AI VPU Redefines Video Encoding

07/24 2026 436

As AI technology advances, video is transforming from a medium for human-to-human connection into a bridge linking the physical and virtual worlds. AI has evolved from being a producer and processor of video to its largest consumer—visual perception modules in humanoid robots, intelligent driving, and other applications heavily rely on video data for AI inference and training. This transformation drives the VPU design philosophy from serving human subjective vision to also serving machine vision, thereby opening up a new dimension of video compression for machines.

 

At WAIC 2026, Dr. Huang Xin, VPU R&D Director at Arm China, delivered a keynote speech titled "AI VPU: Ushering in a New Era of Video Compression." He systematically elaborated on the evolution of the "Linglong" VPU encoding technology, unveiled the comprehensive upgrades of the new-generation "Linglong" VPU product, "Wudang," in terms of PPA, encoding specifications, and encoding quality, and shared insights into the future R&D progress and direction of AI end-to-end video encoding technology, continuing to provide robust video technology support for various AI applications.

 

In response to the new demands for video processing in the AI era, the "Linglong" VPU product is evolving from intelligent encoders to content-aware encoding (CAE) and AI end-to-end encoding technologies. According to Huang Xin, the upcoming "Linglong Wudang" VPU has achieved upgrades across three key dimensions:

Significant PPA Improvements: A single core delivers 8K 30fps+ performance at 1GHz, meeting the performance demands of diverse AI applications. Meanwhile, the encoder's PPA saves over 40% compared to the previous generation, reducing chip area and power consumption while enhancing computing power.

More Comprehensive Encoder Specifications: New support for YUV422 encoding format caters to professional videography (camera) scenarios; direct compression of Raw Data eliminates format conversion, more efficiently serving AI training; support for ultra-high bitrates up to 1Gbps+ meets the high-bandwidth demands of Physical AI scenarios like humanoid robots.

Enhanced Encoding Quality: Fully inherits the previous generation's CAE technology, which integrates lightweight AI to significantly reduce bitrates and improve video quality. Additionally, it further enhances motion feature extraction and matching, delivering an extra 5%–10% improvement in encoding quality for AI-typical scenarios involving irregular motion.

 

Looking ahead, Huang Xin stated that the "Linglong" VPU AI encoding technology is advancing from locally optimized CAE to globally optimized AI end-to-end solutions. The AI end-to-end video encoding technology is based on a new AI VPU training framework. During the training phase, a pre-posed network is combined with a virtual encoder and then integrated with specific AI applications, enabling full-link joint training through backpropagation to achieve task-specific optimization. During the deployment phase, the optimized pre-posed network is directly coupled with the Linglong VPU encoder, forming a truly task-aware encoder that significantly enhances compression rates while maximizing AI task accuracy.

 

According to Huang Xin, this framework has been validated at the engineering level. For instance, in two typical AI tasks—object tracking and action recognition—it achieved over 10% improvement in encoding quality while maintaining accuracy. Moreover, the framework demonstrates strong generalization capabilities. In the future, it can flexibly adapt to different AI scenarios by incorporating pre- and post-processing networks and multimodal configuration parameters, ultimately achieving bidirectional improvements in both human subjective vision and AI-specific tasks—ensuring clarity for humans and precision for machines.

 

From optimizing for human eyes to empowering machines, every technological leap of the "Linglong" VPU responds to the most fundamental video processing demands of the AI era. As the engine of AI vision, the "Linglong" VPU is using cutting-edge technology to answer a defining question of our time—enabling humans to see the world clearly and AI to understand it accurately.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.