09/15 2026
556
Produced by Zhineng Zhixin
The mobile chipset landscape in 2026 is shaping up to be thrilling, especially with Xiaomi's recent unveiling of several new models. The Xuanjie O3 has set a new benchmark, achieving a score exceeding 5.22 million.
Qualcomm has also upped its game, releasing a series of three blog posts ahead of the September summit. These posts fully unveiled the Oryon CPU, Adreno GPU, and Hexagon NPU of the next-generation Snapdragon platform.
Featuring a 5GHz mobile core, FlexCache shared cache, GPU-integrated Matrix Core, and NPU support for 30B MoE models, these technological advancements are tailored for on-device agentic intelligence.
Part 1: Qualcomm Unveils Its AI Strategy, Charting a New Course
Traditionally, Qualcomm has released all details of its new Snapdragon generation at the Hawaii summit in one fell swoop.
This year, however, the company has taken a different approach. Prior to the summit, from late August to early September, Qualcomm released three official blog posts, offering separate technical insights into the CPU, GPU, and NPU. Now, the complete closed loop of intelligent agents can operate locally on the phone, without relying on the cloud.
The CPU orchestrates the operations, the GPU handles rendering and AI graphics, the NPU is responsible for inference with an always-on Sensing Hub, and there's a cloud-side fallback path. Together, they form the platform foundation for on-device agentic AI.

The most striking feature of the Oryon CPU this time is its 5GHz frequency.
With two Prime cores reaching this frequency, Qualcomm believes this is the first time a mobile CPU has crossed this threshold. The company achieved this by optimizing its self-developed microarchitecture, physical implementation, and subsystem design to push the Prime cores to 5GHz.
Oryon has evolved from Kryo, and Qualcomm has invested nearly a decade in CPU self-development, reaching a point where it no longer relies on Arm's public architecture in this generation.
Note: Arm provides services to a wide range of companies, so Qualcomm needs to showcase its own innovations to stand out.

FlexCache is a cache pool shared by heterogeneous cores and allocated on demand.
When the Prime cores require a large working set, they can utilize the entire FlexCache pool without repeatedly accessing system memory for data.
In traditional designs, large and small cores have their own fixed caches. FlexCache changes this by adopting a shared and dynamic allocation approach. The cores receive as much cache as they need from the pool.

The new generation of Adreno GPU consists of three slices, each running at 1.45GHz, accompanied by a command processor and an 18MB Adreno HPM (High-Performance Memory).
HPM keeps graphical work data, such as tiles and framebuffers, local to the GPU, reducing the need to access system memory.
Power efficiency is 12% lower than that of the Snapdragon 8 Elite Gen 5. Matrix Core has been integrated into each slice, allowing matrix and AI operations to be performed directly within the GPU pipeline.
In collaboration with Neural Fusion for AI rendering and super-resolution, it has already been integrated with Unity and Unreal Engine's upscale frameworks. In the Dragon Alley demo, it achieved a 40% reduction in power consumption.

The Hexagon NPU now features an Element Accelerator specifically designed for element-level operations in Transformers, along with vector and scalar extensions for AI math and agent decision-making and routing.
It supports a maximum context length of 32K and KV-cache acceleration, with shared memory expanded by 50%, keeping model states, context, and KV-cache local to the NPU.
INT4 model prefill is 50% faster than the previous generation, with the first token time reduced to 1.5 seconds. The MoE model can handle a total of 30B parameters, with each token activating approximately 3B parameters, managed and cached through flash-to-memory expert management.

With a complete agentic loop, voice input comes through the always-on Sensing Hub, orchestrated by the Oryon CPU for planning, routing, tool invocation, reflection, and response. Inference workloads are distributed to the NPU or GPU, with results transmitted back. Any overflow is handled by the cloud, enabling the complete closed loop of agentic AI to run locally on the phone.

Part 2: Xiaomi Xuanjie O3, Xiaomi's Innovative Leap
Xiaomi recently unveiled the Xuanjie O3 at its technical communication event on August 24. Built on a 3nm process, it features 24 billion transistors and a die size of 133 square millimeters, achieving an AnTuTu score of 5.228 million, making it the first Android mobile SoC to surpass 5 million points.
This is no small accomplishment. The score of 5.22 million firmly places Xiaomi's self-developed SoC in the top tier of flagship chips.
The CPU is a deca-core all-big-core design, comprising two 4.35GHz C1-Ultra cores, four 3.68GHz C1-Premium cores, and four 3.15GHz C1-Pro cores, eliminating small cores entirely.
The NPU delivers 200 TOPS, optimized specifically for Xiaomi's MiMo on-device large model, with AI performance improved by 45% over the previous generation. The GPU is a custom G2-Ultra NX, boosting graphics performance by 85%. In GeekBench 6, its multi-core score of 15,221 is approximately 40% higher than that of the Apple A19 Pro.
Comparing Xiaomi and Qualcomm, we can see two distinct approaches.

◎ Qualcomm pursues full-stack self-development.
Its CPU is the self-designed Oryon, not based on Arm's public architecture. Its GPU is the proprietary Adreno, and its NPU is Hexagon. From microarchitecture to software stack, everything is in-house. The advantage lies in deep scheduling and collaboration, allowing for joint optimization of hardware and software. The downside is the significant investment and long iteration cycles, requiring building from the ground up.
◎ Xiaomi follows the Arm public architecture with deep customization.
Its CPU cores are based on Arm's latest C1 series architecture, with customizations in frequency, cache, and scheduling. The NPU is self-developed, and the GPU is also customized. The advantage is speed; as soon as Arm updates its architecture, Xiaomi can adapt and release a flagship chip. The downside is that the CPU core IP is not owned by Xiaomi, with Arm laying the foundation.
Arm has made its public architecture sufficiently flagship-worthy, allowing companies to create top-scoring chips without designing CPUs from scratch like Qualcomm.
Xiaomi has released two generations in two years, with the O1 to O3 benchmark scores jumping from 3 million to 5.22 million, faster than Qualcomm's normal iteration cycle. This fast-fashion approach, building on a public foundation and adding custom elements, allows for controlled release schedules.
The Xuanjie O3 does not integrate a 5G modem, relying on an external solution for connectivity. Its 3nm process is at a disadvantage compared to Qualcomm's rumored next-generation 2nm process.
Xiaomi plans to debut the Xuanjie O3 in high-end products with controllable shipment volumes, such as the 18 Fold foldable phone and tablets, adopting a cautious approach of verification before wider adoption. The digital flagship 18 series will continue to use Qualcomm chips, with the Xuanjie not expected to be widely adopted in the short term.
Qualcomm has shifted its focus away from CPU frequency, emphasizing the NPU's Element Accelerator, 30B MoE inference, 32K context, and 1.5-second TTFT for on-device agentic capabilities.

Xiaomi's Xuanjie O3 boasts an NPU performance of 200 TOPS, while Qualcomm highlights prefill speed, TTFT, and model support. These are two different evaluation dimensions. High TOPS does not necessarily mean faster large model inference, just as high paper FLOPS do not guarantee cost-effective inference. One focuses on peak performance as a selling point, while the other emphasizes actual smoothness, efficiency, and scalability. Given the limited battery and cooling capabilities of phones, impressive specs mean little if they cannot deliver real-world performance.
Qualcomm's Next-Gen Snapdragon Platform vs. Xiaomi Xuanjie O3 Specifications Comparison
◎ In the short term, Xiaomi offers attractive benchmark scores and cost-effectiveness, clear advantages on the surface.
◎ In the long term, Qualcomm's full-stack self-development accumulates advantages in software ecosystem and architectural iteration that cannot be easily matched in two or three generations.
One approach manages the foundation, materials, and finishing entirely in-house, while the other refines a base laid by others. The former's potential is reflected in its long-term accounts, not something a single benchmark score can offset.
The competition in mobile chipsets will revolve around running larger models locally, faster, and more power-efficiently.
If Arm continues to elevate its public architecture to flagship levels, Qualcomm's differentiation through self-developed CPUs will face sustained pressure.
If on-device AI applications truly take off, the scheduling advantages of full-stack self-development will regain value. Numbers like 30B MoE running on the NPU and 1.5-second first token generation will only matter when terminals and apps actually utilize these intelligent agents.
Conclusion
In my view, Qualcomm and Xiaomi's current competition shows Qualcomm demonstrating its capabilities in full-stack self-development. Xiaomi leverages Arm's public architecture to quickly build refined products, achieving strong shipments but paying royalties. The true differentiation in mobile chipsets lies in getting on-device agentic intelligence up and running.