Observations on the First Day of WAIC 2026: From Agent Phones to Embodied Robots, AI Explodes in a New Way

07/20 2026 539

Over the past few years, the WAIC World Artificial Intelligence Conference has not only become increasingly lively but also more representative of the global evolution of AI. At this year's WAIC 2026, it is evident that agent phones, AI glasses, robots, chips, and more are stepping more prominently into the spotlight.

Image Source: Leikeji

Both Jieyue and Nubia are developing agent phones, but while one aims to rebuild the operating system, the other wants AI to directly take over existing apps. Honor took a more straightforward approach by equipping its phones with a rotatable mechanical gimbal. As for the highly anticipated smart glasses, Leqi Rokid developed an AIOS for them and introduced the new-generation spatial computer Rokid AR.

Robots are no longer satisfied with just walking, dancing, and doing somersaults. Some can transform between humanoid and quadrupedal forms, others have entered automotive production lines, and three robots even teamed up to run a pharmacy. Compute vendors are also no longer content with just displaying a single card; instead, they brought 100,000-card clusters, supernodes, near-memory compute chips, and robot development platforms into the exhibition hall.

WAIC 2026 resembles a massive AI product and technology testing ground. According to official figures, over 1,100 enterprises showcased more than 3,000 exhibits, with over 300 products making their global debut. We have selected 20 of the most noteworthy products. While not all may become part of our daily lives in the future, they at least make one thing clear:

AI is surging into the physical world with increasing intensity.

The divergence in foundational models has been an interesting trend, not just at WAIC 2026. On one hand, models like Kimi and MiniMax continue to expand context, multimodality, and agent capabilities. On the other, Minimax compresses models to the edge, enabling phones, cars, and robots to run on-device without cloud connectivity.

These two approaches are not contradictory. Cloud models handle the upper limits of capability, while edge models address cost, latency, and privacy. Both are necessary.

Minimax's MiniCPM approach has always been clear: instead of competing with cloud models on parameter count, it focuses on increasing the "knowledge density" of each parameter in edge models. However, at WAIC this year, Minimax's main focus was no longer the model itself—the core keyword at its booth was "entering the real world."

The new MiniCPM5-1B continues this philosophy, achieving global leadership in its size category on the Artificial Analysis leaderboard and ranking first among models under 2B parameters. More importantly than the leaderboard, this capability is being integrated into real devices and gaining recognition from global leading manufacturers—Minimax has partnered with Samsung, and its on-device models will be featured in several Samsung flagship models.

Image Source: Leikeji

The value of on-device AI lies in the fact that phones, cars, and robots cannot wait for cloud responses for every action; local models must understand instructions, assess environments, and make decisions independently.

This is also the confidence behind Minimax's formal foray into embodied AI this year. At WAIC, Minimax and Legolo Robotics jointly showcased an exhibition hall guide agent solution that runs purely on-device without network connectivity, capable of real-time environment understanding, route planning, obstacle avoidance, and navigation. During park inspections, it can also analyze sensitive footage on-site, keeping data sovereignty with the client.

The same logic is being applied in more scenarios: the SuperMate intelligent cockpit product has launched with Chang'an Mazda EZ-60 and Geely Galaxy M9; the on-device AI development board "Songguo Pai" enables rapid AI hardware development for those without technical backgrounds.

From phones to cars to robots, Minimax is evolving from a "pioneer" in on-device intelligence to a "definer"—bringing intelligence out of the cloud and into the world.

The officially released Kimi K3 boasts 2.8 trillion parameters, making it the world's first open 3-trillion-parameter model, while supporting native vision and 1-million-token context.

However, not all 2.8 trillion parameters work simultaneously. K3 adopts a sparse Mixture of Experts (MoE) architecture, activating only 16 out of 896 experts at a time, and improves information transfer efficiency through Kimi Delta Attention and Attention Residuals. According to official data, its overall scaling efficiency is about 2.5 times higher than K2.

Image Source: Kimi

According to official benchmark results, Kimi K3 ranks just below the top proprietary models, GPT-5.6 Sol and Claude Fable 5.

Moonshot AI has focused Kimi K3's main strengths on long-term programming, knowledge work, and deep reasoning. In official demonstrations, K3 ran continuously for 48 hours, using open-source tools to complete the design and verification of a chip, demonstrating its long-term agent capabilities to some extent.

Additionally, K3 is now available on Kimi, Kimi Work, Kimi Code, and via API. The competition among open models is also shifting from leaderboard scores to who can truly complete complex tasks.

Robots have always been a major crowd-puller at WAIC, and this year was no exception, with a large gathering of next-generation robots, especially humanoid robots with embodied AI. While people used to enjoy watching them dance, fight, and do somersaults, this year's clearer trend is that manufacturers are striving to prove that robots are not just for videos.

Some are entering homes and outdoor environments, others are going into factories, some are responsible for dispensing medication, and others are no longer sticking to a single physical form.

The Qiyuan T1 seems like the robotics industry's rebuttal to the question, "Why must robots be humanoid?"

Indoors, it adopts a wheeled humanoid form, moving quietly through living rooms and studies, with a zero-turn radius ideal for narrow spaces. When encountering grass, gravel, slopes, or steps, it automatically switches to a quadrupedal form, using a lower center of gravity and more stable structure to navigate complex terrain.

This is not simply combining a humanoid robot and a robotic dog. Qiyuan uses a cross-morphology motion control system to uniformly manage joints, power, perception, and balance, allowing the robot to continue its original task after switching between forms.

T1's "human" and "dog" dual forms. Image Source: Qiyuan

Besides companionship and following, T1 can also connect to action cameras, executing voice-controlled filming, tracked shots, and multi-camera setups. In other words, it aims to be not just a household robot but also a mobile camera car that finds its own shooting positions.

It does sound a bit crazy. However, choosing a more suitable body based on the environment aligns more with machine logic than the idea that "robots should look human."

At WAIC this year, Zhiyuan showcased five new products, including the Expedition A3 Ultra, Elf G2 Max, Lingxi X2 EDU version, Critical Point OmniHand 3 Ultra-M, and Kuto, the world's first cycling robot.

Among them, the Expedition A3 Ultra adds an embodied processor (700TOPS), LiDAR, multi-part fisheye cameras, and a high-degree-of-freedom dexterous hand. As the world's first mass-producible, commercially deployable full-size humanoid robot, the A3 Ultra provides stronger support for deployment scenarios and was selected as a WAIC "Hall of Fame" exhibit.

Zhiyuan's positioning for the Expedition A3 Ultra is clear: it is a full-size commercial robot ready for deployment in exhibition halls, hotels, stores, and factories.

Image Source: Zhiyuan

It has a 174 cm human-compatible body, uses 360° vision fused with LiDAR for perception, and can perform autonomous navigation, reception, explanations, and guided shopping. At WAIC, Zhiyuan had the Expedition A3 Ultra make tea and play table tennis.

At the venue, Zhiyuan even brought an entire chip processing production line from Universal Robots into the WAIC exhibition hall, having robots continuously perform chip loading, finished product boxing, and full-container transportation in a super-long workflow.

"Mass producibility" is a keyword repeatedly emphasized for the A3 Ultra. Compared to a one-time successful demo, commercial robots require long-term operation, rapid deployment, on-site maintenance, and safe takeover. They don't necessarily need to perform the most daring actions but must, like a real employee, do the same job well every day.

From this perspective, the true coming-of-age ceremony for humanoid robots might be receiving an employee ID and joining a work schedule.

The robotic smart pharmacy may lack the visual impact of humanoid robots but could be one of the most realistic embodied AI products at this WAIC.

Developed through a collaboration between Ant Lingbo and Guoda Pharmacy, three robots of different configurations work together to receive orders, dispense medications, and package them, publicly stated to complete an order in 90 seconds. They are all connected to the LingBot-VLA cross-platform embodied base model, enabling collaboration on the same task.

Image Source: Ant

More importantly, this system integrates robots into online consultations and prescription workflows. Users can consult, receive prescriptions, purchase medications, and collect them without jumping between fragmented systems.

Of course, when robots enter medical scenarios, speed is not the only consideration. Prescription reviews, abnormal medications, stockouts, and dispensing failures all require strict human oversight. However, it at least demonstrates that the true value of embodied AI lies in reconnecting repetitive, standardized, and error-prone processes.

To clarify, HGR is a collaborative system connecting humans, AI glasses, and robotic dogs, rather than a single robot.

The human eye is at about 1.6 meters height, while a robotic dog's camera might be only 0.4 meters high—they do not see the world the same way. When a person wearing glasses stares at a location and says, "Go there," the system must first align the two perspectives, then convert the human's gaze point into a navigation target the robotic dog can understand.

In the demonstration, staff only needed to look at the location of a food delivery and issue the command, and the robotic dog would stand up, go to the pickup point, and deliver the food to the designated area.

While fetching food delivery may not seem complex, HGR demonstrates a more natural form of robot interaction: humans no longer need to learn remote controls and coordinate systems but can simply use gaze and language. Truly useful robots must understand what humans are looking at.

Over the past two years, AI hardware has undergone a wave of experimentation, from earbuds and pendants to various wearable devices. However, at WAIC 2026, the most promising carriers for personal AI agents remain phones and glasses: one with a mature computing, display, and app ecosystem, and the other closest to human eyes and ears.

The new products unveiled this time are also not satisfied with just adding another AI assistant. Agent phones aim to let AI understand the screen, invoke apps, and complete tasks; AI glasses are beginning to compete in operating systems, spatial awareness, and long-term memory. While these two types of hardware take different approaches, they compete for the same position: who can become the closest computing entry point to humans in the AI era.

It is easy to understand why Jieyue Stars chose phones as the first hardware for its STEPX brand. It needs portability, a screen, and sufficient on-device computing power. Combining these three requirements, the most mature carrier today remains the phone.

However, what truly makes the STEPX Neo noteworthy are the Step AOS and the personal AI agent Amoo behind it. Jieyue's vision is to have the AI agent enter the system layer, understand screen content, plan tasks, invoke tools, and complete tasks across multiple apps.

Jieyue STEPX Neo. Image Source: National Business Daily

For example, users do not need to open maps, hotel, and travel apps one by one but can simply tell Amoo their destination and preferences, with the AI agent planning the remaining steps. According to Jieyue, the STEPX Neo has passed the L3-level test of the "Artificial Intelligence Terminal Intelligence Grading" standard and is currently the only agent phone with this level of certification.

Of course, for a system-level AI agent, demonstrations are never the hardest part—the real challenge lies in building the entire ecosystem behind it, with countless issues to resolve.

Late last year, Nubia and Doubao's first-generation collaborative product, the M153 Doubao Phone Assistant technical preview, did not require each app to specifically open interfaces. Instead, it used visual understanding of the phone screen, simulating human taps and swipes to complete tasks like price comparisons, dining reservations, and order placements.

The advantage is that theoretically, it can operate any app, but the problem is equally clear: platforms may not be willing to let a third-party AI agent move freely within their interfaces.

This second-generation product unveiled at WAIC is no longer a limited engineering prototype but a flagship phone designed for mass production. It continues the system-level GUI Agent approach and adds a prominent AI button, hoping to advance from "help me check" to "do it for me."

"Second-Generation Doubao Phone" Nubia Navi X Ultra. Image source: Xiaguang Society

To some extent, what the second-generation Doubao Phone truly aims to solve is the ecosystem issue for intelligent agents. As long as WeChat, Taobao, payment platforms, and content platforms remain isolated, no matter how intelligent AI becomes, it may still be blocked at the final step.

The Honor Robot Phone is one of the most intuitive products in this batch. It incorporates a retractable three-axis mechanical gimbal on the back of the phone, allowing the camera to move actively, track people, and adjust shooting angles instead of being fixed in one direction.

This gives the phone a bit of a "bodily sense" for the first time. Placed on a table, it can follow a person during video calls; when capturing motion, it can automatically track like a small gimbal camera; when the camera faces the user and responds to external movements, the phone even produces a subtle sense of being "alive."

Honor Robot Phone. Image source: Leikeji

Honor has also introduced ARRI's Log-C encoding and LUT color grading files to the mobile side, attempting to ensure that this movable eye is not just a gimmick but also a truly functional imaging system.

However, the most interesting aspect of the Robot Phone remains its exploration of new phone forms. While foldable screens address screen size issues, the Robot Phone poses another question: Now that AI can perceive the environment, should phones also gain mobility?

Compared to giving a phone a movable eye, AI glasses go even further by directly occupying the user's first-person perspective. After cameras, microphones, speakers, and translation functions gradually converge, manufacturers are now vying for two more critical issues: who will define the operating system for glasses and when glasses should actively work.

The new generation of Rokid AR is not satisfied with being just a portable large screen for display. It supports 6DoF spatial positioning, gesture recognition, and spatial audio, while simultaneously perceiving user movements and the external environment through dual cameras.

More importantly, it is the first to feature Qualcomm's Snapdragon Supreme Spatial Computing Coprocessor. Instead of offloading all computational pressure to the phone, independent spatial computing capabilities allow the glasses to more stably understand head position, gestures, and the surrounding environment, then anchor virtual content in real space.

Rokid AR Next-Generation Spatial Computer. Image source: Leikeji

This also signals the re-convergence of AI glasses and AR glasses. Glasses that only capture but do not display are lighter but have limited interaction capabilities; glasses that can display and understand space are heavier but closer to being a true next-generation computing platform. Rokid has clearly chosen the latter.

Beyond hardware, Rokid has also introduced YodaOS—the first AI operating system for smart glasses.

Today, most AI glasses are essentially accessories for phones: captured content is transmitted back to the phone, models are processed in the cloud, and application capabilities are limited to a few pre-installed functions by manufacturers. YodaOS aims to change this relationship, enabling agents, spatial awareness, cameras, displays, and third-party services to collaborate directly within the glasses' system.

For developers, what truly matters is creating independent applications by leveraging cameras, microphones, spatial coordinates, and interaction components. Only with a system and development ecosystem can AI glasses evolve from a hardware device with limited functions into a platform that can continue to grow.

Of course, every company wants to become the Android of the glasses era. The question is, will developers come?

The deeper AI products venture into the real world, the less underlying computing power resembles an isolated chip. Robots require real-time on-device computing, large model training demands super nodes and large clusters, and scientific research also needs to connect models with experimental systems.

The RDK S600 is a development platform for embodied intelligence, powered by Digua Robot's self-developed Xuri S600 chip, delivering up to 560 TOPS of on-device inference computing power and equipped with an 18-core Arm Cortex-A78AE processor.

Image source: Leikeji

Its design logic is "integrated computing and control." In the past, robots often used one high-computing-power platform for vision and large models and another real-time controller for joint and motion management, resulting in bulkiness and complex software-hardware collaboration. The RDK S600's heterogeneous architecture aims to handle environmental perception, model inference, and real-time motion control simultaneously, integrating the robot's "brain" and "cerebellum" onto a single platform.

The development board also provides 6 MIPI camera interfaces, 6 USB 3.0 ports, and 4 PCIe 3.0 interfaces, making it easy for developers to connect cameras, lidars, and various actuators. More importantly, the Xuri S600 has entered mass production verification with over 20 leading clients.

What the robotics industry truly lacks is not just more powerful demos but also a standardized, affordable, and stably mass-produced computing base. The RDK S600 aims to be the developer motherboard for the robotics era.

The Dongfang Suanxin DF1000 is defined as a software-defined near-memory computing 3D chip. Its most unique feature is that instead of continuing to pursue more advanced manufacturing processes along traditional lines, it brings computing and storage closer through 3D integration.

Image source: Leikeji

Today, the bottleneck for AI chips lies not just in how fast they compute but also in whether data can be delivered to computing units in time. The DF1000 reduces interconnection spacing from tens of micrometers to sub-micron levels, improving interconnection density and bandwidth density, attempting to bypass the "memory wall" that plagues large model computing.

"Software-defined" means the chip can adjust its computing approach for different tasks without requiring a fixed-function chip to be redesigned for each algorithm. Of course, this claim needs validation through real tasks, but the underlying idea is intriguing: as advanced manufacturing processes become more expensive and harder to obtain, chip competition can shift from planar processes to packaging, memory, and architectural innovations.

The most intuitive figure for Huawei's Ascend 950 supernode is that a single supernode can accommodate 1,024 Ascend cards.

However, the significance of supernodes is not simply stacking cards together. During large model training, chips need to frequently exchange data, and if interconnection speeds cannot keep up, even more computing power will be wasted waiting. Supernodes aim to use high-bandwidth, low-latency interconnections to make thousands of cards appear as a single massive computer in software.

Image source: Huawei

Building on this, the Ascend 950 SuperCluster (supernode cluster) can further scale to 500,000 cards. Huawei's goal is clear: single-card performance may not be the only answer; as long as interconnection, memory, scheduling, and software stacks are sufficiently complete, systemic capabilities can also form a competitive edge.

The AI computing power war is shifting from "who has the strongest single card" to "who can make hundreds of thousands of cards truly work together."

Among a pile of phones, glasses, and robots, Tianwu Technology's "Xiaowu" may seem less flashy but could be closer to how AI truly transforms scientific research.

In simple terms, "Xiaowu" is a conversational protein design agent. Researchers can directly describe their goals in natural language, such as designing a more efficient plastic-degrading enzyme. The system then calls on protein models to generate candidate structures and connects automated experiments and data feedback to continuously screen and iterate.

Traditional protein research often involves Repeated handover (repeated handoffs) between computational design and wet lab experiments, with cycles spanning years. Xiaowu aims to string these steps into a closed loop, allowing models not just to "guess an answer" but to continue refining based on experimental results.

The biological recycling of PET waste showcased at WAIC is a concrete application: AI modifies naturally occurring plastic-degrading enzymes to improve degradation efficiency, then converts waste plastics and textiles into reusable raw materials.

This may not be as eye-catching as robots doing backflips, but it can handle part of the experimental pathway for discovering new materials and drugs for researchers.

From FaceMind AI's new-generation "Little Cannon" to Tianwu Technology's "Xiaowu," the changes at WAIC 2026 are not just about advancements in foundational models but also comprehensive transformations from agents to hardware.

Phones are the most mature "bodies," so Jieyue, Doubao, and Honor are all vying for them; glasses are closest to human eyes and ears, so Rokid and Moonix want to turn them into new perceptual entry points; robots can directly alter the physical world, so Qiyuan, Zhiyuan, and Ant Lingbo are exploring homes, factories, and pharmacies, respectively. Further down are the models, chips, development boards, and supernodes providing brains for these bodies.

But with more "bodies," the problems become more specific.

Agent-enabled phones must address app permissions, active-recording glasses must confront privacy issues, home robots must prove endurance, safety, and reliability, while supernodes must demonstrate that domestic cards can serve as the cornerstone of a truly efficient AI computing power system.

After AI enters the real world, every misoperation, disconnection, or mechanical failure will be directly exposed to users. But the direction is already clear; what remains to be seen is how many of these new species emerging from exhibition halls can truly stay in our lives.

WAIC 2026, themed "Intelligent Partners · Creating the Future Together," officially opens today!

The AI narrative has shifted from stacking model parameters to deploying agent productivity; heterogeneous collaboration and photonic computing continue to push computing limits; embodied intelligence accelerates applications, with robots entering homes and factories to make physical AI a reality.

The Leikeji WAIC exploration team has arrived in Shanghai to witness the annual peak moment of AI industrialization. Stay tuned!

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.