WAIC 2026 Roundtable Discussion - From Virtual to Reality: How World Models Drive Embodied Intelligence

07/20 2026 470

At the main forum of this year's WAIC 2026 World Artificial Intelligence Conference and High-Level Conference on Global AI Governance, the buzzword "world models" in autonomous driving and embodied intelligence was highlighted in a roundtable discussion titled "From Virtual to Reality: How World Models Drive Embodied Intelligence."

The roundtable was moderated by Jiang Yugang, Vice President of Fudan University, with dialogue guests including Yao Maoqing from Zhiyuan Robotics, Chen Yilun, Founder and CEO of Tashi Zhihang, and Yao Song, Founder of Liangyuan Xinchuang in the physical AI industry.

The unconventional aspect of this roundtable discussion was that the participants did not simply view world models as "better video-generating models" in the digital world, but rather discussed them within the context of Physical AI's embodied intelligence and autonomous driving industries.

At the same time, their respective companies are all actively practicing and exploring the application and implementation of Physical AI, making their insights worthy of exploration and learning. This article summarizes and shares their views and experiences on world models from the roundtable discussion, hoping to provide some information and inspiration.

Yao Maoqing: World Models Predict States, Not Images The first response in the roundtable discussion came from Yao Maoqing of Zhiyuan Robotics. He defined world models quite directly: in the field of Physical AI, the essence of a world model is an AI system capable of understanding the operational laws of the physical world. Its most crucial function is to predict the next state of the world, rather than simply selecting the next frame or generating a seemingly plausible video.

In his view, a true world model must possess several capabilities:

Understanding multimodal information and integrating inputs from different dimensions of the real world;

Grasping physical laws, including dynamics and spatial relationships;

Possessing causal reasoning abilities to understand "why objects are pushed or fall";

Being capable of long-term, extended reasoning without collapsing during continuous deductions.

These points highlight the challenges in exploring world models within the industry.

They also define the boundaries of world models: video generation is merely a surface-level capability, while physical state prediction is the underlying ability. Yao Maoqing's key judgment can be summarized in one sentence: World models predict the next state of the world, not the next image.

Regarding implementation, he believes that within the next three years, embodied intelligence is more likely to first land in scenarios that are "high-frequency, rigidly demanded, environmentally controllable, and highly certain." Zhiyuan has already conducted multi-day live validation in production line scenarios, with Yao Maoqing mentioning "six days, over ten hours per day, over 60,000 production line operations, with an overall success rate of 99.99%." His judgment is pragmatic: open environments like homes require stronger general-purpose generalization capabilities, while deterministic tasks in industrial production lines will become earlier testbeds for embodied intelligence.

Chen Yilun: World Models Must Answer "What Happens to the World After an Action"

Chen Yilun from Tashi Zhihang framed world models more in terms of robot control. He argued that while VLA can tell robots "what action to take next," world models must further answer: "If this action is taken, how will the world state change? Is this change desirable?"

Of course, VLA also involves training and reasoning, as seen in our previous sharing of NVIDIA's article on Alpamayo, "NVIDIA Alpamayo: A Comprehensive Analysis of the Design and Mass Production Deployment of a Reasoning-Based Autonomous Driving Large Model."

However, he broke down the world model problem into State, Action, and Reward: state, action, and feedback. World models do not isolate state prediction or action output but predict the joint relationship between Action and State. This allows them to combine supervised and reinforcement learning, enhancing task completion, reliability, and success rates for robots.

The most powerful statement in his speech was: Neural networks should not just provide an Action; they should provide what lies behind the Action.

Chen Yilun also pointed out that the current biggest bottleneck is data. Public videos primarily provide geometric and visual information, which may help non-contact systems but fall short for robotic operations involving contact systems. Simply watching videos makes it difficult to obtain key physical quantities like force, touch, flexibility, and friction. In other words, videos can tell models "what appears to happen" but not necessarily "why it happens physically."

Therefore, he believes training truly usable world physical models requires three conditions:

Comprehensive modalities, including not just video but also force and touch;

Data from high-frequency interactions, where robots continuously change states through actions;

Data from real-world scenarios and tasks.

His quantitative judgment is also clear: if autonomous driving requires approximately one million hours of data to reach industrial-grade usability, then more complex embodied operations may require tens of millions of hours.

Jiang Xu: World Models Alone Are Not Enough; Embodied Intelligence Needs Three Paradigm Shifts

Jiang Xu's perspective resembles viewing world models through the history of large model industry development. He believes the core tasks of world models can be divided into two parts: predicting the next state (or the next frame for video models) and predicting the next action. If a choice must be made, he prioritizes "predicting the next action" because the ultimate goal of training world models is to control robots.

His judgment comes from an analogy: the development of large models in recent years has essentially involved large-scale imitation of humans, ultimately surpassing human capabilities in certain areas. Embodied intelligence may follow a similar path. However, when humans act in the physical world, they do not constantly and explicitly predict the next state; instead, they intuitively act based on thorough observation. Therefore, he proposes a clear division of labor: states should be perceived, and actions should be predicted.

Jiang Xu also cautions that while world models have suddenly gained attention, "world models alone cannot achieve embodied intelligence." He attributes the explosion of large models to three paradigm shifts: scalable pre-training, scalable alignment, and scalable commercialization and deployment. This holds true for language models, code models, video models, and will likely apply to embodied intelligence as well.

In terms of commercial implementation, he continued (continues with) the thought process (logic) of large model product history: AI capabilities often lack precision and reliability in their early stages, so they first emerge in high-tolerance scenarios. ChatGPT had hallucinations early on but was tolerable as an assistant; code models initially only provided completion suggestions, chosen by programmers; video models could initially be used through "trial and error." Embodied intelligence also needs to find similar high-tolerance scenarios rather than immediately entering zero-error, high-responsibility tasks.

Consensus: The Key to World Models Is Not "Generation" but "Understanding Before Action"

From the perspectives of the three guests, several consensuses on world models emerged at WAIC.

World models must align with physical laws. Models that only generate images or videos do not equate to the world models needed for embodied intelligence. They must understand dynamics, spatial relationships, causality, and long-term evolution.

World models must be tied to actions. The truly valuable aspect is not "what the world will become" but "how the world will change after I take a certain action." This is also what distinguishes them from ordinary video generation models.

The biggest bottleneck for world models is data, especially real physical interaction data. Public video data is valuable but far from sufficient. Robots require a data loop composed of state, action, force, touch, and result feedback.

World models are just one part of embodied intelligence. For robots to truly enter industries, a flywheel effect must be created involving the robot body, data, models, scenarios, and commercial deployment.

This aligns with our previous article, "World Models 101: Concepts, Technical Debates, and Scientific Applications." Thus, it is evident that world models are rapidly being explored, applied, and iterated in the industry, serving as a crucial algorithmic core for current Physical AI.

Final Summary: World Models Are Moving from "Imagining the World" to "Intervening in the World"

If discussions of world models in the digital world previously focused on "whether models can generate a realistic world," then this WAIC roundtable advanced the question to the real physical world: Can world models enable AI to act in the real world?

Yao Maoqing emphasized state prediction and physical laws, Chen Yilun emphasized joint Action-State modeling and contact data, and Jiang Xu emphasized action prediction, scalable training, and commercial implementation scenarios. While their expressions differed, they all pointed to the same judgment: The endpoint of world models is not generating the world but helping robots understand, predict, and complete tasks in the world.

This also reflects the industry's direction. World models are transforming from "imagination" in the virtual world to "actionability" in the physical world. They are no longer just a pair of eyes for AI to see the world but an internal physical theater in AI's mind where it rehearses before entering the world. Whoever can connect this theater with real data, real robots, and real scenarios will be closer to the true ChatGPT moment for embodied intelligence. With the prominence of "world models" at this WAIC, it is undeniable that the true ChatGPT moment for embodied intelligence is not far off.

*Unauthorized reproduction or excerpting is strictly prohibited.-

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.