NIO, Huawei, and Li Auto are All Talking About the 'World Model'. Where Are the Differences?

08/20 2026 523

Produced by Zhineng Technology

The term 'World Model' is rapidly becoming a high-frequency buzzword in the field of intelligent driving.

NIO has NWM, Huawei incorporates the World Engine into its WEWA architecture, and Li Auto has built a cloud-based World Model behind its VLA Driver. Other players, such as XPeng, Horizon Robotics, and overseas companies like Waymo and Wayve, have also integrated it into their respective systems.

While these companies use similar names, their technical focuses differ significantly. Judging solely by the names announced at product launches, it may seem like they are all following the same path, but a closer look reveals substantial differences.

VLA, World Models, and Reinforcement Learning: Not an Either-Or Choice

VLA, World Models, and Reinforcement Learning, often discussed together in the industry, operate at different technical levels.

◎ VLA addresses input-output issues: How vehicles convert visual, linguistic, and other information into driving actions.

◎ World Models tackle environmental modeling: Whether the system can understand objects, relationships, and changes in the road environment and predict the consequences of different actions.

◎ Reinforcement Learning is a training method where models act in simulated or closed-loop environments and adjust their strategies based on evaluations and rewards.

These three can be combined rather than being mutually exclusive. When comparing different companies, the focus should not be on which term they 'choose' but rather on where the World Model is positioned, what tasks it undertakes, and how it contributes to the formation of final driving capabilities. Viewed through this lens, NIO, Huawei, and Li Auto represent three distinct system combinations.

NIO: Letting the World Model Participate in Capability Iteration

In late May 2025, NIO began gradually rolling out the first version of its NWM. By January 28, 2026, a new version was pushed to over 460,000 vehicles equipped with the Banyan system.

Within NIO's ecosystem, NWM is not just for generating driving videos or reconstructing scenes; it is placed at the core of intelligent driving capability iteration and incorporates closed-loop reinforcement learning into the training system.

The key here is 'closed-loop.'

Traditional simulation is like giving the system pre-prepared exam questions to see if it can complete them. In closed-loop training, the model first takes action, and the simulation environment changes accordingly. The evaluation system provides feedback, and the training system adjusts the strategy.

In a pedestrian crossing scenario, whether the model chooses to accelerate, decelerate, or detour will result in different vehicle trajectories and collision risks. The World Model must respond to these changes; otherwise, it is merely a scene-generating tool and not a complete training environment.

The focus of NIO's approach is to have NWM participate in both environmental understanding and capability iteration simultaneously. However, a complete simulation closed-loop does not necessarily equate to stronger real-world road capabilities. It also depends on whether the scenarios are realistic, whether the evaluations are reliable, and whether virtual training can transfer to reality.

Huawei: Building the World in the Cloud, Handling Actions in the Vehicle

Huawei has opted for a more explicit system split (explicit system division).

◎ In the ADS 4's WEWA architecture, WE refers to the World Engine running in the cloud.

◎ On the vehicle side, the World Action Model generates driving behaviors based on real-time environments.

The cloud-based World Engine primarily reorganizes real-world data and generates complex training scenarios. For example, a segment of ordinary road data can be modified in the cloud to change vehicle positions, participant behaviors, and scenario trigger timings, generating scenarios such as cut-ins, sudden brakes, and pedestrian crossings. The density of effective scenarios is key, with virtual mileage being a secondary consideration.

In April 2026, Huawei unveiled ADS 5, upgrading the architecture to WEWA 2.0 and introducing multi-agent gameplay and online reinforcement learning.

Multi-agent gameplay ensures that other vehicles in the simulation no longer follow fixed scripts: when the main vehicle changes lanes, the side and rear vehicles may yield or accelerate; if the main vehicle hesitates, surrounding participants will also alter their actions.

According to Huawei's announcements, multi-agent gameplay increases training intensity tenfold, and online reinforcement learning boosts training efficiency tenfold. These figures reflect Huawei's internal training system and cannot be directly compared with other companies' computational power or simulation mileage.

Huawei's approach is characterized by clearly defined modules for the cloud-based world, vehicle-side actions, and training feedback, all connected through the WEWA architecture.

Li Auto: Vehicle-Side VLA-Centric, Cloud-Based World Model for Training

Li Auto is more easily categorized under the VLA route.

VLA Driver integrates visual, linguistic, and driving behaviors into a single model, not only recognizing vehicles and pedestrians but also understanding lane rules, traffic signs, and the driver's linguistic commands.

According to public information, VLA Driver is supported by a self-developed reconstruction-based and generative unified cloud-based World Model that enables large-scale closed-loop reinforcement learning. In September 2025, VLA Driver was fully rolled out to AD Max models.

This system is divided into two segments: the vehicle-side VLA Driver understands the environment and generates actions, while the cloud-based World Model reconstructs and generates scenarios, allowing the vehicle-side model to undergo training and evaluation. VLA is closer to the vehicle-end model that users interact with daily, while the World Model remains hidden behind the R&D and training system.

Blurring Boundaries Across the Industry

Beyond NIO, Huawei, and Li Auto, the routes taken by other players further illustrate that VLA, World Models, and Reinforcement Learning are converging, though each company integrates them differently.

◎ XPeng: VLA Defined as a World Model

XPeng unveiled its second-generation VLA in November 2025, adopting a 'vision, implicit representation, action' approach that eliminates explicit linguistic translation and directly generates actions from visual signals.

Its self-definition is intriguing: it is both an action generation model and a World Model capable of understanding and predicting the physical world.

On the training side, XPeng disclosed a 72 billion-parameter cloud-based foundation model, a 30,000-card cluster, and nearly 100 million clips. In April 2026, it released the X-World technical report, applying generative World Models to VLA 2.0's closed-loop simulation, online reinforcement learning, and model evaluation.

According to official data, the number of simulated scenarios grew from 30,000 a year ago to over 500,000, with daily simulated mileage equivalent to 30 million kilometers of real-world driving.

XPeng's approach demonstrates that as VLA begins to acquire the ability to predict the future and understand physical laws, the boundary between it and World Models becomes increasingly blurred.

◎ Xiaomi: Shifting from Data Imitation to Cross-Domain Cognition

Xiaomi started late but has made significant iterative strides. The end-to-end assisted driving system in the YU7 used 10 million clips, primarily relying on professional driving data to learn behaviors. In 2026, the new-generation SU7 introduced the XLA architecture, based on the MiMo-Embodied foundation model, aiming to unify the cognitive capabilities of assisted driving and robotics while enhancing spatial understanding, physical world comprehension, and risk prediction.

From public information, Xiaomi is transitioning from 'large-scale driving data imitation' to 'a shared foundation model for driving and embodied intelligence.' However, details on XLA's closed-loop methods, World Model structure, and mass-production effects remain incomplete.

◎ Horizon Robotics: Positioning the World Model as a Supplier Capability

Horizon Robotics faces a different challenge: how to adapt a single technology to multiple brands, models, and hardware configurations. HSD employs a one-stage end-to-end approach while incorporating World Models and Reinforcement Learning. The former understands and predicts scenario changes, while the latter adjusts behaviors through environmental interactions.

As a solution provider, Horizon must also consider computational costs, model adaptation, engineering delivery, and scalable mass production. This makes Horizon's World Model more akin to a platform capability: it serves not just one vehicle but is reusable across different automakers' data and product boundaries.

◎ BYD: Prioritizing Fleet and Data Scale

BYD's strength lies first in vehicle scale. As of May 28, 2026, its assisted driving models exceeded 3.15 million units in circulation, with the 'Divine Eye' system generating over 200 million kilometers of data daily. BYD plans to upgrade its Physical AI Large Model and integrate it with the Xuanji Architecture 2.0.

It is important to note that '200 million kilometers daily' does not equate to 200 million kilometers of effective closed-loop training completed each day, nor can it be directly compared with other companies' simulation mileage.

BYD offers an alternative path: first expanding real-world scenario access through a massive production fleet, then using Physical AI to enhance data utilization and generation efficiency.

For World Models, model structure determines what it can learn, while vehicle scale determines how many real-world scenarios it can access. BYD's focus leans more toward the latter.

Overseas Players Place Greater Emphasis on Training and Validation

Among overseas companies, Waymo and Wayve more explicitly utilize World Models for simulation, evaluation, and safety validation.

◎ Waymo's Foundation Model serves three components: Driver (responsible for driving), Simulator (generating closed-loop environments), and Critic (evaluating results). The internal training loop then uses reinforcement learning to refine strategies.

◎ Wayve's GAIA series is closer to a generative simulation platform. GAIA-2 generates controllable multi-perspective driving scenarios, while GAIA-3, released in December 2025, further focuses on system evaluation and validation.

Both companies emphasize that the value of World Models lies not in generating 'realistic-looking' videos but in their ability to stably control scenarios, reproduce dangerous conditions, and provide comparable test results.

Tesla's public stance differs. It has not consistently branded 'World Models' as a core component of FSD, instead emphasizing real-world fleet data, unified models, training computational power, and reinforcement learning.

In 2026, Tesla's disclosed FSD v14.3 upgraded the reinforcement learning training phase to handle more long-tail scenarios. This indicates that even without prominently featuring the term 'World Model,' Tesla is using real-world fleet data and reinforcement learning to address long-tail problems.

How Far Has a Company's World Model Progressed?

The term 'World Model' has become a broad industry term, and its progress can be evaluated from at least four perspectives:

◎ Can it predict the consequences of actions? Generating the next frame based solely on historical footage is insufficient; the system must understand how other participants will react to the vehicle's acceleration, deceleration, or turning.

◎ Can it sustain closed-loop operation? Generating a few seconds of realistic video does not mean the driving model can continuously operate in the environment, receive evaluations, and refine its strategies.

◎ Can scenarios be controlled and reproduced? Models used for training must be able to fix road conditions, weather, vehicle behaviors, and dangerous scenarios. Only with controllable scenarios and repeatable results can meaningful comparisons be made across versions.

◎ Can simulation results transfer to real-world roads? This is the most challenging step: realistic visuals do not equate to accurate physical relationships, and high virtual scores do not guarantee real-world safety. Ultimately, the proof lies in the mass-produced vehicles' disengagement rates, misjudgment rates, and collision data.

Conclusion

When examining these various approaches, a clearer trend emerges: World Models are evolving from environmental prediction tools into training infrastructure for intelligent driving.

Automakers are concerned with whether it can improve the mass-produced vehicle experience, solution providers with whether it can be reused across models, and Robotaxi companies with whether it can provide repeatable safety validation. Different business models lead to different system constraints, explaining why everyone talks about World Models but makes public metrics difficult to compare directly.

What can be confirmed at this stage is that vehicle-end action models, cloud-based generated environments, evaluation systems, and reinforcement learning are being connected into closed loops. What remains uncertain is which closed-loop approach yields the greatest real-world benefits.

The answer lies not in simulation mileage itself but in whether long-term road data can demonstrate reduced disengagements and misjudgments while lowering collision risks.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.