Has End-to-End Technology, Having Evolved to Its Present State, Still Left Edge Scenarios as an Insurmountable Challenge for Autonomous Driving?

09/28 2026 477

End-to-end technology is widely acknowledged as the optimal route to achieving Level 3 autonomous driving. When this technology first emerged, many within the autonomous driving industry considered it the most promising path to attain Level 3 or even Level 4 autonomy.

However, due to its "black-box" nature, end-to-end systems can only make guesses when encountering edge scenarios. In extreme situations, such as construction zones or overturned vehicles, this limitation can have dire consequences.

As end-to-end technology has advanced, it has given rise to VLA (Vision-Language-Action) and world model technologies. The question now is: will these innovations enhance the ability of autonomous driving systems to handle edge scenarios?

01 VLA Equips Autonomous Driving with Understanding Capabilities

The VLA, or Vision-Language-Action model, integrates natural language reasoning capabilities into an end-to-end visual architecture.

The development of VLA was driven by the recognition that pure vision-based end-to-end models possess only perception capabilities and lack true understanding. To address this, a thinking module was introduced.

The VLA model comprises three core components: a visual encoder, a large language model backbone, and an action decoder. The visual encoder transforms camera footage into high-dimensional feature vectors, which the large language model processes logically. The action decoder then translates the reasoning results into physical actions, such as steering and acceleration.

By integrating these three components within the same Transformer framework, perception, reasoning, and execution are aligned within the same semantic space. This architecture enables the system to not only detect obstacles and lane markings but also to comprehend the semantic logic underlying these elements.

For instance, at a bustling urban intersection with mixed pedestrian and vehicle traffic, a traditional end-to-end model might merely detect moving objects, whereas a VLA can infer a pedestrian's intention to cross the road.

Image Source: Internet

MindVLA-o1, unveiled by Li Auto at GTC 2026, represents an effort focused on 3D spatial understanding, multimodal reasoning, and unified behavior generation. NVIDIA's 32-billion-parameter Alpamayo 2 Super also emphasizes reasoning capabilities as a core feature of its VLA model.

Nevertheless, the VLA approach is not without its flaws. Zhan Kun, head of Li Auto's base model, highlighted three key pain points in current industry VLA solutions at GTC 2026:

Insufficient alignment efficiency between 3D spatial understanding and semantic reasoning, excessive decision-making delays due to long visual-language-action transmission chains, and difficulty in covering long-tail scenarios solely through real-world data expansion. Li Chuanhai, CTO of Geely Automobile Group, also identified three limitations of VLA:

It can only match standard answers and lacks systematic cognitive abilities, relies on limited driving operation data rather than vast internet video resources, and struggles to model the operational rules of the physical world. Additionally, weak 3D spatial perception is another issue that VLA needs to overcome.

Many early solutions directly adopted 2D visual-language models, inherently limiting spatial perception. Even when 3D spatial representations are introduced to enhance spatial understanding, balancing them with the original reasoning capabilities of language models becomes a challenge. Reasoning delays are also problematic, as autoregressive reasoning processes account for most of the delay, and millisecond-level delays in high-speed scenarios can significantly impact driving behavior.

Although architectural innovations from 2025 to 2026 have mitigated this issue, delay remains one of the core obstacles to VLA deployment.

Moreover, the high cost of covering long-tail scenarios with data, which is difficult to achieve through real-world testing alone, and the massive parameter count of VLA models, which puts pressure on in-vehicle computational power and system costs, pose significant challenges.

Some academic studies also suggest that these models are highly sensitive to input perturbations, and their reasoning credibility remains to be thoroughly verified.

These issues indicate that while VLA has opened up new directions, many problems still need to be resolved before it can be reliably deployed on a large scale.

02 World Models Enable Systems to Learn Anticipation

World models take a different approach by bypassing the linguistic intermediate layer and directly modeling and predicting in 3D space.

They enable the system to proactively reason by constructing a dynamic, physically plausible virtual traffic scenario internally. Based on physical causality, such as inertia, friction, and motion trajectories, they predict the future motion paths of surrounding vehicles, pedestrians, and obstacles, allowing for optimal path planning seconds in advance.

This capability is particularly crucial for edge scenarios.

For example, if a truck suddenly overturns ahead, a system equipped with a world model can begin reasoning about how the obstacle will move and what the consequences of different braking intensities will be while simultaneously identifying it, leading to more reasonable decisions.

In contrast, VLA solutions in such scenarios rely more on behavioral patterns learned from human driving data and lack direct modeling and reasoning capabilities for physical processes.

This is the fundamental difference in reasoning paths between the two when dealing with unknown scenarios.

Image Source: Internet

Xpeng showcased the X-World world model at CVPR 2026. Its inputs include historical multi-view videos and future vehicle actions, and its output is the potential future scenes the vehicle might encounter.

This allows the autonomous driving system to predict how the surrounding world will change if the vehicle executes a certain action.

Waymo's Waymo World Model, released in early 2026 and based on Google DeepMind's Genie 3 architecture, can generate simulation environments containing rare scenarios such as tornadoes, road flooding, and even elephants entering the road. This generative capability allows the system to encounter extreme scenarios in a virtual world that are almost impossible to encounter in real-world testing.

However, world models are not all-powerful. Their strength lies in bottom-up physical reasoning—judging how objects will move—but they lack high-level semantic reasoning and understanding of social rules, meaning they cannot determine how to react appropriately.

They know that a vehicle ahead might slow down but are uncertain about the correct traffic rules to follow in such a situation.

03 The Two Paths Are Converging

In the second half of 2026, the integrated approach of VLA and world models has surpassed the competing approach between the two.

At CVPR 2026, Xpeng explicitly proposed a dual-pillar architecture, where its second-generation VLA and world model jointly form the two pillars of the physical world base model.

The logic of VLA is to learn from humans—acquiring decision-making habits in complex road conditions from driving videos and instructions—while the logic of world models is to learn from the world by making frame-by-frame predictions on massive unlabeled videos to gradually acquire the dynamics and causal structure of the physical world.

The former provides sparse but high-density behavioral supervision, while the latter provides dense physical prediction signals, making them highly complementary. Li Auto has taken a different integration path.

In 2025, Li Auto unified spatial understanding, language understanding, and action decision-making into a single model framework, constructing the VLA Driver Large Model based on three technology stacks: VLA, world models, and reinforcement learning.

MindVLA-o1, released at GTC 2026, was further upgraded by introducing a predictive latent world model on top of the language model's semantic understanding capabilities, efficiently simulating future scene changes in latent space.

Li Auto defines this architecture as a general-purpose agent for the physical world, where the same VLA model can control both vehicles and robots.

Image Source: Internet

From a technical perspective, this integration indeed enhances the system's ability to handle edge scenarios.

Traditional end-to-end models can only make guesses when faced with unseen scenarios, while the combination of VLA and world models provides the system with two independent sources of information: one from semantic understanding of human driving data and the other from causal reasoning based on physical laws.

When both paths provide consistent conclusions, decision confidence increases; when they diverge, the system has more redundancy and verification space.

However, the integrated path still faces many issues. Real-time responsiveness, physical realism, and lightweight deployment of world models are still being addressed; the hallucination problem of large models has not been fully resolved.

Additionally, both VLA and world models require massive computational power and data support, which is a significant barrier.

04 Conclusion

So, will the maturation of VLA and world model technologies make autonomous vehicles more stable in handling edge scenarios? From the perspective of technological evolution, the answer is yes.

VLA addresses the shortcoming of end-to-end models' inability to understand, while world models address their inability to anticipate. Their integration provides the system with more reliable reasoning paths when facing unknown scenarios, rather than relying solely on statistical patterns in training data.

However, Zhijia Zuiqianyan believes that higher stability does not mean sufficiently high stability. The essence of edge scenarios is their rarity and inexhaustibility. No model, regardless of the size of its training data or the number of parameters, can have encountered all possible extreme situations.

The role of VLA and world models is to enable the system to make relatively reasonable judgments based on a deeper understanding of semantics and physics when faced with unseen scenarios, rather than relying entirely on pattern matching to guess a result.

In this sense, the value of these two technologies lies not in eliminating edge scenarios but in giving the system more thinking capabilities and less blind guessing when facing the unknown. This transformation visibly enhances safety, but there is still a long way to go before true stability and reliability are achieved.

#AutonomousDriving #EndToEnd #EdgeScenarios

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.