10/08 2026
439
With the development of end-to-end technology, VLA models have found their place in autonomous driving.
The VLA model, or Vision-Language-Action model, often sparks curiosity about the role of language within it. What exactly does the 'L' in VLA do?
Is it similar to the voice assistants we are familiar with in the cockpit?
In fact, in VLA models, the role of language far exceeds mere human dialogue. It is not a voice assistant serving passengers but rather the cognitive hub of the autonomous driving system itself.
Between perception and action, language provides crucial understanding and reasoning capabilities.
The integration of this capability has enabled autonomous driving to evolve from statistical conditioned reflexes to cognitive-based, informed decision-making.
01 From Seeing Pixels to Understanding Scenes
Traditional end-to-end autonomous driving models can accurately perform object detection but struggle to grasp the deeper meanings behind scenes. The language module in VLA fundamentally changes this.
Through the deep integration of visual and language encoders, the model can align the spatial information perceived visually with the semantic knowledge embedded in language.
For example, the system can not only recognize a traffic cone ahead but also infer, based on construction signs and lane markings, that it is entering a construction zone and should slow down and prepare to detour.
The key to achieving this understanding lies in the fusion of multimodal information.
Language encoders (such as GPT-like models) can convert natural language instructions or scene descriptions into high-dimensional semantic vectors, which interact with the spatial features extracted by the visual encoder in a unified representation space.
This fusion endows the VLA model with the ability to handle long-tail scenarios. Scenes like pedestrians suddenly emerging from between parked cars in narrow residential areas, judging the speed and yielding relationships of oncoming traffic at unprotected left turns, and temporary detours in construction zones are difficult to address with direct mapping from visual input to driving actions alone.
The internet-scale common-sense knowledge and causal reasoning capabilities embedded in the language module precisely fill this gap.
02 How Language Drives Planning
The most profound impact of language on driving is reflected in the decision-making and planning processes.
The language module in VLA can perform chain-of-thought reasoning, engaging in a series of structured thoughts like a human driver.
The AutoVLA model, accepted by NeurIPS 2025, combines chain-of-thought reasoning with tokenization of physical actions to directly generate planning trajectories through a unified autoregressive generation process.
The model features two modes: fast thinking (outputting only trajectories) and slow thinking (incorporating chain-of-thought reasoning).
In complex scenarios (e.g., construction zones, ambiguous intersections), the model activates slow thinking mode, generating internal reasoning chains like 'The traffic light ahead is about to turn yellow, and there are pedestrians waiting on the left, so I should slow down and prepare to stop' before generating driving trajectories.

Image source: Internet
In simple scenarios, fast thinking mode is employed to enhance efficiency. The driving intentions inferred by the language module in VLA can also directly participate in and guide trajectory generation.
The MindVLA-U1 model, jointly proposed by the MMLab of The Chinese University of Hong Kong, Li Auto, and Tsinghua University, utilizes the Classifier-Free Guidance (CFG) mechanism to enable the driving intentions predicted by the language side to directly participate in continuous trajectory generation.
Experimental data shows that the model achieved an RFS (Route Completion Score) of 8.20 on the validation set of the WOD-E2E autonomous driving benchmark, while the RFS for human driving reference trajectories was 8.13.
In other words, the quality of trajectories generated by the model in open-loop evaluation surpassed human driving reference standards for the first time.
03 A Bridge to Interpretability and Human-Machine Trust
The black-box nature of autonomous driving has long hindered public trust and regulatory implementation. The introduction of language modules provides a feasible technical pathway to address this issue. VLA models can output their internal reasoning processes in natural language, generating the basis for decisions.
NVIDIA unveiled the open-source VLA reasoning model NVIDIA DRIVE Alpamayo-R1 (AR1) at the NeurIPS conference in December 2025, describing it as the world's first industrial-grade open-source reasoning Vision-Language-Action autonomous driving model. This model innovatively integrates chain-of-thought AI reasoning with path planning technology.
In areas densely populated with pedestrians and adjacent to bicycle lanes, vehicles equipped with AR1 can reason through chain-of-thought processes, collect driving path data, integrate reasoning trajectories (i.e., the system's explanations for taking specific actions), and then plan subsequent routes.

Image source: Internet
Built on NVIDIA Cosmos Reason, AR1 represents an interpretable AI driver, which holds significant importance for safety verification and regulatory review.
The VLA large model equipped in the Wey Blue Mountain Intelligent Advanced Edition also provides a CoT reasoning card function, which can present the reasoning process of assisted driving to users in real-time, making the reasons for each brake or detour clearly visible.
04 Final Thoughts
The integration of language essentially answers the most fundamental question in driving: whether the vehicle truly understands its environment and the situations it must navigate.
The push for VLA models toward mass production indicates industry recognition of this pathway.
The discussion now shifts from whether VLA can be used to how to use it more intelligently—that is, how to make more critical inferences with limited computational resources, rendering decisions both interpretable and trustworthy, and transitioning from mere usability to true practicality. #AutonomousDriving #VLA