Detailed Explanation of the Architecture and Training Methods of NVIDIA's VLA Autonomous Driving Model Alpamayo

09/30 2026 450

This year, the second sub-forum of NVIDIA's Alpamayo Summit focused on the core of the open ecosystem—the reasoning model itself. The speaker, Yurong You, is a Senior Research Scientist in NVIDIA's Autonomous Driving Research Group and one of the main authors of the Alpamayo model. In this talk, he shared NVIDIA's Alpamayo model and training recipe.

He also revealed several topics of particular concern for autonomous driving algorithms:

A 4.7% improvement in closed-loop simulation scores after a separate round of RL on the diffusion action expert

A 2 to 4 times acceleration in reasoning after moving from text to latent space

A reduction in the number of tokens used for the same decision, from 90 to 38, when comparing text-based reasoning to latent space reasoning

Only 24 GB of VRAM is required to run Alpamayo 1.5 reasoning

The Alpamayo 2 Super, to be released this summer, will have 32 billion parameters

More importantly, during training, NVIDIA found that the model now reflects more frequently in increasingly difficult scenarios. In a CVPR paper presented by You, the model first proposed 'accelerating through the roundabout' but then paused—'Wait, if that pedestrian continues walking, accelerating now will cause a collision'—and revised its decision to decelerate and wait. The key takeaway: This generation of Alpamayo is not about 'making the car talk' but about 'enabling the car to correct itself during the reasoning process'—reasoning has transformed from an explainability Decoration (decoration) into a mechanism for improving action quality. Of course, achieving this does not rely on a larger model but on a carefully designed, downloadable RL post-training recipe.

Therefore, this article follows the chapters of the original speech, first clarifying the information from the PPT and the speaker's original words in each chapter, and then provides insights from the perspective of industry learning and commentary. It shares the architecture and training methods of NVIDIA's latest model, hoping to provide information and inspiration for the deployment of autonomous driving and humanoid robots.

Title Slide: Yurong You, Senior Research Scientist, NVIDIA Autonomous Driving Research Group, Alpamayo Summit, June 4, 2026

I. First, the Architecture: Reasoning and Action Layers Are Separate Agenda: Overview of Alpamayo Model and Recipe; Alpamayo 1.5 Demonstration; SFT/RL Recipe

You opened with a single sentence: Driving a car requires more than just perception and trajectory prediction; we need a paradigm that can 'scale and generalize to solve long-tail problems.' He then reviewed the architectural diagram introduced by Marco Pavone at the opening.

Three-layer architecture.

Action Layer: Reasoning-guided trajectory diffusion, lightweight conditional flow matching decoding, and RL post-training to align reasoning with actions.

Reasoning Layer: Internet-pretrained Cosmos-Reason backbone, improved causal reasoning using RL with verifiable rewards.

Vision Layer: Multi-camera, multi-timestep tokenization. Inputs: Multi-camera images, user instructions, navigation, and vehicle history. Training signals: IL, SFT, RL

Alpamayo is a VLA that connects reasoning and action prediction, designed specifically for autonomous driving. It is not a standard VLA—which directly maps vision and text to actions; instead, it separates the task by first reasoning about the scene and then outputting a trajectory. Inputs are multi-camera, multi-timestep images, plus text-based user instructions and navigation; outputs include three things: reasoning trajectory, meta-action, and trajectory. You said this approach is both scalable and generalizable, significantly improving action prediction and adding something new: 'We have a way to introspect what the model is thinking, which provides a path for safety monitoring.'

Vehicle Commentary

First, the most important information in this architectural diagram is the parallel (parallel) presence of three training signals: IL, SFT, and RL.

These are all ways to provide 'learning signals' to the model during autonomous driving (and large model) training, differing in where the signals come from and how the model is corrected.

IL — Imitation Learning: The model observes human experts' driving records (sensor inputs → expert trajectories/actions) and learns to output similar actions. The signal is 'what the expert did at this moment,' and the model copies it. The most common form in autonomous driving is Behavior Cloning, which is essentially the pretraining for Alpamayo. Advantages: Data is readily available, and training is stable. Disadvantages: As discussed in Igl's talk, expert data lacks errors, so the model cannot learn to recover from its own deviations (covariate shift); it also cannot distinguish whether it stopped because of a red light or the car ahead (causal confusion).

SFT — Supervised Fine-Tuning: A pretrained model is retrained with a batch of carefully annotated 'input → standard answer' data, where the signal is provided by humans or high-quality models. Mathematically, it is the same as IL (both are supervised learning), but the difference lies in usage: IL typically refers to learning actions from scratch, while SFT refers to targeted corrections to an existing large model. In Alpamayo, SFT data consists of human-written reasoning chains ('There is a pedestrian crossing ahead, so decelerate') paired with trajectories, teaching the model to provide reasons before taking actions. The 'closed-loop supervised training' (RoAD) mentioned in the reasoning model talk also falls into this category—when the model deviates in simulation, an expert model provides corrected trajectories, and these pairs are used as SFT data.

RL — Reinforcement Learning: No standard answers are provided; only a score (reward) is given. The model generates actions, and the environment or scorer tells it whether they are good, prompting it to adjust toward higher scores. The signal is 'how good the outcome is,' not 'what the correct answer is.' This is the only way to teach the model to recover from errors and learn causality, as it must experiment on its own. NVIDIA divides it into two types:

Open-loop RL: The input is still a fixed frame from the dataset, the model outputs a trajectory, and the scorer assigns a score (e.g., deviation from expert trajectories, consistency of reasoning chains) without interacting with a simulator. It is cheaper, as discussed in Yurong You's talk.

Closed-loop RL: The model is placed in the AlpaSim simulator, runs a full segment, and is scored based on actual collisions, progress, and comfort. It is much more expensive (several GB per rollout), but the signal is the most realistic, and AlpaGym was built for this purpose.

The sequence of the three in Alpamayo's training pipeline is: IL pretraining as a foundation → SFT to teach reasoning format → RL (first open-loop, then closed-loop) to correct behavior. This follows the same recipe as LLMs' 'pretraining → SFT → RLHF,' but replaces 'text' with 'trajectories' and 'human preferences' with 'whether collisions occur in simulation.'

Most domestic end-to-end solutions only use the first method—imitation learning. NVIDIA's architecture reserves slots for the latter two from the design stage: SFT teaches the model 'how to think,' and RL teaches it 'whether its thinking is correct.' Without SFT and RL, an end-to-end model is essentially a larger imitation learner, and long-tail problems will not disappear simply by increasing model size.

Second, 'introspection' as a path for safety monitoring is a missing piece in domestic L3 compliance discussions. Domestic discussions on safety monitoring typically refer to rule-based fallbacks and redundant perception. NVIDIA treats 'reading the model's reasoning trajectory' as a monitoring tool—if reasoning says 'no car ahead' but perception detects a car, this is a signal that can trigger a degradation. This approach may not be mature, but it gives 'explainability' an engineering purpose beyond product marketing.

II. Reasoning Takes Multiple Forms: From Causal Chains to Latent Space

Multiple representations of reasoning trajectories, all aimed at 'reasoning like humans.' Structure: Chain of Thought, Causal Chain, Counterfactual Reasoning (action → reflection → action → reflection). Modality: Multimodal reasoning (thought → image), latent space reasoning (latent thought → latent thought). Left image: Pedestrian by the roadside and a small animal crossing

Alpamayo currently uses pure English causal chains for reasoning—the model lists causal relationships. But You said reasoning need not be limited to this, presenting options in a table. By structure: Chain of Thought, Causal Chain, Counterfactual Reasoning. By modality: Multimodal reasoning, latent space reasoning.

He focused on explaining counterfactuals: Partial observability is common in driving. In the image, a cat or rabbit crosses the road, and a pedestrian by the roadside may or may not follow. If the car can first hypothesize an action, reflect on its consequences, and then act accordingly, it would perform much better. 'This is not a linear process; it goes back and forth.' As for modality: Humans can reason visually or entirely in their minds—in latent space—without verbalizing.

Counterfactual reasoning adaptively occurs in difficult scenarios and reduces trajectory errors. Left: Horizontal axis shows scenario difficulty (following → lane change), bar charts show thinking rates, lines show trajectory errors before and after counterfactual reasoning. Right example: The system first proposes 'hold 0–1.3 s, accelerate 1.3–6.4 s,' then reflects—'A pedestrian may cross; decelerating and waiting is safer'—and revises to 'decelerate 0–5.0 s, wait 5.0–6.4 s.' CVPR 2026

Counterfactual reasoning is from NVIDIA's CVPR paper this year. In the right example: A pedestrian may step in front of the car. The system initially suggests 'accelerate' but then self-reflects—'Wait, if I accelerate now and the pedestrian continues walking, I will collide; let me adjust'—and changes to decelerate and wait. Two findings: Counterfactual reasoning adaptively emerges in harder scenarios—the harder the scenario (horizontal axis), the higher the occurrence rate; and the harder the scenario, the better the outcome when counterfactual reasoning is applied. You emphasized: 'This is not something we explicitly programmed; it is emergent behavior from the model.'

Latent space chain of thought. Top: Text-based reasoning uses ~90 tokens; bottom: Latent space reasoning uses ~38 tokens, alternating between latent world model and actions. Limitations of text-based causal chains: Slow reasoning in long sequences, weak spatiotemporal representation, weak action-text alignment. Advantages of latent space: 2–4 times faster reasoning, better trajectory quality, more suitable for RL post-training. CVPR 2026

Latent space reasoning is also from NVIDIA's CVPR paper this year. Instead of reasoning trajectories in text, it uses latent variables to represent the surrounding world model, thinking about how actions cause state changes. The PPT comparison is intuitive (intuitive): For the same decision, text-based reasoning uses 90 tokens, while latent space reasoning uses 38. Results: 2 to 4 times faster reasoning, better trajectory quality, and because it is a continuous space, more suitable for RL post-training.

Vehicle Commentary

First, the phrase 'reflection is emergent' is the most useful signal from this talk for domestic VLA teams. It means you do not need to handwrite rules for 'when to reflect'; simply include reflection samples in your training data, and the model will learn to think one step further in difficult scenarios. This transforms 'long-tail handling' from rule engineering to data engineering. Domestic teams should check whether their reasoning annotations include structures like 'first propose, then negate, then revise'—without them, the model cannot learn reflection.

Second, latent space reasoning addresses the top concern of domestic in-vehicle teams: latency. The fatal issue with text-based reasoning in vehicles is that it generates tokens sequentially—90 tokens take hundreds of milliseconds on automotive-grade chips. Latent space reasoning halves the token count and eliminates the need for decoding into text. However, it also compromises 'explainability'—thoughts in latent space are unreadable by humans. NVIDIA's solution is likely to use text-based reasoning in the cloud (readable, annotatable, evaluatable) and latent space reasoning in the vehicle (fast), connected via distillation. Domestic teams aiming for both 'explainability' and 'low latency' should acknowledge the trade-off between these goals.

III. Five Stages: How a Reasoning Model Is Trained

Five stages: VLM Training (Cosmos-Reason2, general world knowledge) → Pretraining (driving data, inject action modality) → SFT (reasoning trajectories, reasoning cold start) → RL Post-Training (open-loop RL; closed-loop RL 'coming soon') → Distillation & Quantization (in-vehicle deployment)

Here is the skeleton diagram of the entire presentation. VLM training directly uses the Cosmos Reason backbone, which has been pre-trained on large-scale internet data, bringing general world knowledge. Pre-training involves one round on driving data to handle conventional driving. You said the essence of this step is to inject the action modality into the model—letting the model know what actions look like and how to interact with the environment. SFT uses reasoning trajectories produced by an automated labeling pipeline for supervised fine-tuning, stimulating reasoning. His statement was, "Use CoT to cold-start the model, teaching it what to think and how to think." RL post-training further enhances reasoning capabilities and trajectory alignment. Distillation and quantization bring all capabilities to edge devices.

Then came the core argument of the presentation: "The biggest gain of Alpamayo 1.5 comes from scaling post-training." On the 10B model—used as a general-purpose teacher—both reasoning quality and trajectory-reasoning alignment improved significantly. The next step is closed-loop RL, currently under construction, with AlpaGym, released at this CVPR, prepared for it.

Vehicle's Commentary

First, the definitions—"Pre-training injects action modality, SFT cold-starts reasoning"—are worth domestic teams writing on their walls. They answer a common confusion: How can a language model drive a car? The answer is in two steps: First, expose it to enough "image → trajectory" pairs to know what actions are; then expose it to enough "image → reasoning → trajectory" triplets to know how to think. Skipping the first step and doing SFT directly teaches the model to "generate a plausible statement" rather than "drive."

Second, the conclusion that "the biggest gain comes from post-training" challenges how domestic teams allocate compute budgets. Most domestic teams spend their compute on pre-training and data scale, treating post-training as a finish. NVIDIA's main improvements from Alpamayo 1 to 1.5 were not changing the backbone or adding data but scaling and deepening RL post-training. This implies that the same compute, spent on post-training, may yield higher marginal returns—provided you have the capability in reward design and RL infrastructure.

IV. Scaling RL: Three Axes, 64x, 4.7%

You used three slides to explain how Alpamayo 1.5's RL post-training was scaled. The three scaling axes: learning scale, task breadth, and RL curriculum.

Left: Three scaling axes—learning scale, task breadth, RL curriculum. Top right: Asynchronous infrastructure—Rollout Workers execute policies, Policy Trainer updates policies, Orchestrator coordinates; weights via NCCL, coordination via Redis, load via HTTP. Bottom right: Learning experiences per step from 64 to 4096, 64x throughput with only 3x more time.

Learning scale: A lesson from Alpamayo 1: Model capability strongly depends on the amount of learning experience available during training. For this, an asynchronous actor-learner infrastructure was built on the open-source Cosmos-RL framework—a group of rollout workers continuously run the current model, collecting reasoning trajectories and trajectory predictions; another group of model replicas learn from them; self-generated data is evaluated, scored, and fed back as RL signals; a central controller coordinates the two groups. Result: Over 4,000 experiences per step, learning throughput increased 64x, with training time per step only 3x longer.

Task breadth: RL improves across reasoning, planning, and grounding benchmarks. Six comparisons (After RL vs Base model): Driving reasoning score +9.8%, Trajectory ADE 18.9%, Trajectory comfort +5.4%, Lingo-QA +4.5%, Lane following 22.4%, Navigation condition ADE 2.1%.

Task breadth: Alpamayo is evolving into a driving foundation model with a growing set of capabilities, and RL must keep up. 1.5 expanded reward signals to reinforce four capabilities: reasoning, driving-specific Q&A, grounding, and trajectory planning. You said, "After carefully designing reward models that provide meaningful learning signals, RL can now broadly improve Alpamayo across different tasks." All six metrics on the slide improved, with trajectory error reduced by nearly 19% and lane-following error by 22%.

RL curriculum: Interleaved SFT → RL. Phase one: RL on the autoregressive VLM; Phase two: Freeze VLM, SFT then RL on the diffusion action expert. "RL signals Throughout the entire model stack (penetrate the entire model stack)." Bottom right: AlpaSim closed-loop score—baseline 0 → Action expert SFT +1.2% → Action expert RL +4.7%.

The model has two parts—VLM and action expert—so RL must adapt to this structure. The approach is interleaved: First, do a round of RL on the VLM; then freeze the VLM, train the action expert; then do another round of RL on the action expert alone. The bar chart in the bottom right of the slide is the hardest data in this chapter: AlpaSim closed-loop score, with the action expert improving by 1.2% after SFT alone and to 4.7% after RL. "RL signals penetrate the entire model stack."

Vehicle's Commentary

First, the numbers behind "64x throughput, 3x time" reflect infrastructure, not algorithms. Asynchronous rollouts, Redis coordination, NCCL weight transfers—these are standard in LLM training infrastructure, and NVIDIA brought them to autonomous driving RL. Most domestic autonomous driving teams have not discussed this setup. To scale RL post-training, the first priority is not designing rewards but building a system that can run 4,000 rollouts per step. This is an organizational capability issue: Autonomous driving teams and large model infrastructure teams must either merge or collaborate deeply.

Second, the conclusion that "freezing VLM, training action expert alone, +4.7%" is a directly replicable experimental finding. It shows that RL for the reasoning head and action head should be done separately and in stages—doing them together interferes. If domestic teams already have a reasoning VLA, the cheapest next improvement is to follow this curriculum. Our previous article, "The DeepSeek Moment for Autonomous Driving Algorithms: Alibaba's Open-Source Qwen-Drive?" shared that Qwen-Drive also used a method of freezing certain modules during training, so this approach should see rapid adoption in production.

Third, among the six metrics, "lane following 22.4%" deserves the most attention from domestic teams. Lane following is the most basic driving capability, and a reasoning model still having 22% room for improvement here shows that the SFT-trained model is not solid in fundamentals—it has learned to "talk" but not to "act." RL fills this gap.

V. Alpamayo 1.5 and a Live Demo on an H100

Alpamayo 1.5: Flexible sensor configurations. Four capabilities: RL post-training, navigation guidance, interactive Q&A, scene grounding. No. 1 in LingoQA autonomous driving reasoning; No. 2 robot model in Hugging Face downloads.

Alpamayo has gained strong attention since CES this year, with the 1.5 release at GTC adding navigation guidance and VQA capabilities, ranking first on the LingoQA autonomous driving reasoning leaderboard and second in downloads on Hugging Face. Then You switched to the demo.

Demo scope: What the data looks like; Causal chain reasoning with navigation guidance; VQA; Checking outputs. Requirements: Python 3.12, GPU ≥24 GB VRAM (RTX 3090/4090, A5000, H100), Linux.

The environment was an H100 (80 GB) server, but the slide specified that a 24 GB VRAM consumer-grade GPU would work. He had pre-downloaded the model and a snippet from the Physical AI dataset.

Input grid: Four timesteps (t = 0.3 to 0.0 seconds) × three cameras (left cross, front wide, right cross), nighttime snow.

Two input types: images and trajectories. Ego-vehicle history uses the past 1.5 seconds. Images take four frames per camera at 0, 0.1, 0.2, 0.3 seconds, with cameras being left cross, front wide, right cross, front telephoto. Token sequence = system prompt + placeholder tokens per camera per frame (replaced with embeddings after image encoder) + trajectory history tokens; then the model outputs the chain of thought for driving and the future trajectory.

Bird's-eye trajectory distribution for navigation "turn right in 30 meters": Blue for right-turn navigation, green for counterfactual left-turn, red for no navigation, black for ground truth. Without navigation, the model's output is multimodal.

Navigation guidance: A new capability in 1.5: Telling the model "turn right in 30 meters" like Google Maps. The tokens only add "route start / turn right in 30 m / route end," with the rest unchanged. He also tried counterfactual navigation—swapping left and right. Since action generation uses a flow-matching action expert, CFG (classifier-free guidance) can amplify the navigation signal. The result is intuitive (intuitive): Blue (right) and green (left) separate, red (no navigation) scatters into multiple modes. Increasing CFG makes the turn more pronounced.

VQA notebook: question = "Describe the scene.", constructing multi-turn messages, applying dialogue templates, sequence length 1791. Same model, same data, different tasks.

VQA: The same model, same data segment, but this time outputting answers instead of trajectories. The scene is a construction zone. Asked "Describe this scene," the model replies: This is a work zone cordoned off with cones, with work vehicles and workers occupying the right side of our lane, a vehicle approaching from behind, and we should slow down. Asked "What are the key traffic elements and how should they affect driving?" the model lists construction elements, cones, and workers. Both notebooks are on GitHub.

Vehicle's Commentary

First, the implementation of navigation guidance—three tokens—shows that "guidability" essentially means having corresponding samples in the training data. The model does not understand the semantics of "turn right" but has seen enough pairs of "route start / turn right / route end + corresponding trajectory." This also explains the complaint from the Honda researcher in the Q&A: Saying "turn right in 100 meters" is ignored because the model cannot see 100 meters ahead—there are no such samples in the training data. This approach is also starting to land in domestic mass production: During discussions with Horizon HSD, they mentioned incorporating human navigation instructions into autonomous driving algorithms to improve navigation understanding, offering inspiration for other domestic teams working on navigation-conditioned VLAs. Navigation instruction distance distributions should match actual usage scenarios, not just label "next intersection."

Second, the phrase "a vehicle is approaching from behind, and we should slow down" in the VQA demo is noteworthy. The model automatically adds action suggestions when describing the scene—a habit from the SFT data format: The causal chain label format is "observation → judgment → suggestion." If domestic teams train with their own annotation formats, the output style will follow. Annotation formats define how the model speaks.

VI. Recipe, FAQs, and Alpamayo 2 Super

Available today: SFT formulations for Alpamayo 1 and 1.5, RL formulations. Planned: Model quantization formulations, AlpaGym closed-loop training formulations

How to add custom training data: Convert it to a format similar to the PAI dataset, referring to src/alpamayo/data/pai.py. How to use different rewards in RL: Implement custom rewards and integrate them into aggregated rewards, referring to recipes/alpamayo1_x_rl/rewards/aggregated_reward.py

The open-source recipe repository now offers SFT and RL fine-tuning formulations for Alpamayo 1 and 1.5, with plans to add quantization formulations and AlpaGym closed-loop training formulations. Developers most frequently ask two questions: How to add their own data by converting it to the Physical AI dataset format, following the example in pai.py for data loading. For changing rewards, refer to aggregated_reward.py, implement your own reward function, and integrate it into the aggregated rewards.

GTC Taipei Stage: NVIDIA DRIVE Hyperion Robotaxi platform, announcing the Alpamayo 2 Super Robotaxi inference model. Screen: Alpa Simulator, Alpamayo and OmniDreams open models, training data generation, Halos OS, DRIVE AGX Thor

Alpamayo 2 Super: A 32B-parameter driving foundation model, releasing this summer. Full 360° surround perception; stronger reasoning and causal chain output; meta-action output (lane changes, yielding, stopping); reasoning-based auto-labeling and visual grounding; SOTA for reasoning, prediction, and alignment tasks.

Finally, a preview. Alpamayo 2 Super, announced by Jensen Huang at GTC Taipei and releasing this summer: 32B parameters, full 360° surround perception, stronger reasoning and causal chain output, meta-action output, reasoning-based auto-labeling, and visual grounding capabilities—used for large-scale data labeling.

Vehicle Reviews

First, the aggregated_reward.py file in the recipe repository is the most valuable piece of code for domestic teams to study in the entire open ecosystem. The reward function defines 'what constitutes good driving'—it serves as the labeling guideline during the RL phase. By open-sourcing it, NVIDIA has publicly disclosed its quantitative definition of 'good driving.' Domestic teams should examine in detail what it rewards and penalizes, then ask: Are these weights appropriate for Chinese road conditions? For example, the weights for 'comfort' and 'traffic efficiency' during Beijing's morning rush hour will differ from those in Silicon Valley.

Second, in Alpamayo 2 Super's capability list, 'reasoning-based auto-labeling' is prioritized over 'SOTA performance.' This ordering is no accident. The primary use of the 32B model is to label data—it serves as the engine for automated labeling pipelines, with distillation as a secondary role. This aligns perfectly with the keynote's positioning of 'cloud-based large models acting as referees first.' When evaluating Alpamayo 2 Super, domestic teams should prioritize its labeling quality over its policy quality.

VII. Incremental Information from the Q&A Session

The ten-minute Q&A session featured eight questions with high information density.

Open-Source vs. In-Vehicle Versions: The first question asked whether the open-source model includes maps. You clarified: The open-source model is not the one currently being tested in vehicles. It accepts lightweight navigation inputs similar to Google Maps—'turn right, turn left'—but does not accept map inputs.

Do SFT and RL use the same data? Answer: There is overlap, but to 'validate the role of RL,' the data distributions for the two stages were deliberately different.

Generalization to Unseen Countries: Can it work out-of-the-box? What if one camera is missing? You said, 'We were surprised by its generalization ourselves'—Japanese users tested it, and both reasoning and trajectories were reasonable. However, to achieve optimal performance in new domains, a round of SFT on custom data is recommended. Camera configurations are flexible: two cameras (front + front telephoto) or four cameras both work.

Navigation Sensitivity: Honda Research Institute's Piyush noted that the model was insensitive to navigation instructions and output multiple trajectories. You responded candidly: Navigation serves only as a hint for the model and does not guarantee strict execution; navigation must be placed at the correct distance to be effective—'turn right in 100 meters' is ignored if the model cannot see 100 meters ahead; accuracy improves when the distance is correct. If not, use CFG. CFG is a single fixed value selected using a held-out validation set.

RL Outputs for VLM: What outputs are used as rewards when performing RL on VLM? Answer: Reasoning outputs, meta-actions, and trajectories are aggregated. Follow-up: So VLM also outputs action tokens? Yes.

Counterfactual Ground Truth: How is ground truth constructed for counterfactuals that 'never happened'? A colleague added: A teacher model evaluates action predictions and provides correct reasoning chains, which are distilled into the student model.

Perception Heads: Will perception output heads or occupancy network heads be added like Tesla's? Answer: Probably not—the goal is a foundation model that handles various tasks. To detect objects, prompt it; to add custom perception heads, decode from the final latent variables—'completely feasible.'

Vehicle Reviews

First, the clarification that 'the open-source model is not the one being tested in vehicles' draws NVIDIA's boundary. The open-source model is a research tool and ecosystem entry point, while the production model is a separate entity—likely with map inputs, more sensors, and Halos safety layers. When assessing 'how far you can go with Alpamayo's open-source model,' domestic automakers should understand: You may need research practices but cannot rely on production use.

Second, 'not adding dedicated perception heads' represents a clear architectural stance, contrary to domestic mainstream approaches. Most domestic end-to-end solutions retain perception heads—for regulatory visualization and safety fallback. NVIDIA is betting on 'one model, prompt-driven, capable of answering anything.' These two paths will diverge over the next two years: If VQA-style perception queries can meet regulatory visualization requirements, domestic multi-head architectures become redundant; if not, NVIDIA will need to catch up. Domestic teams need not choose sides now but should test the accuracy limits of 'prompt-based perception' on their own models.

Third, the Honda researcher's feedback was the most authentic information from the entire Q&A. A user who had actually tested the model said, 'It's insensitive to navigation'—something NVIDIA's PPTs wouldn't mention. Your response effectively acknowledged the limitations of navigation guidance: It relies on training distribution and fails outside it. When developing navigation-conditioned VLA systems, domestic teams should incorporate this into their requirements: The distance, type, and timing distributions of navigation instructions must cover real-world usage scenarios.

VIII. Final Thoughts: What the Autonomous Driving and Robotics Industry Can Take Away from These 33 Minutes

Summarizing the seven sections, You essentially explained one thing: From Alpamayo 1 to 1.5, the model didn't grow larger, nor did the dataset expand—what changed was post-training. By scaling RL to be larger, broader, and more systematic, the model learned to think one step further in challenging scenarios and drive well beyond just following instructions. This post-training recipe is now available for download. From the perspective of China's automotive industry, I believe five points can be directly applied:

1. Allocate a separate, substantial compute budget for post-training. NVIDIA's main improvements from 1 to 1.5 came from scaling RL post-training, not pre-training. Domestic teams should reevaluate their compute allocation strategies.

2. Perform RL for reasoning and action heads separately and in stages. First train the VLM, freeze it, then train the action expert, and finally apply RL to the action expert—this curriculum yielded a 4.7% closed-loop improvement and can be directly replicated.

3. RL infrastructure is a prerequisite, not an option. Achieving 4,000 rollouts per step and 64x throughput relies on asynchronous infrastructure from LLM training. Autonomous driving teams must either build this in-house or collaborate deeply with large model teams.

4. Study aggregated_reward.py in detail, then adjust weights for Chinese road conditions. The reward function serves as the 'labeling guideline' during RL. NVIDIA's version reflects Silicon Valley priorities—comfort, efficiency, and rule-following weights must be recalibrated for China; this is where each team needs to customize its 'recipe.'

5. Include 'reflection' samples in reasoning labels. Counterfactual reasoning emerges, but only if training data contains structures like 'propose, negate, correct.' Labeling format determines whether models learn to reflect.

Finally, a detail: When discussing latent space reasoning, You contrasted: The same decision required 90 tokens for textual reasoning but only 38 tokens in latent space. What he didn't elaborate on is that these 38 tokens are uninterpretable to humans. The first half of the talk argued for the importance of 'making models explain their reasoning'—for introspection, monitoring, and interpretability; the second half argued for 'making models not explain'—for speed. These directions are not contradictory but complementary: The cloud-based 32B model clarifies its reasoning for human review, labeling, and evaluation; the vehicle-distilled model reasons faster in latent space to drive well. Interpretability remains in the cloud, while efficiency is deployed to the vehicle—this may be the final form of reasoning-based autonomous driving. When domestic teams debate 'whether VLA needs interpretability,' they should first ask: Interpretable on which end?

References and Images

【PPT】autonomous driving with reasoning models / Yurong You* Unauthorized reproduction and excerpting are strictly prohibited.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.