Detailed Explanation of Zhuoyu ZYT World: While the Industry is Still Reviewing Recordings, Zhuoyu has Already Thrown AI into the 'GTA World' for Driving Practice

09/28 2026 557

Recently, Zhuoyu Technology, a leading domestic mobile physical AI company, unveiled its closed-loop world model, ZYT World. Its distinctive feature is that any decision made by the AI within the model receives a genuine response. It's as if the AI's steering wheel is electrified—when the AI turns left, the world turns left, and the new visuals are fed back to the AI.

To achieve this, four long-standing industry challenges that are difficult to balance simultaneously must be addressed:

Realism: Ensuring that what the AI sees matches the visuals from mass-produced vehicle cameras.

Controllable Interaction: The environment instantly responds to changes in the AI's direction and supports dynamic script modifications.

Rapid Response: Making decisions within a hundred milliseconds.

Memory Retention: When the AI returns to the same intersection, the real environment is accurately reproduced.

ZTY World has become a key tool for Zhuoyu's iterative evolution towards mobile physical AI. Zhuoyu describes its role as follows:

Downward, it provides reinforcement learning and simulation evaluation for Zhuoyu's mass-produced native multimodal foundational models, contributing over 200,000 scenario closed-loop cases for each model release.

Upward, it serves as a real-time training ground for native multimodal foundational models, enabling them to learn causal understanding and counterfactual reasoning in a multimodal parallel world.

Outward, its unique camera control capabilities extend existing data to various vehicle platforms, such as commercial heavy trucks and logistics vehicles.

Giving the World 'Memory' and Enabling AI to Understand 'Causality'

World models are not new; the novelty of ZYT World lies in its 'memory' and 'closed-loop' capabilities.

Compared to 3D reconstruction, the primary advantage of world models is scene editability. However, models can hallucinate, resulting in different environments each time the AI traverses the same route.

Zhuoyu's solution is to endow the world with memory. Its pioneering 'memory recall module' ensures that when the AI retraces a previously traveled route, the model reproduces road signs and parked vehicles accurately instead of generating random variations or skipping them.

The other half of the equation involves changing evaluation methods. The industry typically employs open-loop evaluation: the AI merely watches recordings and manipulates an unpowered steering wheel. When it says, 'I will turn left,' the recording remains unchanged, and the next frame still depicts a human driver's actions.

This training and evaluation method suffers from a critical flaw: the generated content is uncontrollable, and the reconstructed world is inflexible. Without hands-on experience, the AI never has the opportunity to make mistakes—and thus never learns to correct them.

Zhuoyu implements closed-loop reinforcement learning: every AI decision leads to different consequences, and through a reinforcement learning reward-and-punishment mechanism, the AI independently develops high-scoring strategies. The shift from open-loop to closed-loop doesn't alter simulation precision but marks the first time the AI must take responsibility for its decisions.

What Can It Achieve? AI Has Its Own 'GTA'

The GTA series is known for its realistic visuals—the cars you hijack don't change color or disappear in subsequent scenes, and every action you take has consequences. Zhuoyu has similarly created a parallel world for AI drivers that shares physical logic with the real world.

Aligned with What the AI Actually Sees:

Vehicle sensors vary in reality—some offer long-range vision, while others provide wide-angle views. Conventional approaches fall short of comprehensive coverage, resulting in generated visuals that differ from what the AI driver perceives in reality.

ZYT World achieves consistency between the virtual and real worlds, known as 'multi-view heterogeneity'—fish-eye lenses remain fish-eye, and pinhole cameras remain pinhole, all strictly aligned to the same physical moment. The same data, when processed with different camera positions and parameters, can generate perspectives for heavy trucks, logistics vehicles, and more, enabling cross-vertical expansion.

Controllable Environmental Interactions Within the World

Control your own vehicle: At the same starting point, making a U-turn versus braking creates two distinct worlds.

Control surrounding vehicles: Sudden braking by the vehicle ahead, cut-ins by adjacent cars, and e-bikes emerging from blind spots.

Control road conditions: Converting a straight road into a fork or changing a green light to red prematurely.

A Realistic, Stable Model World That Doesn't 'Lag'

Real-time performance with just two GPUs. Compressing 40 steps into 1, combined with a full suite of inference acceleration: the generator is 107.7 times faster than the 40-step teacher model, the decoder is 59.8 times faster, and video memory usage is reduced by approximately 27 times.

The bounded KV-cache mechanism prevents video memory usage from increasing over time. Training data covers only 8-second sequential fragments yet enables minute-long continuous output with stable lane markings, buildings, and lighting effects.

How Is This Achieved?

The foundation is 'frame-by-frame sequencing': each step generates a multi-view frame for a single moment, writes it to the KV-cache, and proceeds to the next step. Two Plücker adapters encode the ray geometry of fish-eye and pinhole cameras separately, while the vehicle's motion, surrounding vehicles, and lane layouts are injected frame by frame.

Compression itself is not difficult; the challenge lies in maintaining stability post-compression. Reducing 40 steps to 1, Zhuoyu accomplishes this in four rounds:

First, the model observes only past frames to align with initial conditions (causalization).

Next, the denoising path is divided into 48 small stages for gradual approximation (consistency distillation).

Then, the model uses its own output to continue for 20 moments, with the teacher model providing immediate corrections (self-rolling DMD).

Finally, a cross-camera validation step (RigCritic) is added. Image quality retains over 90% of the original 40-step performance.

The last two rounds ensure the world remains distortion-free over time and that multiple cameras remain consistent—otherwise, the AI would be training in a fake world. The memory module is a zero-initialization plugin. Connecting it has almost no impact on the base model; if disconnected, the entire layer is skipped.

Training data is generated by an end-to-end model creating alternative trajectories in real-world scenarios, followed by reconstruction and rendering. Nearly a million pairs have been accumulated. The remaining time is optimized through the inference stack. A 19M-parameter TinyVAE, W8A8 low-bit matrix multiplication, and a self-developed inference framework with hardware-aware optimizations make single inferences 1.6 times faster. Two GPUs working in tandem achieve a 1.7x speedup. Combined with the latency optimization of 1-step generation, this results in a 107.7x speedup, supporting real-time inference.

Where Does It Stand Compared to Waymo, Tesla, NVIDIA, and Huawei?

Earlier, Li Feifei categorized world models into three functions: renderers (outputting visuals for humans), simulators (outputting geometric and physical states for programs), and planners (outputting actions).

More precisely, ZYT World creates a parallel world that can be driven into, rewritten, and remembers the roads.

Several players operate in this space: Waymo World Model builds on Genie 3 for domain-specific post-training, directly generating camera and LiDAR data; Tesla uses generative Gaussian splatting for counterfactual replay of 'what if' scenarios; NVIDIA employs Cosmos for data generation and Alpamayo for action output; Huawei ADS 5 uses a world engine in the cloud to create scenarios.

ZYT World's differentiation lies in its ability to be directly used for real-time training and learning (closed-loop reinforcement learning) while genuinely integrating into mass-production R&D workflows. The figure of 'contributing over 200,000 scenario closed-loop cases per model release' speaks more to its practical value than any acceleration ratio.

*Reproduction or excerpting without permission is strictly prohibited-

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.