Battle for the Gateway to the World Model: ByteDance Leans Left, JD.com Leans Right

09/17 2026 466

The New Game

On one side, ByteDance is developing a real-time spatial video generation model based on Seedance, with founder Zhang Yiming personally coordinating across departments to align model development, computing power, and hardware resources. On the other side, JD.com unveiled its JoyAI-EchoWM world model at the JDD Conference, announcing plans to build a cluster of 100,000 domestically produced GPUs.

These two events, separated by just a few days, may seem like major tech firms following the latest AI trend, but they actually point to a deeper shift: as the parameter race in large models hits a bottleneck, the gateway to the next era of computing is shifting from “text dialogue” to “spatial interaction.”

Over the past two years, competition in China’s AI industry has centered on large language models, focusing on parameters, context windows, and inference speed—all aimed at “better answering questions.” But the emergence of world models is fundamentally changing the competitive landscape: it’s no longer about generating text, images, or videos, but about creating spaces that users can enter, interact with, and see change in real time based on their actions.

This is not just a functional upgrade for large models but a paradigm shift in computing—from “information retrieval and generation” to “spatial understanding and experience.” ByteDance and JD.com have made simultaneous moves but chosen entirely different entry points, reflecting their distinct visions for the “gateway to the spatial computing era” and highlighting the divergent paths of China’s tech industry in the next wave of technological innovation.

World Models: Not Just “Animated Large Models” but the Gateway to Spatial Computing

To grasp the significance of this competition, we must first move beyond the misconception that world models are simply advanced video generators. Many equate real-time spatial video generation with an “upgraded version of AI video,” but their core logic is fundamentally different.

Traditional video generation follows an “input command, output clip” model, serving primarily as a content production tool. In contrast, world models operate on an “input behavior, output real-time space” model, functioning as an interactive system. Their core strength lies not in generating visuals but in understanding spatial patterns, predicting physical changes, responding to user actions, and continuously producing coherent visual experiences.

From a technical standpoint, ByteDance’s real-time spatial video generation model aims for an end-to-end latency of about 50 milliseconds and a frame rate of 20 frames per second. These numbers are more than just about “smoothness”—the human perceptual threshold for interaction latency is roughly 50-100 milliseconds. Achieving 50 milliseconds means users perceive virtually no delay between their actions and the visual feedback, making the virtual space feel “real” to the brain.

This threshold has eluded the VR industry for decades. Traditional VR relies on pre-rendered content and local computing power, leading to either prohibitively high content production costs or bulky, expensive devices—keeping it confined to niche markets. By contrast, cloud-based, real-time world models offload rendering and computation to servers, with terminals handling only display and motion capture, effectively transforming VR from a “hardware device” into a “cloud service.”

This is why Zhang Yiming personally stepped in to coordinate resources. The challenge lies not in the model itself but in cross-system engineering collaboration: the model team must solve spatial understanding and real-time generation, the cloud team must ensure low-latency computing power allocation, the Pico team must adapt terminal interactions, and the content team must implement scenarios like live streaming, short dramas, and games. A breakdown in any of these areas would make the 50-millisecond experience unattainable.

ByteDance’s strength lies in its technical foundation with Seedance, which has evolved from early multi-angle storytelling to Version 2.5’s 30-second long video generation, multimodal referencing, and editing capabilities, establishing stable visual generation abilities. Combined with Depth Anything 3’s spatial reconstruction and Seed3D’s 3D asset generation, ByteDance has effectively built a technological ladder from “generating images” to “generating spaces.”

Globally, this approach is not unique. Meta’s former Chief AI Scientist Yann LeCun has repeatedly argued that language models alone cannot enable machines to truly understand the physical world, predicting that world models will become the dominant architecture for next-gen AI. Startups like World Labs (founded by Li Feifei) and Google’s Genie are also exploring real-time, interactive virtual spaces.

What sets Chinese players apart is that while overseas giants focus on foundational technologies and general-purpose capabilities, Chinese firms are tying their efforts to specific scenarios and commercial closed loop (closed loops) from the start. ByteDance is not building a general-purpose world model platform but is directly targeting content consumption and VR terminals, aiming to solve its core challenges: how to evolve Douyin’s content from 2D screens to 3D spaces and how to differentiate Pico in a crowded hardware market.

In other words, world models are not standalone products but the “operating systems” of the spatial computing era. Their value lies not in the model’s raw power but in the applications, content, and users they can support. Just as Windows defined the PC era and iOS defined mobile, world models are now vying to define the entry point and ecological rules of spatial computing.

ByteDance’s “Content Closed Loop”: Bringing People into Virtual Worlds, Breaking Ground on the Consumption Side

ByteDance’s world model strategy is deeply rooted in its “content DNA.” Its logic is clear: build on Seedance’s video generation capabilities, extend into real-time spatial interaction, and deploy through Pico headsets and the Douyin ecosystem, forming a closed loop of model, computing power, hardware, and content.

Zhang Yiming’s personal coordination aims to unify these four previously independent business lines into a cohesive force, rapidly translating technical capabilities into user experiences. The first pillar of this approach is Douyin’s massive content ecosystem and user base. ByteDance is not positioning world models as a B2B technical service but as a direct C-end consumption scenario, including interactive live streaming, interactive short dramas, and immersive gaming.

These scenarios share three key traits: high demand for visual quality, sensitivity to interaction latency, and mature monetization models.

While AI video generation has primarily served creators by reducing content production costs, real-time spatial models directly serve consumers by transforming content consumption. Users no longer passively watch videos on screens but step inside them, becoming part of the scene. For example, in interactive short dramas, user choices dynamically alter the narrative; in live streaming, audiences can “enter” the host’s virtual space, achieving true “presence.”

The second pillar is Pico’s hardware positioning. Since acquiring Pico in 2021, ByteDance has continued investing in VR hardware. The planned September release of Pico Space Pro was delayed to Q4 2026, with official statements citing “a major software experience upgrade”—widely believed to be related to world model integration.

Behind this lies a shift in VR hardware competition: from “spec wars” to “experience wars.” Past VR headsets competed on resolution, refresh rate, and weight, but without compelling content and interactions, even high-end hardware remains a “viewing device.” World models provide Pico with an “infinite content generator,” enabling real-time virtual space generation based on user demand, eliminating the need for Manufacturer (manufacturers) to pre-produce VR content.

This “cloud generation + terminal display” model also significantly reduces headset hardware costs, making immersive experiences accessible without high-end GPUs—crucial for VR adoption.

However, this path has clear boundaries. ByteDance excels in consumer internet and content distribution but lacks expertise in physical scenarios and industrial data. Its world models primarily serve virtual experiences like entertainment, socializing, and content consumption, struggling to penetrate industrial, logistics, or robotics applications.

This means ByteDance’s world models essentially upgrade content consumption rather than transform the physical world. They redefine video and VR formats but fall short of directly tapping into the vast industrial market. Meanwhile, content consumption scenarios prioritize visual effects and creative expression over physical accuracy and task reliability—requirements that diverge sharply from industrial needs.

From an industry perspective, ByteDance’s content-first approach aligns with its strengths. It doesn’t need to build scenarios from scratch but can migrate its existing content ecosystem and user base into the spatial computing era, rapidly forming a commercial closed loop . The challenge lies in the relatively low ceiling of content consumption, whereas the long-term value of spatial computing extends far beyond entertainment.

JD.com’s “Industrial Closed Loop”: Bringing AI into the Physical World, Trading Supply Chain Expertise for Entry Tickets

In stark contrast to ByteDance’s consumer-focused strategy, JD.com’s world model is anchored in the physical world. At the JDD Conference, the launch of JoyAI-EchoWM was just the tip of the iceberg. The real differentiators are the planned 100,000-card domestic GPU cluster, the collection of 10 million hours of real-world scenario data, and industrial applications spanning logistics, manufacturing, and healthcare.

JD.com’s logic is that world models’ ultimate value lies not in creating virtual spaces for entertainment but in simulating the physical world, enabling AI and robots to train, validate, and refine their behaviors in virtual environments before deployment in real-world settings—achieving a closed loop of virtual simulation and physical execution.

The first core of this approach is building a data moat with real-world scenario data. Unlike large language models that rely on internet text, world models and embodied AI require first-person perspective (perspective) data from human operations in physical spaces—sorting packages, assembling parts, organizing shelves. These seemingly simple actions embody complex physical laws that AI struggles to grasp.

JD.com’s advantage lies in its unparalleled supply chain ecosystem: warehouse sorting robots, autonomous delivery vehicles, and unmanned distribution devices generate vast amounts of real-world physical data daily. Building on this, JD.com has initiated a multi-million-hour collection of human operation data, with its first open-source dataset, EgoLive, covering 346 real-world tasks.

This data isn’t scraped from the web but sourced directly from industrial scenarios—a barrier internet companies struggle to replicate.

The second core is the deep integration of computing power and scenarios. JD.com’s planned 100,000-card domestic GPU cluster isn’t about brute-forcing computational scale but meeting the low-latency, high-reliability demands of physical AI. Unlike large language models, where a text generation error is negligible, a robot’s misstep in a warehouse or a self-driving car’s judgment error can cause real damage.

This requires computing power that is not just large but stable and fast, capable of supporting large-scale real-time simulations and parallel scheduling. Moreover, JD.com’s computing cluster isn’t just for internal use—it’s open to the entire industry, aiming to transform its supply chain-validated AI infrastructure into a public service for industrial players.

The third core is the industrial closed loop from simulation to deployment. While JoyAI-EchoWM topped the WBench Navigation benchmark with 81.6 points, for JD.com, model benchmarks are just the starting point. The real focus is vertical scenarios like logistics, manufacturing, and healthcare: Logistics Brain 3.0 compresses path planning for billions of packages from minutes to seconds, achieving a 96.7% success rate in multi-task logistics scenarios for embodied AI; in manufacturing, natural language inputs can now generate mechanical designs and directly output 3D-printed prototypes; in healthcare, the model assists with medical imaging and diagnostics.

Here, world models act as “virtual training grounds,” where robots and AI repeated practice (repeatedly practice) and validate solutions before real-world deployment, drastically reducing trial-and-error costs.

This path offers high industrial barriers and long-term value but also significant shortcomings. JD.com lacks C-end entry points and content ecosystems, making it harder to reach mass users quickly compared to ByteDance. Model iteration and experience optimization are constrained by scenario limitations. Industrial deployments also have long cycles and high customization needs, making it difficult to achieve the rapid scalability seen in consumer internet. However, this also means that once a closed loop is established, the competitive moat becomes much deeper.

Two Paths, One Long-Term Race for the Gateway

ByteDance and JD.com’s divergent strategies reflect their unique strengths and choices, not superiority. ByteDance enters through content consumption, aiming to make spatial computing the gateway to next-gen entertainment and socializing. JD.com enters through industrial scenarios, aiming to make world models the infrastructure for physical AI. Though their track (tracks) appear different, both are vying for the same prize: ecological dominance in the spatial computing era.

The previous large model competition was essentially a “capability race” focused on parameters and performance. The world model competition, however, is an “ecosystem race” centered on integrating models, computing power, data, scenarios, and terminals into closed loops. The model itself is no longer the barrier—the real barrier (moat) lies in the scenarios and data behind it.

ByteDance has content and users but lacks physical scenarios; JD.com has industrial data but lacks C-end entry points. Both are extending from their core strengths to define the rules of next-gen computing.

It’s worth noting that this race has just begun. China’s world model sector is still in its early stages of technical validation and scenario deployment. Neither ByteDance’s real-time spatial video nor JD.com’s physical world simulation has achieved large-scale commercialization yet.

Overseas giants like Google and Meta are also accelerating their deployments, with their technical routes and business models still taking shape. For Chinese companies, the real opportunity does not lie in following overseas technical routes, but in leveraging their own scenario-based advantages to forge a path suited to the Chinese market.

From a longer-term perspective, world models will bring not only technological transformation but also a fundamental shift in human-computer interaction. When AI can understand, generate, and interact with spaces, the way we use computers will transition from 'operating on a screen' to 'acting within a space.' This represents the third computational paradigm shift after PCs and smartphones, and the competition for entry points often determines the industrial landscape of an era.

The paths taken by ByteDance and JD are just the beginning of this great migration. Over the next three to five years, more players will enter the fray, and more scenarios will be redefined. Ultimately, victory will not be determined by the size of model parameters, but by who can first make spatial computing a part of ordinary people's lives and penetrate every corner of industry.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.