09/30 2026
421
Preface:
Having mastered the art of conversation, large models are now venturing into the realm of 'real-life' experiences. Recently, JD.com has made its real-time interactive world model, JoyAI-Echo WM, available on open-source platforms. Meanwhile, ByteDance has reportedly earmarked the highest funding among all model directions for a singular focus.
Author | Fang Wensan
Image Source | Internet

AI Gains Ground in Physical Common Sense, World Models Become the Coveted Prize
The essence of world models is straightforward: empowering AI to grasp the fundamental laws of the physical world. A cup will fall when released, and obstacles lurk around corners. While language models excel in textual understanding, they've never physically held a cup of water—this is where their limitations in the real world become apparent.
Genie 3 crafts explorable environments, V-JEPA 2 excels in robot planning, and Cosmos positions itself as the bedrock for Physical AI. These advancements should not be merely lumped into the video generation category; video generation predicts only the next frame, whereas spatial intelligence must also discern object locations and the effects of actions.
Fervor cannot obscure distinctions; visual plausibility does not equate to physical accuracy. A door may appear openable but might not account for hinges and friction.
Li Feifei冷静分析 (calmly analyzes) that rendering, simulation, and planning constitute a trinity, with simulation posing the toughest challenge. Today's world models are roughly where language models stood in 2019—the code remains unbroken.
Calm analysis cannot deter the influx of hot money. World Labs secured $1.23 billion in funding at a $5 billion valuation, with NVIDIA and AMD rarely appearing together on the same investor list.
Domestically, $66.6 billion was invested in the primary market over the first seven months, with 58 out of 64 companies successfully raising funds.
Money is pouring in, but the yardstick for success is still being forged. Even the definition of world models remains a subject of debate; qualitative change may not arrive until 2028. On the revenue front, no one has yet disclosed large-scale income (scaling revenue); the technology remains in demo form, but tickets (for future gains) are already being sold—industry jargon refers to this as 'futures.'
Yet, the allure of potential rewards is undeniable. IDC projects that China's spending on embodied robots will skyrocket from $1.4 billion in 2025 to $77 billion in 2030, marking a 94% compound annual growth rate (CAGR). The State Council Development Research Center is even more optimistic, forecasting that this figure will exceed RMB 1 trillion by 2035.

ByteDance Doesn’t Lack Models—It Lacks the Next Screen
ByteDance's foray into world models is quintessentially ByteDance: acknowledge the late start, then go all in.
In 2024, Zhou Chang joined from Alibaba’s Tongyi; internally, the path was unclear, so they initially focused on video models. By 2025, two VLA (Vision-Language-Action) teams had formed—one utilizing simulated data, the other natural data.
After the 2026 Spring Festival, the landscape shifted. Seed established a new world model research group led by former Meta FAIR researcher Fan Haoqi, pursuing a 3D simulation route targeting gaming and entertainment. The two VLA teams merged under Zhou Chang, followed by the integration of Seed Robotics.
The financial commitment tells a compelling story: the world model data budget reached tens of millions of yuan, 3–4 times that of competitors. By year-end, it had closed the gap with Google Genie 3 to approximately 10%. Doubao’s DAU surpassed 200 million, forming the world’s largest AI traffic pool.
ByteDance's anxiety stems from structural challenges. Douyin has pushed 2D information flows to their attention limits; user duration (time spent) and ad load rates face physical ceilings. On text platforms, it ranks 8th; its recommendation engine needs a new screen for the next phase of growth.

Hardware: A Hard-Learned Lesson
ByteDance learned tough lessons in the hardware arena. In 2021, it spent $9 billion to acquire Pico; in 2022, the PICO 4 launch aimed for 1 million units sold. Three months later, ChatGPT emerged, shifting the winds, and the following year's target was halved to 500,000. In Q2 2023, only 156,000 VR/AR headsets were sold in China, down 50% year-over-year. Each headset generation shares the same tombstone: content drought.
World models could offer a solution. Seedance 2.5 extends single-shot generation to 30 seconds, digesting 50 reference materials. Seed3D 2.0 outputs 3D assets with joints, modularity, and simulation engine compatibility—dubbed 'half a world model.' September’s target metrics are revealing: end-to-end latency around 50ms, approximately 20 FPS, all rendered in the cloud.
50ms is a critical threshold; human tolerance for latency sits at 50–100ms. Below this line, the brain misjudges virtual spaces as physically present—a challenge the VR industry has grappled with for a decade.
Cloud rendering revolutionizes cost structures; headsets become display-and-sensor devices, shifting VR from hardware to cloud services. The Pico Space Pro, originally slated for September, was delayed to Q4—clearly awaiting this breakthrough.
Only by transitioning recommendation engines from 2D screens to 3D spaces and generating content in real-time, like 'running water,' can the perennial issue of content drought be eradicated. This represents less a technical ambition than a morphological self-rescue for a content empire.

JD.com Moves Its Supply Chain into the Simulator
JD.com's approach to world models is distinct. Its arsenal lacks DAU myths but boasts unparalleled assets: over 20 years of supply chain experience, 600,000 frontline workers, and China’s densest physical scenarios.
At the JDD Conference, JD.com open-sourced JoyAI-Echo WM. It continuously generates scenes as users move, synchronously outputting 720P video, ambient sound, and speech while supporting multi-round interactions. It topped the WBench Navigation benchmark with 81.6 points.

The model is just the tip of the iceberg. JD.com plans to collect over 10 million hours of human-real-world scene videos within two years, building the world’s largest embodied AI data collection center. Its open-source, first-person dataset EgoLive covers 1,969 object categories and 1,796 action types, attracting applications from hundreds of universities across eight countries.
The economics are clear: robots’ most expensive lesson is falling. Real-world trial-and-error costs by the hour; teleoperation data collection runs tens of thousands of dollars per hour. Simulation moves falls into the virtual world, reducing costs to electricity bills.
JD.com possesses China’s most complete simulatable reality: warehouses, sorting lines, delivery stations, and autonomous vehicles churn out physical data with action labels daily—data that gaming videos cannot produce.
Its positioning is impeccable. Over the next five years, it will procure 3 million robots, 1 million autonomous vehicles, and 100,000 drones, deploying 80 RoboBase hubs nationwide. JoyInside partners with nearly 200 hardware brands, connecting over 10 million terminals this year. Its investment list spans Zhiyuan, Unitree, Zhongqing, and Paxini, covering bodies, components, and embodied models.
The company is revaluing its assets, transforming 20 years of accumulated supply chains from cost centers into AI-era training grounds.

The Entrance Is the Facade, the Right to Collect Rent Is the Core
Public discourse often fixates on 'entrance battles,' a term worth dissecting.
Betting on headsets as the entrance repeats the 2016 VR hype script. The Vision Pro remains a toy for the wealthy after two years; Quest clings to a niche, and Pico just recovered from winter. The hardware door remains uninstalled.
The real target is defining production costs in the spatial computing era.
ByteDance bets on reducing spatial content production costs to zero, then taxing the content industry post-scarcity.
JD.com bets on reducing physical trial-and-error costs to zero, training all robots in its simulations before deployment, then taxing the robotics industry.
Both are lining up at the same ticket booth, selling tickets to 2028.
Ecosystem moves pave the road to rent collection. ByteDance builds via self-research and investments, adding Shadow Technology, Shengshu Technology, and Autovariable Robotics to its portfolio.
JD.com combines investments, channels, and real estate: 200+ brands connected to JoyInside, 80 bases claiming land, and 10 million hours of data continuously feeding models.
Both calculate clearly: when world models mature, the platform hosting others’ businesses becomes the de facto standard.
Vulnerabilities should be laid bare. ByteDance’s 2D information flows and 3D real-time interactions face a technological generation gap; Pico’s history warns that its hardware patience is thin.
JD.com’s scenarios are all in-house, with a beautiful data closed loop—but an OS’s value depends on third-party adoption. Its world model’s largest current customer may be itself.
The underlying competition is the same: sustained spatial understanding. ByteDance holds models, content ecosystems, and terminals; JD.com holds supply chains, logistics, and verifiable physical scenarios. Neither has closed the loop yet.

Conclusion:
The phase of judging models by their resemblance to the world is ebbing; the phase of competing to host others’ businesses in one’s world is rising.
The strategic value of world models likely far exceeds model APIs themselves.
Partial sources: JD Explore Academy: From AI Models to the Physical World: Highlights from JDDiscovery 2026, ByteDance Seed: Seed Research: BAGEL, The Open-Source Unified Multimodal Model, An All-in-One Model, Meta AI Research: V-JEPA 2: Advancing World Models for Physical AI