AMD Bets $8.2 Billion on Fei-Fei Li, ByteDance is Also 'Creating Worlds,' Is XR About to Change?

09/30 2026 444

AI's generative capabilities expand from text, images, and videos to spaces that can be entered, explored, and interacted with.

By/VR top

On September 28, AMD announced the acquisition of World Labs, a world model company co-founded by Fei-Fei Li, with the transaction completed entirely in stock and valued at approximately $8.2 billion. The deal is expected to close by the end of 2026, pending regulatory and other conditions.

Why can an AI company founded just over two years ago prompt a chip giant like AMD to invest $8.2 billion? The answer may lie in a shift happening within AI.

Over the past two years, AI has rapidly advanced in generating text, images, and videos, but real-world environments are not merely a series of independent frames. Objects exist in space, with positional and structural relationships to each other; when a user opens a door, picks up a cup, or moves forward, the environment continuously changes.

The core goal of 'spatial intelligence' proposed by World Labs is to enable AI to deeply understand the relationships between objects, spaces, and users in the real world. XR devices naturally become the key hardware entry point for world models to transition from 'generating spaces' to 'entering spaces.'

01 Why is AMD acquiring a company that 'creates worlds'?

World Labs has consistently been a well-funded startup in the AI field.

Upon its inception in 2024, the company secured approximately $230 million in funding; this February, it completed a new $1 billion funding round, with participation from companies like AMD, NVIDIA, and Autodesk. Previously, AMD had already collaborated with World Labs on model training and inference optimization for GPUs.

AMD's interest in World Labs centers on computing platforms. AMD noted in its announcement that as AI further advances into areas like reasoning, robotics, simulation, and Physical AI, computing demands will evolve accordingly. By incorporating World Labs, AMD can align model research with chip, software, and system development, enabling earlier involvement in defining computational requirements for new models.

Fei-Fei Li has also explained that for World Labs to expand model scale and applications, closer proximity to hardware is necessary. A lack of targeted hardware investment would limit model efficiency and scale, as well as hinder AI's extension into spatial and physical worlds.

Behind this lies a shift in AI's computational focus. Large language models primarily process tokens, while world models must also handle images, videos, 3D geometry, spatial states, and physical changes. Models need to determine what lies beyond a camera's view, the positions of objects after movement, and how user actions alter the environment. Consequently, chip manufacturers are now engaging earlier in defining model capabilities.

China's primary market is also rapidly following suit. According to a report by Vision Pro Analysis, from January to July this year, there were 105 investment events related to world model concepts in China's primary market, with cumulative financing reaching 66.6 billion yuan. The same report cautions that technical routes have not yet converged, and market enthusiasm is outpacing technological maturity.

02 Generating videos versus generating spaces: Two technical paths for world models

Currently, world models have formed two representative technical directions.

The first approach emphasizes real-time scene generation. Google DeepMind's Genie 3 can generate explorable environments in real-time based on textual descriptions. As users move forward or change direction, the model continues to predict subsequent scenes while maintaining environmental consistency for several minutes. Unlike video models like Sora or Veo, Genie 3's environment is driven by user actions rather than following a fixed timeline.

The second approach focuses on generating editable 3D spaces. World Labs' Marble, released last year, can generate spatially consistent 3D worlds from text, images, and videos, exporting them as Gaussian Splat, triangular meshes, and collision meshes for further use in other 3D workflows. Its Spark renderer can run these scenes on desktops, mobile devices, and VR headsets.

In September this year, World Labs unveiled Atlas. It simultaneously processes text, images, videos, and 3D information, placing them within a unified spatial context. Atlas covers world generation, spatial reconstruction, and spatiotemporal simulation, enabling the recovery of real-world environments from a few images or understanding spatial changes from videos.

World Labs has also proposed '3D as code.' Just as code serves as the interface between humans and machines for describing software logic, 3D could become the interface for describing spaces between humans and AI. The editable structures output by models can be further utilized by game engines, physics engines, and robotic systems.

Both approaches focus on the same question: Can AI generate a space that users can enter, explore, and interact with? They are also bringing world models closer to XR.

03 World models begin entering the XR production chain; PICO may lead the way

For the XR industry, 'worlds' are the content itself.

Traditionally, XR content involves developers creating the world first and then placing users inside it. The locations of scenes, the paths of roads, the arrangement of objects, and most interaction rules must be predefined.

World models will alter part of this process. Users provide an image, a video, or a natural language description, and the model first establishes a basic space; after the user enters, the model continues to expand the environment based on their actions. Real-world spaces can also be scanned and reconstructed first, then restyled, supplemented with content, or used to generate new scenes. For MR, once the model understands the structure of the real-world space, generated content can establish relationships with walls, tables, and the user's position.

Another change is that content does not need to be fully completed before user entry. In traditional VR games, maps, tasks, and environmental changes are mostly pre-made. When models can continuously generate environments based on player behavior, the same base content can yield more variations, with new areas dynamically extending and different players potentially facing different spaces.

The long-standing issue of insufficient content supply in XR may find new solutions through this approach.

ByteDance may be attempting to integrate models, content, and hardware. It is reportedly developing a real-time spatial video model based on Seedance, with Zhang Yiming personally coordinating cross-departmental teams, AI resources, and computational investments. The model could debut as early as October this year, aiming to create interactive virtual worlds that respond to user behavior for applications like live streaming, short dramas, and gaming.

Notably, beyond world models, PICO has also begun incorporating AI into spatial content development.

In September, PICO officially launched PICO CLI 0.5.0, a command-line tool specifically designed for AI development and a new entry point for spatial application development in PICO Space Pro and PICO OS 6. Developers can directly use natural language to describe requirements, with AI agents handling project creation, coding, building, debugging, and simulator operation. The system can also read performance and error information, then automatically modify the code, forming a 'develop—run—inspect—repair' loop. Currently, PICO CLI supports development frameworks like PICO Spatial SDK, Unity SDK, and SpatialML.

PICO CLI is not a world model itself, but it reflects another concurrent change: AI is not only beginning to generate 3D assets and spaces but is also entering the XR application development process. Tasks that previously required developers to configure SDKs, write code, build scenes, and repeatedly debug are now being repackaged through natural language and AI agents.

From this perspective, ByteDance is forming a unique technological chain: upstream, there is Seedance and the real-time spatial video model under development; in the middle, there are PICO OS 6, Spatial Engine, and AI-assisted development tools; downstream, there is XR hardware like PICO Space Pro.

Another point worth observing is the timeline. On August 24 this year, PICO announced the postponement of the PICO Space Pro, originally scheduled for release on September 2, to the fourth quarter. Officials stated that while the product's hardware and software had reached a high level of completion, the team decided at the final stage to significantly upgrade the software experience, thus pushing the release to the fourth quarter.

While there is no public evidence directly linking the Space Pro delay to this model, as ByteDance's spatial generation model, PICO OS 6, AI spatial development tools, and Space Pro all accelerate in the second half of this year, a question lingers: Will ByteDance's 'interactive virtual world' under development first appear on its own XR platform?

The likelihood is high, especially since PICO previously experimented with user-created virtual worlds through 'Light World.' If ByteDance's spatial generation capabilities ultimately integrate into PICO, it may not simply add an AI feature but revive 'Light World's unfinished vision of 'everyone creating virtual worlds' through AIGC.

04 Numerous challenges remain for XR scenario implementation

World models are still far from replacing game engines and 3D development workflows. While visual effects are easy to showcase, a truly functional world must address several fundamental issues.

The first is spatial consistency. After a user wearing a headset moves and then turns back to their original position, the table must still be there, the door must remain open, and objects on the ground must not disappear. Genie 3 can currently maintain consistency for several minutes, but gaming, social, and spatial applications may require sustained consistency for hours or even long-term preservation of world states.

The second is physics and interaction. Generating an image of a chair is not difficult, but for a user to sit on it, the system must understand the chair's structure, collision range, and weight-bearing capacity. When a user picks up a cup, the model must also comprehend whether the cup is empty, how the liquid moves, and the outcome when the cup touches the table. These capabilities require simulators.

The third is real-time performance and computational power. XR is highly sensitive to latency. Traditional game engines can stably provide high frame rates because most assets and rules are pre-completed. When every user movement requires large models to re-infer the environment, model scale, inference costs, and latency will directly impact the user experience. AMD's interest in World Labs is also tied to this workload. Future world models will need to simultaneously process high-resolution visuals, 3D structures, physical states, and real-time interactions, with computational demands differing significantly from those of large language models.

The fourth is developer controllability. Gaming and film and television (Note: ' film and television ' means 'film and television' in Chinese, retained as is for context) content cannot be entirely random; creators need to control visual styles, story pacing, level structures, and player experiences. Models must provide robust generative capabilities while adhering to the developer's design.

A more practical approach is to pair world models with traditional engines. AI handles spatial generation and understanding, while game engines continue to manage stable rules, deterministic states, and high-real-time tasks, with structured 3D information like Gaussian Splat, meshes, and collision bodies connecting the two systems.

World Labs' '3D as code' concept adopts a similar strategy. The model generates the world, while traditional systems handle state management, physics, collisions, and synchronization.

04 Epilogue

The objects generated by AI continue to expand. After text, images, videos, and 3D assets, space has become the new target.

XR devices have traditionally run content pre-made by developers. When spaces can be generated, reconstructed, and altered in real-time by AI, VR, MR, and spatial computing devices will gain new content sources. Users will enter a continuously changing world.

AMD's $8.2 billion acquisition of World Labs indicates that chip manufacturers are already betting on the next generation of computational demands. While it will take time for world models to integrate into XR production systems, they have clearly opened a new path forward.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.