10/08 2026
435

Editor: Lv Xinyi
The world model has transformed from an 'end' to a 'means.'
This shift is not a natural extension along the same path. In the technological context, the world model is the goal, with researchers debating routes and defining schools of thought, all ultimately pointing toward 'modeling the world.' However, when it enters the industrial discourse, the world model becomes a means—so vast that it can accommodate any narrative, allowing all stories to fit within it.
The world model no longer answers 'how to do it' but instead answers for companies, 'where else can I go?'
After reviewing dozens of companies, we found an interesting phenomenon: everyone talks about world models, but the way they approach the concept is entirely different. Some have grown out of their existing businesses, some were born to do this, and others have naturally extended into it from other fields.
The first type is major companies that have 'grown' into world models from their existing businesses.
They are not short on technology or users; what they lack is the next growth story. The world model provides them with a reason to extend their existing capabilities further.
The most typical example is ByteDance. On September 16, the Bloomberg Billionaires Index showed that Zhang Yiming, 43, had a net worth exceeding $105 billion, making him Asia's richest person for the first time. As ByteDance's commercial value is being reassessed, its next technological move has also come to light. Bloomberg, citing sources familiar with the matter, reported that Zhang is personally driving the development of a real-time spatial video model, set to launch as early as October. This plan builds on the video generation model Seedance, aiming to extend ByteDance's existing video generation capabilities into a virtual world where users can participate in real time.
For ByteDance, the world model is not a second curve that emerges out of nowhere but rather the next step forward from what it is already doing with video generation.
The second type is native companies that have been 'about' world models since day one.
World Labs, founded by Li Feifei, is the purest example. Spatial and environmental modeling has been its core mission since the company's inception—it doesn't need to borrow or package the concept because it is inherently a world model.
For these companies, the logic is simple: the world model is their main business. They don't need it to explain who they are; they just are.
The third type is embodied AI companies that have 'naturally extended' into world models from the technical side.
This is the largest group and the one most adept at establishing new concepts. They insert many verbs or modifiers into the term 'world model,' creating variations like 'world XX model' or 'XX world model,' which appear in their technical announcements and funding PPTs.
For them, the world model serves two purposes: on the surface, it enhances robots' perception and decision-making capabilities; at a deeper level, it gives the company a broader identity than just a 'robotics company.'

Three approaches, one term. The industrial buzz around world models has been steadily amplified by these three entirely different motivations. Major companies use it to tell extension stories, native companies use it to defend their turf, and embodied AI firms use it to reach the next stage of business and technology. While all are betting on the same concept, they are not all vying for the same thing—some want users, some want orders, and some want a favorable position in the future industrial chain.
As a result, the world model has gained weight beyond its technical essence. It has become a public springboard: some use it to leap toward a distant future, some stand on it to prove they've always been here, and others use it to jump into a larger narrative.

Major companies are the most 'expected' players in the world model race.
From CV to AIGC to LLM and now world models, major companies have consistently been the primary drivers of AI commercialization. The logic is straightforward: AI is infrastructure and a strong hand for the next era. Failing to hold it means letting others take it. To stay at the table, they must continuously add new cards.
However, AI commercialization has always been accompanied by significant uncertainty, which naturally limits the imaginative boundaries of major companies' involvement. The ROI of investing in existing areas is far higher than going all-in on new fields.
Take ByteDance as an example: its vision for world models still revolves around content.
According to Bloomberg, the model is planned to serve live streaming, short dramas, and gaming, connecting ByteDance's models, cloud computing, content platforms, and Pico hardware. The sounds and movements users make while using Pico can also serve as inputs for changes in the virtual environment.
Based on this product vision, ByteDance could expand how users interact with content. The model provides video generation capabilities, the content platform offers distribution channels, and Pico provides a spatial interaction entry point. The world model allows these previously disperse (scattered) businesses to form new product combinations.
Moreover, ByteDance can use this to make changes in content forms an extension of its ability to firmly grasp (firmly grasp) a vast user base. Today's platforms merely arrange and organize content for people to watch, but when content becomes an environment that can be entered and participated in, platforms will become new competitive nodes for distributing tickets and participation rights. The world model pushes this competition one step forward.
For Tencent, which has a gaming business, how content is made and whether generated results can enter the development process are also compelling reasons to invest in world models.

Tencent's HY-World 2.0, announced in April this year, emphasizes outputting editable and savable 3D assets. After users construct scenes through text, images, or videos, they can import the results into game engines like Unity and Unreal Engine for further development. The scenes delivered by the model can thus remain in the production pipeline, allowing developers to modify and use them further.
This design corresponds to specific needs in game production. A single exploratory scene still requires work on characters, tasks, rules, and performance optimization before it can become a formal game. By being compatible with existing tools, the world model can handle part of the scene construction work, gradually entering the production pipeline without waiting for 'generating a complete game with one sentence' to become a reality.
Thus, Tencent is not betting on a yet-to-materialize end product but rather on maintaining its position in production tools and development processes as world models enter the 3D content industry. For Tencent, this provides a path to connect with its gaming business: first, change how content is produced, then push for new product forms.
Alibaba's HappyOyster also targets interactive content, but its approach to providing capabilities to developers differs.

HappyOyster allows users to change the camera, character behavior, and plot during scene operation and freely explore the virtual environment. The use cases listed by Alibaba at launch include real-time film and television production, interactive short dramas, and game concept prototypes—areas that overlap to some extent with the content markets ByteDance envisions.
What better illustrates its commercial choice is the subsequent open approach. In July, HappyOyster 1.0 entered grayscale testing for enterprises and developers, providing Android, iOS, and Web development kits as well as server-side interfaces, allowing developers to build applications like interactive plots and virtual companions. Alibaba also disclosed that Reactor had become one of the first partners to integrate.
Compared to Tencent handing 3D assets to production personnel, Alibaba hopes to open up continuous interaction capabilities to application developers, letting downstream companies find specific scenarios. This aligns with the business logic of cloud service providers: even if it remains uncertain what product forms interactive short dramas and virtual companions will ultimately mature into, as long as developers start building applications around the model, Alibaba has the opportunity to convert the world model into new cloud service and model invocation demands.
For NVIDIA, the scenarios and workflows that the world model needs to connect with change further.
Cosmos 3 already incorporates environmental understanding, generation, simulation, and actions into a unified model. Around these capabilities, NVIDIA also provides robot simulation and training tools: Isaac Sim, built on Omniverse, can create test environments and generate training data in combination with Cosmos. The world model is placed within a toolkit for robotics R&D.
This toolkit also extends toward the execution end of machines. Cosmos 3 Edge, released in July, supports running on edge computing platforms like Jetson. That same month, NVIDIA disclosed that SoftBank was developing a physical AI platform based on Cosmos, Omniverse, and Isaac Sim, while Fujitsu was exploring a collaborative control platform with Fanuc, Yaskawa Electric, and Kawasaki Heavy Industries that incorporates NVIDIA's technology.
These collaborations reflect NVIDIA's different industrial position compared to content platforms. By serving multiple robotics and manufacturing companies, NVIDIA can connect model development, simulation testing, and hardware deployment. It doesn't need to bet on a specific robot or end product but rather hopes that no matter which physical AI applications become mainstream, developers will need to use the corresponding models, simulation software, computing hardware, and deployment platforms.
After all, being the 'shovel seller' remains NVIDIA's tried-and-true classic strategy. In the world model field, the industrial position that can most attract NVIDIA is still one that keeps more companies conducting physical AI R&D within NVIDIA's entire technological ecosystem.
The product choices of these companies outline the different paths major companies take to enter the world model space. Content platforms focus on new user experiences, the gaming sector focuses on production tools, cloud service providers open up capabilities to developers, and computing platforms aim to occupy more industrial R&D and deployment processes. From this perspective, the world model has become both a means for major companies to continue (continue) their eligibility for the next round of competition and a tool to connect with existing users, content, developers, and industrial platforms to open up new markets and growth narratives.
The advantage for major companies lies in the fact that before releasing new models, they already have initial users, tools, and cooperative relationships. The world model can leverage these accumulations to enter the market and, in turn, has the opportunity to expand the boundaries of their original businesses.
However, getting a new card does not equal finding a new business. The world model can help major companies construct the next round of growth narratives, but these narratives must ultimately withstand real-world scrutiny: whether users are willing to enter and participate, whether developers can create effective products, and whether industrial clients will incorporate relevant capabilities into their daily R&D. Only when these connections truly materialize will the world model become more than just eligibility to stay at the table—it will also become a tool to expand existing markets.
(See the table below for relevant companies in this section, their representative models, projects, etc.)


While major companies integrate world models into their existing businesses, native world model companies lack such a history. For them, the world model is not an added capability for an existing product, nor is it a new card to help the company enter the next stage. Instead, it is the direction the company chose from its inception.
This 'native' identity eliminates the trouble of conceptual grafting but introduces another pressure. Major companies can rely on existing users, channels, and industrial relationships to find a foothold, while world model companies must start from scratch in defining their products, establishing tools, and building customer relationships. They must not only prove that the model can understand and generate the world but also demonstrate that this capability can become an independent business—transitioning from a demonstrable model to a foundation that others continuously use. Product definitions, tools, and scenario providers can all become hurdles.
A 'foundation' is not an identity automatically obtained upon model release but rather an industrial position that naturally emerges when other products can stand firmly on it.
World Labs is breaking this path down into specific products. Marble generates 3D worlds for creators, allowing users to edit, expand, and combine scenes or export the results for subsequent production processes. It first addresses a specific need: reducing the cost of constructing 3D spaces so that content producers in film, television, gaming, and other fields can obtain usable and modifiable scenes more quickly.

Atlas, announced on September 1, attempts to establish more unified capabilities beneath these specific use cases. It incorporates text, images, videos, and 3D information into a single model, covering scene generation, spatial reconstruction, and simulation. World Labs stated that Atlas will support subsequent versions of Marble and other products, with early access applications currently open.
The relationship between these two layers of layout (strategy) is crucial: Marble ensures that creators actually use the world model, while Atlas tries to bring the same set of capabilities into more products. For film, television, and gaming, the model can generate and reconstruct editable scenes; for robotics, similar capabilities can be used to construct training environments. In research published by World Labs in July, 'reconstructing simulated environments from reality and bringing training results back to reality' was already identified as a Entry Path (entry path) for robotics.
From this layout (strategy), the company is not vying for production rights over a specific type of content but rather for the foundational capabilities that different applications might invoke when constructing, understanding, and using 3D spaces. This resembles the position once occupied by large language model companies: model companies don't need to make all end products themselves but can allow developers to build their applications on top.
However, the editable scenes needed by creators and the reliable training environments required by robots cannot be automatically connected simply by 'the same model.' The more world model companies aim to become universal foundations, the more they need to enter the actual processes of different industries and prove the model's effectiveness one by one.
The research by Jijia Vision has precisely pushed this question into a more specific realm. From the perspective of embodied intelligence, how can a generated world be truly useful for robots?
Its GigaWorld-0 is designed as a training data engine, expanding the scenarios and interaction data accessible to robots through video generation and 3D modeling. GigaWorld-1, on the other hand, focuses on strategy evaluation: it allows robot control models to operate in generated environments and then examines whether virtual testing can reflect real-world performance.
The product logic here lies in the fact that researchers need more usable training data, more effective strategy screening, and fewer yet more efficient real-world tests. Research from GigaWorld-1 also points out that the key to determining evaluation quality lies not just in how realistic the visuals appear in a short time but in whether the model can accurately respond to actions during continuous simulation.
Jijia Vision's approach demonstrates that world model companies do not necessarily have to start with a grand platform covering all industries. Instead, they can first address a specific problem in downstream research and development, use data generation, strategy evaluation, and other products to prove the model's value, and then extend from a single link (component) to a broader technological foundation.
Going deeper into robot control, world model companies may also directly participate in action generation.
Shengshu Technology's Motus2 incorporates this possibility into the same model: the strategy component proposes candidate actions, the simulation component predicts changes after execution, and the evaluation component judges the results. These components are then connected to improve the strategy.
"
"
Compared to primarily providing test environments for external control models, this design more tightly binds world prediction to how robots act. Shengshu's technological pursuits now extend to control and learning itself. Its official projects have outlined plans to gradually release pre-trained weights, training code, and related tools in stages.
This also blurs the boundaries between world model companies and embodied intelligence companies. The former, once delving into action generation, will approach the business of robot "brain" companies; the latter, to train and iterate their products, may also independently build world models. The two types of companies may collaborate or encounter each other in certain areas. A more meaningful basis for distinguishing them lies in a company's starting point and the primary products it delivers, rather than just the label it gives itself.
However, robots are only one path through which world model companies seek market applications. Other layout (strategies) more directly point to cross-industry foundational models, using the same underlying capabilities to support more diverse application scenarios such as gaming, education, customer service, and even healthcare.
This cross-industry layout (strategy) is more evident in Runway. Its GWM-1 series simultaneously targets explorable virtual worlds, conversational characters, and robot operations: generated environments can serve gaming and immersive experiences, interactive characters can enter education, customer service, and entertainment, while the robot branch provides action-driven simulation capabilities. Runway's goal is to gradually unify different domains and action forms into a single foundational world model.
The industrial appeal behind this is to support multiple products with the same foundational research and development accumulations, rather than locking the company's market into a single type of endpoint.
Healthcare is also beginning to appear in similar collaborative landscapes. AMI Labs has established a partnership with clinical healthcare AI company Nabla, which will prioritize access to cutting-edge research, including world models, to explore the next generation of healthcare AI. Of course, related projects are still in the stage of inter-industry cooperation and application exploration, but they already reveal a trend of division of labor: model companies develop foundational capabilities, while industry partners are responsible for integrating them into specific businesses.
From 3D creation and robot research and development to interactive entertainment and healthcare collaboration, these companies do not need to redefine themselves through world models, as world models are already their core business. However, they must answer a more practical question: Why are other companies willing to let their models become the foundation connecting their businesses?
After all, being "native" only explains a company's starting point. World model companies need to both expand their model capabilities to prove they can serve different scenarios and form sufficiently concrete, irreplaceable product value in at least one scenario. The more industries they cover, the greater the imaginative space for a universal foundation; the deeper they penetrate actual processes, the higher the likelihood that models can become a business.
(See the table below for relevant companies in this section, along with their representative models, projects, etc.)
"
"
"
"
In January of this year, autonomous driving company Waabi announced a partnership with Uber, planning to expand its business from self-driving trucks to Robotaxi services and proposing to serve both application scenarios with the same AI model. The partnership goals include deploying at least 25,000 Robotaxis equipped with Waabi Driver on Uber's platform in the future. Uber has committed to additional investment but will only follow through if Waabi achieves agreed-upon milestones.
"
"
The technical system supporting this plan is composed of the company's driving model and neural simulator. Waabi World updates the simulated environment based on the vehicle's decisions, such as steering, acceleration, and deceleration, allowing the driving system to continuously make judgments on virtual roads. The company also uses paired testing to compare the performance of the same driving system in real-world scenarios and their digital counterparts, verifying whether the simulation results hold reference value.
Thus, world models become the most convenient and useful "intermediary" and footnote for embodied companies like Waabi in terms of technological prospects and developmental narratives. Technologically, world models provide more training and validation conditions for new vehicle and road scenarios. In terms of company narrative, they support Waabi's transition from a "self-driving truck company" to an autonomous driving technology provider capable of covering multiple transportation businesses.
In the robotics domain, the "intermediary" role of world models returns to specific operations and on-site conditions. When entering a new factory or store, facing changes in objects, tools, operation methods, or even the tasks themselves, companies similarly need world models to adapt existing capabilities to new task requirements.
GE-Sim 2.0, developed with participation from Zhiyuan, attempts to let robots practice in virtual environments first and then test the improved strategies on real machines. Each time the control model outputs a set of actions, the simulator simulates the corresponding scenario changes while predicting the robot's own state and evaluating the task results; the control model continues to act based on this feedback, completing a continuous task attempt. The research team also uses data and feedback from the simulation process to improve the strategy and tests its effectiveness in real-world experiments.
This echoes Waabi's approach. Even if world models do not directly handle final control, they can still transform how companies iterate their control capabilities: by moving work that requires repeated trials, screening, and strategy improvement into virtual environments, reducing excessive reliance on real equipment and human coordination. For companies with physical entities and deployment businesses, the data accumulated on-site thus has the opportunity to serve subsequent research and development, becoming a technological asset for the next delivery.
Some embodied intelligence companies choose to combine whole-machine manufacturing with model capabilities. They develop robot bodies adapted for model research and development and incorporate data collection, models, components, and hardware into their full-stack layout (strategy). These companies deliver specific devices but also hope to establish their core value on a continuously reusable world model, making the model itself the hub for organizing data, knowledge, and action capabilities—and allowing the model's development to, to a certain extent, feed back into and effectively organize the research and development of physical devices and other products.
These paths collectively reflect that world models are simultaneously entering the research and development systems and growth narratives of embodied companies. Companies hope to use them to enhance the capabilities of their existing equipment and to prove that their value lies not just in how many products they have already delivered but also in their future achievements, including the growth of model capabilities, thereby constructing a more attractive technological vision.
In other words, for embodied companies, world models are both an extension of capabilities and an extension of corporate identity. They help companies explore "what else my system can do." However, what truly determines whether this step succeeds is not how grand the model's name becomes but whether the accumulated capabilities can repeatedly transcend the boundaries of physical entities, tasks, and scenarios. Only in this way does the model's progress truly expand the company's business boundaries.
(See the table below for relevant companies in this section, along with their representative models, projects, etc.)
"
"
Companies need growth, they need narratives, and they need springboards.
This is why world models are placed here. Major firms, world model companies, embodied intelligence companies... by leveraging the term "world model," they connect their existing capabilities with future goals, seek new application entry points and industrial positions, and, amid the inevitability of technological development, compete for profitability, expansion, and even the future of AGI.
This surge is driven by technological progress and will be shaped by the choices each company makes. Concepts open up new possibilities, but the specific paths must still be forged by the companies themselves.
The next question to ask is probably whether models have entered practical applications and production processes, whether they have truly reduced training, development, or delivery costs, and whether they have achieved generality so that the next business does not have to start from scratch... Only when these connections truly hold will world models transition from a technological language companies use to depict the future into a new form of industrial division of labor.
After all, being technologically advanced, capable, and versatile is a kind of potential; being continuously adopted is what makes it a business.