Not Just 'Talking the Talk': AutoNavi Enhances AI's Understanding of the Physical World

09/23 2026 556

AI's Next Challenge: From 'Talking' to 'Doing'

The next battleground for AI is not on screens.

This year, large language models (LLMs) have continuously set new records in parameter scale. AI can write poems, code, and even complete complex multi-round tasks in virtual worlds, working continuously for dozens of hours, becoming a crucial tool for enhancing productivity.

However, as AI capabilities advance rapidly, a structural contradiction remains unresolved: once AI enters the physical environment, even simple tasks like opening a water bottle cap become challenging.

LLMs excel at predicting 'the next word,' but AI in the real world needs to predict 'the next action.' The former deals with symbols and texts with sufficient training data, while the latter faces distances, directions, environmental changes, and physical laws. What is missing between the two is a set of spatial cognitive abilities to understand the real world.

The second half of the AI competition is no longer just about model parameters and computational power but about enabling machines to truly understand the real world. Whoever possesses richer spatial data, can build more accurate world models, and can translate these capabilities into actionable industrial infrastructure is becoming key to AI's entry into the physical world.

01 Why Can't Current LLMs Enter the Real World?

In the past few years, AI's greatest breakthroughs have occurred on screens with few physical constraints. AI is nearly omnipotent in the digital world, but when it steps into the real world, even pouring a glass of water becomes difficult.

Professor Su Hao from Fudan University offers an apt metaphor: models can fluently describe 'a cup will break if dropped on the ground' but never have the opportunity to feel the weight of a cup. This means that when AI faces an open environment filled with multidimensional geometry, real-time changes, and physical constraints, it lacks the ability to measure and causally reason about the physical space.

This dilemma has a classic academic expression: 'Moravec's Paradox.' Things difficult for humans (like calculus) are easy for machines, while things simple for humans (like picking up a cup) are extremely difficult for machines.

This occurs because humans continuously build an understanding of the real world from birth. Infants see objects fall and gradually understand gravity; touching different materials forms perceptions of hardness and weight; constant movement builds their sense of space.

Large models predict answers by understanding language, but language itself is not the real world.

Therefore, in structured environments, robots can complete simple production tasks by relying on pre-planned paths and fixed programs. However, when the environment shifts from factories to homes, hospitals, or even city streets, complexity increases rapidly.

But this is precisely the most crucial step for AI to interact with reality.

As a result, the flow of money in capital markets has quietly changed. Over the past 18 months, more than $10 billion has flowed into related sectors. Global AI giant OpenAI shut down its video generation app Sora early on, redirecting its team to world model research, abandoning efforts to let AI 'generate reality' and instead letting it 'understand reality.' This strategic shift is highly symbolic.

AI scholar Li Feifei has made a sharp judgment: AI represented by LLMs can generate text, images, and videos but still lacks understanding of physical space, laws, and causality. It can describe a road but may not know where it leads or what changes along the way will bring.

This is why spatial intelligence has become key to the next phase of AI competition.

In the future, AI will need not only larger models but also a data infrastructure capable of continuously perceiving, understanding, and responding to the real world. AutoNavi stands precisely at the crucial gateway connecting virtual intelligence with the real world.

02 AutoNavi's Spatial Intelligence Advantage: A Trifecta Spatial Intelligence Framework

In 2025, AutoNavi formally set its course toward spatial intelligence. Unlike other explorers in the field, as China's largest navigation app over the past decade, AutoNavi's accumulated spatiotemporal data and engineering prowess ensure that this shift is not built on thin air but represents a natural evolution of its existing technological DNA.

For how AI can enter the real world, AutoNavi's answer is not to create a smarter map but to build a spatial intelligence framework spanning cognition to action.

To support AI's deployment in the physical world, AutoNavi has constructed a trifecta technical framework at the engineering level: spatial representation, dynamic perception, and spatiotemporal reasoning. These address three fundamental spatial questions: What does the world look like? What is happening now? What changes will occur after an action?

Only by equipping AI with the abilities to understand space, perceive changes, and predict outcomes can it truly move from 'seeing' to 'understanding' and then to 'acting.'

Spatial representation serves as the gateway for AI to understand the real world.

When a language model sees 'turn right at the next intersection,' it essentially receives only a command—a string of tokens. But AI in the physical world cannot just understand this sentence; it must know: the intersection's coordinates, the number of lanes, and a series of spatial semantics.

Obtaining this information is a prerequisite for action.

To better help AI understand the physical world, AutoNavi leverages its native 3D world model and high-precision map data to transform real-world elements like roads, buildings, facilities, and terrain into computable spaces with geometric relationships, semantic attributes, and topological connections.

This is not a map for humans but a spatial database for machines. Through it, AI can precisely grasp spatial geometric features like height, distance, and topological relationships, knowing information such as 'where it is now' and 'what is adjacent.'

The dynamic perception system enables the model to understand what is happening in the world at this moment.

The real world is not a static sandbox. Road conditions change, weather changes, and unexpected events occur. Therefore, spatial intelligence cannot be just an offline map; it must be an online capability.

AutoNavi dynamically inputs real-time changes in the physical world into the model by continuously absorbing massive spatiotemporal data such as road conditions, passenger flow, and environmental dynamics. This equips AI with 'eyes' and 'nerve endings': it not only sees the past world but also perceives what is happening now.

This is crucial for embodied intelligence and urban autonomous systems because dynamic perception determines safety baselines. Can robots avoid suddenly appearing pedestrians? Can autonomous vehicles handle temporary construction ahead? These questions cannot be answered with static knowledge but require real-time perception.

However, for AI to truly enter the physical world, it must 'look ahead': if an action is taken now, what will happen in the future? This is the significance of spatiotemporal reasoning.

What AutoNavi does with spatiotemporal reasoning is to engineer this capability: it simulates AI's decision-making and task execution within strict spatiotemporal physical constraints, evaluates behavioral consequences, and makes optimal judgments.

With spatiotemporal reasoning, AI can move from 'perception' to 'decision-making,' from 'predicting the next token' to 'predicting the next action,' and bear consequences in the real world.

Spatial representation, dynamic perception, and spatiotemporal reasoning together form AutoNavi's true advantage. It is not about leading in a single technology but about creating a complete closed loop that enables AI to see, perceive, and reason accurately in the physical world.

But technological value cannot end at conceptualization. How to translate this advantage into a productive tool accessible to all developers and businesses determines how much value spatial intelligence can unleash in the real world.

03 An 'Out-of-the-Box' AI Open Platform

Under traditional models, businesses wanting to leverage spatial capabilities often need to build their own data architectures and algorithmic models. However, spatial intelligence is never as simple as integrating a map interface; it involves extremely complex technical linkages such as location reasoning, environmental assessment, and dynamic spatiotemporal computation. For most businesses, the research and development threshold and trial-and-error costs of starting from scratch are prohibitively high.

At the 2026 Yunqi Conference, AutoNavi showcased its significant open spatial intelligence Layout (layout) and released AutoNavi Qianyu, specifically designed for developers and businesses.

What AutoNavi Qianyu does is package AutoNavi's spatial intelligence into standardized, modular Agent tools. Even without coding, businesses and developers can invoke cutting-edge spatial intelligence services as easily as using general-purpose components. Through this 'out-of-the-box' platform capability, AutoNavi fully opens its foundation, redefining how businesses access the physical world.

To achieve true 'out-of-the-box' usability, AutoNavi Qianyu has completed engineering breakthroughs in two stages: transforming spatial data from 'readable' to 'usable.'

At the 'readable' stage, addressing long-standing industry bottlenecks like difficult access to spatial data and high computational thresholds, AutoNavi Qianyu created a 'digital translation' of the physical world. The platform reorganizes spatiotemporal elements such as 80+ million POIs, 5 million kilometers of road networks, and 7,500 square kilometers of urban 3D models into a 'point-line-surface-dynamic' structure, establishing a physical world data system that AI can understand and compute.

However, having computable data alone is not enough; the key is enabling Agents to possess true spatial reasoning capabilities, making data fully usable. Qianyu builds a semanticized real-world spatial data graph at the foundational level and constructs a three-tier structured reasoning chain connecting data, computation, and business. This allows Agents not only to read the physical space but also to directly invoke data for spatiotemporal anchoring and complex business decision-making.

This means Agents no longer mechanically grab coordinates but truly understand complex 'spatial semantics' and participate in decision-making. Users only need to provide a natural language instruction, and the Agent can autonomously anchor spatiotemporal ranges, organize multidimensional data for cross-computation, and output actionable, traceable decision plans.

In the future, whether for retail site selection, logistics planning, urban management, or even broader industry applications, spatial intelligence capabilities can be quickly obtained through Agents.

Currently, AutoNavi has opened its spatial intelligence to various heterogeneous terminals, injecting spatial intelligence into different devices.

In the two-wheeler mobility sector, AutoNavi's spatial intelligence solution for two-wheelers is fully compatible with mainstream operating systems like Android, Linux, OpenHarmony, and RTOS, and has established deep collaborations with brands like Niu, Yadea, Segway-Ninebot, and Ai Ma.

AutoNavi also collaborates with ecological partners to advance carbon inclusion mechanisms, thereby expanding spatial intelligence into broader industrial value chains such as safety, energy, and green mobility.

In the AIoT field, AutoNavi has partnered with Huawei, Xiaomi, Honor, and terminal brands like Qianwen Glasses, Rayneo, and Rokid to deeply integrate standardized spatial intelligence into hardware terminals such as smart glasses and smartwatches: wearing smart glasses, users can find nearby locations via voice and receive navigation guidance; raising their wrist, a single word can initiate a ride-hailing request.

04 In Closing

When AI truly steps out of screens, it no longer deals with word arrangements but with complex slopes, pedestrian flows, weather, and spatiotemporal laws. Over the past few years, large models have taught machines to 'understand information,' but understanding reality is the crucial leap for AI to evolve from a conversational tool into a productive tool.

The emergence of AutoNavi Qianyu provides businesses and developers with the springboard for this critical leap. It distills complex spatial representation, dynamic perception, and spatiotemporal reasoning into out-of-the-box development infrastructure, thoroughly (completely) breaking down application barriers for spatial intelligence.

For developers across industries, this means they no longer need to 'reinvent the wheel' to equip their Agents with the instinct to interact with the physical world. When complex spatiotemporal cognition is simplified into low-cost invocations, AI finally sheds its image as a 'talking the talk' champion confined to screens.

The path of spatial intelligence paved by AutoNavi Qianyu is enabling more businesses and developers to walk effortlessly, bringing AI firmly into the real world.

- END -

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.