09/18 2026
340

While model capabilities are rapidly approaching application requirements, data is becoming a new bottleneck for large-scale enterprise AI implementation.
AI is entering a somewhat counterintuitive phase: models are getting stronger, yet companies are increasingly finding that what truly hinders AI implementation may not necessarily be the models.
At the recently concluded 2026 Inclusion·The Bund Summit, Wang Nan, Research Manager for Enterprise Software Markets at IDC China, released a study on AI data infrastructure. One key judgment from the report is worth pondering: In an increasing number of enterprise-level applications, the capabilities of large models themselves are no longer the sole major bottleneck. What truly hinders AI from moving from pilot projects to large-scale implementation is becoming data capabilities.

This shift would have been hard to imagine two years ago.
Back then, the industry's attention was entirely focused on models and computing power. Base model parameters kept piling up, inference chips were in short supply, and GPUs dominated corporate procurement lists.
However, by 2026, the real pain points for enterprise AI implementation were changing.
A survey by Dun & Bradstreet covering 10,000 enterprises across 32 countries showed that 97% of companies were advancing AI projects, yet only 5% believed their data was fully prepared. Forrester's research earlier this year provided more specific attribution, with most RAG implementation failures traceable to data quality issues rather than retrieval algorithms or language models themselves.
The industry's bottleneck is shifting. In the first half of AI's development, the focus was on models and computing power; in the second half, it's on data infrastructure.
01 Models Are Sufficient, but Data Is the Bottleneck
A growing phenomenon is that companies acquire models and rent computing power, yet their AI projects get stuck between pilot and production stages.
Databricks' survey revealed that only 37% of executives believed their generative AI applications were ready for production. IDC's research broke down the issue further: 46% of companies stated that AI data preparation was their top priority for 2026. Approximately 90% of newly generated corporate data assets were unstructured, with documents, images, audio, video, and vector data flooding in. Traditional data architectures, primarily designed for transactions, analytics, and reporting, struggled to handle these new data types.
Add to that cross-departmental data silos, inconsistent data standards leading to distorted AI responses, and high latency preventing real-time Agent responses—each issue presented significant challenges.
Companies lacked not the willingness to adopt AI but the usable data to make AI functional.
The issue wasn't just more data but a change in how data was used. Previously, data primarily served business systems and humans; now, Agents had also become data consumers. Databases once handled primarily structured data; now, unstructured data like documents, images, audio, video, and vectors continuously entered production systems. Previously, data only needed to be storable and retrievable; now, AI required real-time, unified, understandable, and callable data.
These changes combined further exposed gaps in traditional architectures.
The growth rate of China's data generation also exacerbated this contradiction. IDC projected that China's data generation would nearly double from 76.05 ZB in 2025 to 146.65 ZB by 2030. As data volumes and forms changed, traditional architectures failed to keep pace in processing capabilities.
Databricks put it more bluntly in a blog post: Banks don't have AI problems; they have data platform problems. Fragmented and inconsistent data on customers, risks, and products prevented AI from delivering results, no matter how powerful.
In other words, AI's bottleneck was shifting from "having models" to "whether AI can access the right data."
The underlying logic was that AI's cost structure had changed. In the first half, costs centered on computing power—buying more GPUs allowed training larger models. In the second half, costs shifted to data engineering—finding, cleaning, managing, and making data understandable and callable for AI accounted for most of the time and budget in AI projects.
As models ceased to be the sole scarce resource, transforming a company's proprietary data into AI-usable resources became the new competitive focus.
02 A Reconstructed Data Foundation
If data needs in the era of large models were limited to RAG retrieval, the Agent era demanded something entirely different.
Agents needed to perform tasks on behalf of humans—checking inventory in real time, adjusting orders, making decisions, and executing operations. This required more than a static knowledge base; it needed a foundation capable of real-time, reliable, unified management of multimodal data that AI could understand and call upon.
IDC defined this emerging market as AI data infrastructure, dividing it into three main areas.
AI data platforms handled unified management and processing of multimodal data, serving as the data foundation for AI applications. AI data intelligence leveraged AI to enhance data governance, analysis, and usage efficiency, representing the fastest-growing segment with a CAGR of 48.2% from 2025 to 2030. Growth drivers shifted from traditional governance tool upgrades to data value migration driven by large models and Agents. AI data services evolved from traditional labeling to ongoing services like knowledge base operations, evaluation data, and synthetic data, with a CAGR of 30.7%.

At the conference, Wang Nan stated that future competition would hinge on who could first form a complete closed loop covering data management, data intelligence, data services, and Agent consumption, with the critical path involving small models plus multimodal foundations.
Industry players had already begun adjusting along this path.
In June, Databricks unveiled the LTAP architecture at the Data+AI Summit, unifying transactional, analytical, stream processing, and operational data into a single storage layer in the data lakehouse, eliminating ETL and data pipelines by design. Databricks CTO Matei Zaharia argued that placing data in the correct location and overlaying general-purpose intelligent agents would enable system operation—but only if the data was properly positioned. The concurrent releases of Lakebase and Agent Bricks essentially aimed to bolster the data platform layer.
Domestic moves were nearly simultaneous.
On June 29, OceanBase launched its lake-warehouse integrated AI database for the AI era during an online conference, along with products like Lakebase, DataStudio, and DataPilot. OceanBase CEO Yang Bing stated at the event that this surpassed traditional database functional upgrades—it represented an infrastructure rebuild for the AI era.
OceanBase positioned this product suite across three layers.
The bottom layer, Lakebase, integrated databases, open storage, analytics, and multimodal data processing to construct a unified foundation for all enterprise data. The middle layer, DataStudio, covered data ingestion, processing, orchestration, semantic modeling, and Agent collaboration, transforming scattered data assets into callable data services. The top layer, DataPilot, enabled business personnel to query data, perform attribution analysis, and generate dashboards using natural language, shifting data capabilities from engineers to business personnel and Agents.
From this path, OceanBase aimed to address not just database issues but how enterprise data could be genuinely consumed by AI.
This reflected a fundamental shift in data architecture.
For the past 40 years, the database industry had adhered to separated OLTP and OLAP architectures—transactional systems handled transactions, analytical systems handled analysis, with ETL transferring data between them.
This architecture sufficed for generating reports for humans but became burdensome in the Agent era. Agents required real-time access to complete business contexts. Fragmentation between transactional and analytical data, structured and unstructured data, and databases and vector databases introduced delays and inconsistencies with each data transfer.
Lake-warehouse integration sought to resolve this by establishing transactional, analytical, multimodal data, and AI applications on a more unified data foundation, reducing data transfers, and enabling AI to access business contexts more directly.
Thus, AI data infrastructure didn't simply add AI functions to traditional databases—it redefined what a data foundation should accomplish.
Databases previously focused on "storing and retrieving data" but now needed to answer: How can AI understand, call, and continuously complete tasks based on data?
03 A Trillion-Dollar Market Is Opening
IDC calculated the figures for this emerging market.
In 2025, China's AI data infrastructure market size was approximately $14.9 billion. From 2025 to 2030, it would grow at a CAGR of 37.7%, reaching $73.8 billion by 2030. By 2035, it could reach $154.8 billion (approximately 1 trillion yuan), growing over 10-fold from 2025.

For comparison, Frost & Sullivan projected China's distributed transactional database market to reach approximately 12.93 billion yuan by 2030, with a CAGR of about 18.5% from 2025 to 2030.
Placing these figures side by side, the more noteworthy aspect wasn't the market size gap but the shifting market boundaries.
Previously, database vendors competed primarily on who could support more core enterprise businesses. In the AI era, the question became who could further handle corporate data processing, data intelligence, and continuous Agent data calls. This meant database vendors faced not just the original database market but a new market expanding toward broader data infrastructure.
IDC categorized market players into three groups.
Cloud vendors extended from models and computing power to data foundations, represented by Alibaba Cloud and Volcano Engine. Their strengths lay in synergy capabilities and developer ecosystems, while challenges included data sovereignty and cross-cloud integration. Database and lakehouse vendors expanded from core enterprise data foundations to semantics, knowledge, and Agents, represented by OceanBase and Databricks. Their strengths included carrying core data and mature HTAP capabilities, while challenges involved traditional architectural burdens and model ecosystems. Independent AI data software vendors entered through vectors, governance, and feature engineering, represented by Zilliz. Their strengths included AI-native designs and rapid iteration, while challenges involved customer base and platform completeness.
These three types of vendors entered the same track ( track : competitive arena) from different control points but ultimately needed to answer the same question: Who could first form a complete closed loop from data management to Agent consumption?

OceanBase's position also warranted reexamination amid these changes.
In distributed databases, Frost & Sullivan reports showed OceanBase leading China's distributed transactional database market with a 20.2% share from the second half of 2025 to the first half of 2026. CCID's report also indicated OceanBase's revenue grew substantially in 2025, ranking first in China's distributed database market and leading in product capability within the leader quadrant of database management system vendor competitiveness assessments.
However, for OceanBase, what truly mattered lay beyond "leading."
During the upgrade cycle of domestic databases, market share reflected the ability to support core enterprise systems. Entering the new cycle of AI data infrastructure, the test became whether this capability could extend to more data types, scenarios, and Agent data consumption.
OceanBase had already secured core clients in finance, government, energy, and other sectors. Data in these industries shared common traits: business-critical, real-time, and with high safety and stability requirements.
This constituted a unique path for database vendors entering the AI data infrastructure market.
According to Bloomberg, OceanBase was pursuing external financing and using data infrastructure companies like Databricks as development references. Sources revealed its 2026 annualized revenue exceeded 1.4 billion yuan, growing approximately 70% year-over-year.
These figures indicated one thing: Market leadership was far from a ceiling—it was merely the starting point for entering the new phase of AI data platforms.
Globally, Databricks expanded from data lakes and analytics toward transactional data, while OceanBase extended from distributed databases and core business systems toward lakes, multimodal data, and AI applications. Both paths converged in the Agent era, aiming for the same destination: enterprise AI's data foundation.
The framework for evaluating such companies also needed to shift from database vendors' market share logic to data infrastructure platforms' TAM (Total Addressable Market) and ecosystem logic. IDC raised the market ceiling from the billion-yuan to the trillion-yuan level, while OceanBase transitioned from China's distributed database leader to a global AI data platform—requiring the market to adopt a new valuation coordinates ( coordinates : framework).
04 Conclusion
I recently came across a statement: Most work in AI projects actually has nothing to do with models—it's all about processing data.
This observation might warrant more attention than "models are getting stronger."
The first half of AI's development focused on models and computing power; the second half might hinge on data and calling capabilities. While everyone rushes to acquire GPUs, companies that quietly solidify their enterprise data foundations will occupy a different position in the next phase.
After all, models can be sourced from multiple providers—but a company's proprietary data exists in only one copy.
This article is an original work by Xinmou. For authorization requests or business cooperation, please contact us.
— END —