AI Data 'Refinery' | Research on China's Data Annotation Industry

09/15 2026 389

The upper limit of a large model's capability depends on the quality of its training data. As an essential upstream component of the AI industry, data annotation is indispensable for large model training, instruction fine-tuning, and RLHF (Reinforcement Learning with Human Feedback). With the rapid development of the domestic large model industry, coupled with globalization trends, the data annotation industry is moving away from traditional labor-intensive models towards a new development stage characterized by 'technology-intensive + global delivery.'

1. Emergence Stage (2010-2015): The AI concept began to rise, with demand primarily focused on simple image classification and speech transcription. Annotation tools were rudimentary, mostly completed by in-house corporate teams. The release of the ImageNet dataset spurred development in the computer vision field.

2. Growth Stage (2016-2019): The outbreak (burst) of deep learning technology, along with the commercialization of autonomous driving and facial recognition, rapidly expanded market demand. Crowdsourcing models emerged, and professional data annotation companies began to emerge. The industry market size reached 2.586 billion yuan in 2018.

3. Explosion Stage (2020-2023): GPT-3 ignited a global large model wave, with RLHF driving high-quality annotation demand. The market size surpassed 7 billion yuan in 2023, with Hithink Royalflush listing on the STAR Market, becoming the first AI training data stock.

4. Intelligent Globalization Stage (2024-present): Large model pre-annotation technology has matured, with human-machine collaboration becoming the industry norm. Domestic policies are being intensively implemented, with leading companies accelerating overseas expansion, and Southeast Asia becoming a key global data delivery hub. The market size reached 11.753 billion yuan in 2025, with the National Data Bureau launching '7+32' pilot cities.

Policies and technology constitute dual drivers of industry development.

At the policy level, data annotation is included in the 'AI Plus' initiative and is a key link in the market-oriented reform of data elements. The state has set a target of over 20% average annual growth for the industry by 2027, with annotation services officially included in government procurement catalogs.

Economically, the scale of China's core AI industry exceeded 600 billion yuan in 2025, with over 400 large model companies, and annotation demand expanding from the internet sector to vertical markets such as manufacturing, healthcare, and finance.

Technologically, the 'large model pre-annotation + manual calibration' model has become widespread, improving annotation efficiency by 35 times. Synthetic data, multimodal annotation, and automated cleaning tools continue to iterate, with speech automation cleaning meeting over 90% of business needs.

It is estimated that the domestic data annotation industry saw a compound annual growth rate of 24.2% from 2018 to 2025, with a market size of 11.753 billion yuan in 2025. It is expected to grow by over 27% year-on-year in 2026, with the overall scale surpassing 15 billion yuan.

Currently, there are 126,000 high-quality datasets in China, with a total volume of 1,815PB. China's data annotation market accounts for approximately 30% of the global share.

Exponential data demand from large model training, inference, and fine-tuning is the core driver of industry growth, while government procurement and vertical industry AI adoption continue to open up incremental space.

- Upstream (Data Providers): Internet platform behavioral data, government public data, sensor-collected data, and AI-synthesized data supply raw materials for the industry.

- Midstream (Annotation Platforms and Tools): Includes annotation tools like CVAT and Wenxin Biaozhi, crowdsourcing platforms like Longmao Crowdsourcing, and cloud-based SaaS systems, while also handling data cleaning and automated quality inspection.

- Downstream (Annotation Services and Applications): Professional data service providers, large model companies, autonomous driving firms, and enterprises in vertical sectors such as healthcare, industry, and finance complete data application.

1. First Tier: Professional Data Service Providers (Hithink Royalflush, DataBaker, Testin Data, Datatang)

These companies specialize in vertical sectors, possess high technical barriers, enjoy strong customer loyalty, and higher gross margins. For example, Hithink Royalflush holds an 18.3% market share, focusing on high-barrier semantic annotation in finance and healthcare, while establishing a thousand-person overseas annotation base in Southeast Asia, generating tens of millions of dollars in overseas revenue and covering 176 countries.

DataBaker focuses on the speech sector, supporting emotional annotation in 32 dialects, deeply serving intelligent cockpits and manufacturing digitalization scenarios.

2. Second Tier: Internet Giant Platforms (Baidu Crowdtest, JD CrowdIntel, Huawei Cloud Annotation, Alibaba Cloud)

Leveraging internal demand from their ecosystems, these platforms embed large model capabilities into annotation processes, offering significant cost advantages. JD CrowdIntel holds a 9.5% market share, with a million-strong crowdsourcing workforce, excelling in large-scale standardized annotation for e-commerce scenarios. Baidu Crowdtest relies on the Wenxin Biaozhi system for automatic rule generation.

3. Third Tier: Crowdsourcing Platform Companies (Longmao Data, Beizhi Data, Huiting Tech)

These companies possess large crowdsourcing user pools, deliver fast response times, and excel in handling large-volume standardized annotation orders. Longmao Data has over 4 million crowdsourcing users, capable of responding to customer needs within 72 hours.

In the future, industry concentration will further increase, with CR5 expected to rise from 45% to over 60%, accelerating the exit of small and medium-sized service providers lacking technical capabilities.

The overseas expansion of domestic annotation companies is undergoing a qualitative transformation: from simply exporting low-cost labor in the past to exporting comprehensive capabilities in technology, data, and services.

Led by Hithink Royalflush, leading companies have established a global network featuring 'Southeast Asian delivery bases + localized sales in Europe and the Americas + subsidiaries in Japan, South Korea, and the EU.' The Southeast Asian base has formed a thousand-person annotation team, contributing tens of millions of dollars in revenue in 2025. The report predicts that the Southeast Asian base will add 300-500 personnel in 2026.

Four mainstream overseas expansion models:

1. Overseas Delivery Base Model: Establishing annotation bases in Southeast Asia to handle global large-scale orders;

2. Multilingual Data Service Model: Building multilingual corpora for ASEAN to serve the localization needs of domestic AI companies going overseas;

3. Technology Export Model: Assisting overseas countries in building large model data infrastructures and participating in global data standard-setting;

4. Direct Global Customer Service Model: Obtaining GDPR and ISO27001 international certifications to directly handle customized orders from overseas tech giants.

While the industry is rapidly expanding, it is also facing multiple real-world pressures amid intense red ocean competition. Market price wars continue to squeeze corporate profit margins, with even leading companies experiencing declines in gross margins. Coupled with high accounts receivable due to long payment terms from B-end clients, many companies face significant cash flow pressures, constraining reinvestment in R&D. Meanwhile, the supply of expert annotators needed for industrial upgrading is insufficient, with high-value-added businesses like RLHF and instruction fine-tuning relying on professionals with industry knowledge. However, domestic talent certification and training systems remain underdeveloped. With stricter domestic and international regulations, compliance requirements for data privacy, cross-border data flows, and GDPR continue to raise operational costs for companies. Additionally, technological iterations in large models and synthetic data pose risks of traditional manual annotation businesses being replaced. Furthermore, the lack of unified industry standards for quality and pricing further disrupts market ecosystems through low-price, disorderly competition.

At this critical juncture of industry transformation, human-machine collaboration will become the foundational production model, with sustained demand for expert annotation and multimodal data. Coupled with overseas business expansion, the implementation of data element marketization policies, and the widespread application of synthetic data as a key supplement in scarce data scenarios like autonomous driving and healthcare, market resources will further concentrate toward technologically capable leading companies. Vertical sectors with high barriers, such as finance, industry, and healthcare, will offer greater pricing power. For industry participants, the core paths to navigate through cycles include increasing R&D in AI pre-annotation and automation tools, Layout (laying out) expert annotation businesses, seizing overseas expansion windows to strengthen multilingual and compliance capabilities, and deep cultivation (deeply cultivating) vertical sectors to accumulate industry knowledge. From an investment perspective, companies with AI pre-annotation technology, expert annotation capabilities, global service capabilities, and vertical industry barriers are more likely to achieve long-term growth potential. The report predicts that the domestic data annotation market size will exceed 30 billion yuan by 2030.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.