08/12 2026
352

Zhang Yiming Won't Let ByteDance Follow the Crowd; It Must Forge Its Own Path
For ByteDance to make strides in large-scale models, simply amassing GPUs may no longer suffice.
According to Enlightened Intelligence, ByteDance has recently established a new top-tier department, "AI Data and Security," on par with departments such as Seed, Flow, and Douyin. Meanwhile, the former head of AI Data and Security, Fu Yue, is set to depart, with Wang Yinglei, previously at the helm of the TikTok Platform Accountability Team, taking the reins.
At first glance, this appears to be just another routine organizational shuffle at ByteDance, but the timing is rather intriguing.
Recently, the Financial Times, citing three informed sources, reported that ByteDance is in the early stages of pre-training a super AI model with a parameter scale that could potentially reach 10 trillion. The final model size will be determined in subsequent training phases. If successful, it would rank among the largest AI models globally in terms of parameter count.
Compared to current mainstream large models, this endeavor is not merely about "scaling up the model a bit" but pushing the boundaries of scalability to new heights.
The larger the model, the more than just GPUs are required to sustain it.
01 The Model Hasn't Hit 10 Trillion Parameters Yet, but Data is Already Elevated in Importance
The establishment of ByteDance's "AI Data and Security" department did not happen in isolation.
According to Enlightened Intelligence, one of the predecessors of this new department was Global Data, founded by Fu Yue in 2023. This team, comprising around a hundred members, initially served international businesses like Dola and TikTok. Later, it gradually took on data procurement and quality control for Seed model training.
As ByteDance's large model business expanded, more data-related teams emerged.
Besides Global Data, the group's data platform DMC, the AI data platform AIDP under Flow, and various model-specific teams within Seed all established their own data teams. Although they served different purposes, their tasks increasingly overlapped: procuring data, organizing annotations, building datasets, and then handling model evaluation and quality control post-training.
Consequently, despite all preparing data for large models, ByteDance internally developed multiple parallel teams.
This year, ByteDance began restructuring this system. Data teams previously scattered across different business lines were gradually consolidated, ultimately forming the new "AI Data and Security" department, which was directly elevated to a top-tier department, on par with Seed, Flow, and Douyin.
However, given the personnel and financial resources ByteDance has invested in data, this elevation was not entirely unexpected. Enlightened Intelligence previously revealed that Seedance's data evaluation team alone has over a thousand members, with each algorithm engineer typically supported by more than ten data personnel. This year, ByteDance's data budget for areas like world models and coding has reached tens of millions of dollars, with the possibility of further increases.
This is no longer about outsourcing annotation work to a few external companies. From data procurement, cleaning, and synthesis to model evaluation and quality control, data is becoming an engineering discipline that, like algorithms and computing power, requires separate organization and sustained investment.
More critically, ByteDance has voluntarily closed off a shortcut for itself.
On August 5, The Information reported that Zhang Yiming, at a Seed all-hands meeting about a month earlier, explicitly opposed catching up by distilling external models. Citing informed sources, the report stated that Zhang Yiming believes ByteDance should be willing to "sacrifice some short-term gains for long-term goals," even if Seed falls behind temporarily, it should not rely on distilling competitors' models to close the gap.
With ambitions for models with 5 trillion or even 10 trillion parameters while refusing to directly copy others' model outputs, ByteDance has no choice but to prepare more of its own "textbooks."
And now, finding these textbooks is no longer just ByteDance's problem alone.
02 Models Keep Expanding, but High-Quality Data is Dwindling
The issue ByteDance faces is one shared by all large model companies still committed to scalability.
In the past few years, the approach to strengthening large models has been straightforward: more parameters, more computing power, and more data. However, these resources are not infinite.
In 2022, DeepMind discovered in its renowned Chinchilla paper that to train models more thoroughly under a given computing budget, both model parameters and training tokens must increase—not just one. According to their empirical findings at the time, if the parameter scale doubles, the training data volume should also increase proportionally.
The problem is that while parameters can continue to stack and GPUs can keep being purchased, the amount of high-quality content written by humans on the internet is finite.
Research institution Epoch AI once estimated that, after accounting for quality and reuse, the global publicly available human text is roughly equivalent to 300 trillion tokens.
Given the past growth rate of large model training data scales, this data could be fully utilized between 2026 and 2032; if models undergo "overtraining" for longer to reduce inference costs, this timeline could move up even earlier.
With the public internet insufficient, large model companies have started looking beyond the walled gardens.
The first targets were content that was previously difficult to scrape at scale or not freely available. Forums, news websites, paid publications, professional databases, and even physical books without digital versions have all been repriced. Thus, data acquisition, once solved by web crawlers, is increasingly becoming a procurement business.
In 2024, Google signed an annual content licensing agreement with Reddit worth approximately $60 million for AI model training. That same year, OpenAI entered a multi-year partnership with News Corp, incorporating historical and real-time content from outlets like The Wall Street Journal and The Times into its usable data pool.
However, spending money on off-the-shelf data is just the first step.
After relatively standardized content like news and forums has been divided among major players, more specialized and scarce data now requires dedicated production. In fields like coding, mathematics, and science, large model companies have begun directly hiring engineers, PhDs, and industry experts to participate in training data production and evaluation.
Going further, some data cannot even be "bought" anymore. Anthropic eventually resorted to physical books...
In January, The Washington Post, citing unsealed court documents, revealed that Anthropic had launched an internal project codenamed "Project Panama" as early as 2024. The project's approach was brutally simple: buy books, cut them up, and scan them.
Anthropic purchased millions of physical books in bulk, used hydraulic cutting equipment to remove their spines, fed the loose pages into high-speed scanners, digitized them, and then recycled the original books. Over roughly a year since the project's launch, Anthropic spent tens of millions of dollars on it. Internal documents even stated the goal as "destructively scanning all the books in the world."
More outrageously, before Project Panama, Anthropic had downloaded millions of pirated books from shadow libraries like LibGen, which later became a direct cause of legal action against the company.
In 2025, a court ruled that using books to train AI could constitute fair use, but Anthropic's practice of maintaining a "central library" of over 7 million pirated books posed infringement risks. To settle the class-action lawsuit, Anthropic ultimately agreed to pay $1.5 billion in compensation. In July of this year, a U.S. federal court formally approved the settlement, making it one of the largest known copyright infringement settlements in U.S. history.
From this perspective, Anthropic's $1.5 billion settlement fee seems like an exorbitant price tag for high-quality data. After all, books can be bought, copyrights can be licensed, and experts can be hired, but truly untouched, high-quality data will not spontaneously emerge alongside parameter growth.
This is precisely ByteDance's most troubling issue now. On one hand, there is a plan for a model with up to 10 trillion parameters; on the other hand, there is a refusal to take shortcuts through distillation. For the model to continue scaling upward, it must ultimately return to the most fundamental question: what data will it be fed?
Even if the 10-trillion-parameter goal remains a plan, data has already become a business ByteDance must independently invest in.
- END -