Zhang Yiming Presses Pause on AI 'Distillation'

08/07 2026 365

▲ All images in this article are from the internet. Please contact us for removal if there is any infringement.

Large Model Distillation May Be Set for Change

This article was first published in Shadow Memo by Mo Yingsheng

Recently, ByteDance founder Zhang Yiming made a rare statement at the Seed team's all-hands meeting: ByteDance will not use distillation as a shortcut to enhance AI model capabilities, even if it means falling behind domestic competitors for now.

Meanwhile, ByteDance has strictly prohibited distillation of open-source models internally, reinforcing restrictions through API detection and other measures.

The timing of this statement is particularly nuanced. Earlier the same day, DeepSeek announced plans to raise API service pricing.

Even earlier, in June, U.S. AI company Anthropic had written to the U.S. Senate Banking Committee, accusing Alibaba of conducting a 'distillation attack' on its Claude model.

At a time when 'ranking chases' and 'iteration races' dominate China's large model landscape, Zhang Yiming's choice is like a stone cast into a lake, sending ripples across the industry.

Before exploring why ByteDance avoids distillation, let's clarify what 'distillation' means for the uninitiated.

Distillation is a widely used technique in AI development where a smaller model learns from the outputs of a more powerful model to enhance its own capabilities.

In simpler terms, a 'student model' mimics the reasoning process and output patterns of a 'teacher model,' acquiring advanced capabilities at a lower cost.

The theoretical foundation of this technology was proposed in 2015 by Geoffrey Hinton, known as the 'father of deep learning.' Over the next decade, distillation became a common training method across the industry. One insider stated bluntly, 'The entire industry uses distillation. U.S. companies distill from each other—it's an open secret.'

However, this neutral technology has increasingly taken on commercial and ethical connotations in recent years.

In February 2026, Anthropic claimed on its official blog that DeepSeek, Yuezhian, and MiniMax—three Chinese AI companies—had launched 'industrial-scale distillation attacks' against its Claude model. By June, it had shifted its focus to Alibaba.

Industry observers generally believe that Anthropic's distillation accusations against Chinese AI companies reflect the intensifying competition in the large model race. When technological advantages become harder to sustain, intellectual property and technical boundaries emerge as new focal points.

Against this backdrop, Zhang Yiming's 'anti-distillation' stance carries significant weight. Interestingly, none of ByteDance's Seed series models appear on Anthropic's published lists.

One TikTok Outperforms a Hundred Top Rankings

ByteDance's choice to 'avoid distillation' is no impulsive decision. Multiple sources indicate that this reflects a calculated choice shaped by external business environments, internal technological beliefs, and strategic resolve.

The most immediate and pressing concern stems from TikTok.

According to The Information, Zhang Yiming's decision was heavily influenced by the persistent regulatory pressures faced by TikTok—a ByteDance subsidiary—in global markets.

Since 2020, TikTok has continuously confronted regulatory challenges in core markets like the U.S. By 2026, these pressures had not subsided but instead grown more complex amid global discussions on AI governance.

Any technical practices that could provide ammunition to competitors—such as accusations of 'unauthorized use' of other companies' model capabilities—risk negatively impacting TikTok's global operations.

Analysts point out that compared to TikTok's global user base of billions, a few percentage points' difference on domestic large model rankings hardly justifies such risks.

When a company's core business faces such stringent global scrutiny, technological route selection transcends pure technical considerations and becomes a strategic gamble with global implications.

If TikTok-related factors represent an 'unavoidable' external constraint, long-termism serves as Zhang Yiming's proactive internal driver.

During the internal meeting, Zhang explicitly stated: Model development should adhere to long-termism and delayed gratification rather than trading others' outputs for temporary ranking gains.

He argued that distillation disrupts genuine long-term technological breakthroughs.

'Delayed gratification'—a value Zhang repeatedly emphasized during ByteDance's early years—has now been extended to the AI domain. According to ByteDance employees, internal debates over whether to adopt distillation have persisted within the Seed team for some time.

These discussions intensified with each release of powerful new open-source models in China.

While Zhang Yiming rarely speaks at Seed team meetings, his clear statement this time—that ByteDance 'should be willing to sacrifice short-term gains for long-term goals'—carries significant weight.

Only those familiar with the intensity of China's AI competition can truly appreciate the implications. This year alone, the domestic AI large model race has grown fiercely competitive, with new models debuting nearly every few days.

In such a 'progress or perish' landscape, voluntarily choosing to 'slow down'—even if it means temporary setbacks—requires not just strategic vision but immense courage.

ByteDance also stands out among China's major AI companies: It is the only one whose large language models are almost entirely proprietary.

This means ByteDance cannot rapidly address capability gaps through distillation, as companies relying on open-source models do. Instead, it must build its technological ecosystem from scratch—a slower, more challenging, but ultimately more robust path.

Zhang Yiming aims to establish the Doubao model from the ground up. This represents not just a technical route but a strategic resolve to choose the long road when everyone else takes shortcuts.

Where Does ByteDance's Confidence Come From Without Distillation?

'Rejecting distillation' is easier said than done. ByteDance's confidence stems from three core assets that cannot be easily replicated.

First is its staggering 'financial firepower.'

The competition for AI large models ultimately boils down to capital and computing power. ByteDance has demonstrated remarkable determination in this regard.

According to Bloomberg, ByteDance is discussing raising its 2026 capital expenditures to up to $70 billion. These funds will primarily target AI chips, data centers, servers, and other infrastructure, with capital sourced mainly from its 2025 profits of approximately $50 billion.

To put $70 billion in perspective: Google's 2026 capital expenditure guidance stands at $180-190 billion, Meta's at $125-145 billion, and Microsoft's AI infrastructure investment plan reaches about $190 billion.

ByteDance's investment scale now approaches that of these global tech giants. Even conservatively estimated, ByteDance has raised its 2026 AI capital expenditure plan to over RMB 200 billion (approximately $30 billion), marking a 25% increase. Some sources suggest ByteDance might even boost capital spending to around $100 billion this year or next.

Such investment levels enable ByteDance to conduct large-scale autonomous training using its proprietary computing power and data resources without relying on distillation.

Second is its unparalleled 'data flywheel.'

ByteDance operates China's largest C-end product ecosystem, including Douyin, Toutiao, TikTok, and others, covering billions of users. These products serve not only as application scenarios for AI capabilities but also as closed-loop systems for data feedback.

As of June 2026, Doubao's large model averaged 180 trillion daily token calls, growing over tenfold in the past year. By early 2026, Doubao's DAU had surpassed 200 million.

QuestMobile data shows that Doubao reached 345 million monthly active users in March 2026.

This scale effect creates a data flywheel that few pure model companies can match. More notably, commercialization is gaining traction. Doubao now offers three membership tiers: RMB 68/month, RMB 200/month, and RMB 500/month.

Volcano Engine leads China's public cloud MaaS market with a 49.5% share. ByteDance's massive user base, high-frequency usage data, and emerging business models form a closed loop that fuels continuous model iteration even without distillation.

Third is its multi-dimensional technological depth.

ByteDance is no 'newcomer' or 'specialist' in AI. In video generation models, its Seedance 2.0 caused a global sensation for producing hyper-realistic videos.

Media reports indicate that Seedance has achieved nearly 95% penetration in the short drama industry and over 80% market share among domestic video generation tools. In AI Coding, ByteDance launched Trae and has large-scale deployed AI Coding production workflows across internal systems.

In terms of model capabilities, Doubao 2.1 Pro ranks among the top tier in multiple benchmarks, scoring 59.8 in SciCode scientific code testing—surpassing GPT-5.5's 58.4—and 71.0 in Terminal Bench 2.1, close to GPT-5.5's 73.8.

In 2026, ByteDance AI will focus on four key areas: increasing investment in world model training, maintaining leadership in video models, strengthening Coding foundations, and enhancing Doubao's commercialization capabilities.

This multimodal, multidirectional technological layout means ByteDance is not merely competing for rankings in a single track but constructing a more comprehensive AI capability system. Avoiding distillation does not equate to stagnation—it simply represents a different path to progress.

An Industry-Wide Reflection on 'How to Win'

ByteDance's 'anti-distillation' choice raises a profound question for the entire AI industry: Should companies prioritize short-term rankings or build long-term capabilities?

Currently, domestic large model competition has entered a 'ranking chase' phase. New models debut nearly every few days, with leaderboards serving as the primary success metric. Zhang Yiming's statement, to some extent, rejects this competition logic.

When everyone fixates on a few percentage points' difference in rankings, are foundational research and original innovation being neglected?

ByteDance's latest large model, Seed 2.1 Pro, ranks 19th on Arena AI's Web development coding leaderboard—far below other Chinese models. According to ByteDance employees, reluctance to use distillation is cited as one reason for its large language model's lagging performance.

This represents a short-term cost. However, in the long run, if ByteDance can establish genuine technological barriers through independent innovation, its moat will prove far deeper and wider than competitors relying on distillation.

Amid sensitive accusations of distillation attacks by Anthropic against Chinese companies, ByteDance's 'no distillation' stance carries special signaling value.

As mentioned earlier, none of ByteDance's Seed series models appear on Anthropic's published lists, keeping ByteDance relatively low-profile.

This proactive declaration of 'no distillation' not only preempts external doubts but also secures greater strategic space amid complex international competition.

Notably, the distillation controversy itself is not black-and-white. Analysts point out that Anthropic itself has been accused of 'distilling' Alibaba's Qianwen model. A clear example: Its newly released flagship model, Claude Opus 4.8, claimed to be Alibaba's Tongyi Qianwen when asked about its identity.

Some experts argue that this represents a case of 'accusing others after secretly distilling their technology first,' turning technological competition into a public relations battle.

Against this backdrop, ByteDance's choice to 'avoid distillation' partially serves to delineate technological boundaries amid complex industry dynamics.

ByteDance's decision does not imply that distillation is 'wrong' or that other companies' choices are 'mistaken.' Distillation remains a neutral technology whose legitimate use drives industry progress.

ByteDance's significance lies in offering an alternative path: Building technological systems from scratch without relying on others' outputs. This route is slower and more difficult, but if successful, will create far more formidable barriers.

Some industry insiders analyze that ByteDance is following a path similar to Anthropic's. Leveraging its Volcano Engine, ByteDance can better lock in domestic large enterprise clients—a commercialization advantage over Anthropic, which lacks cloud ecosystem support.

The Costs and Vision Behind This High-Stakes Gamble

What does ByteDance's 'no distillation' choice truly entail?

The costs are visible. According to ByteDance employees, reluctance to use distillation is seen as one reason its large language model lags behind competitors.

In an industry where 'daily model updates and weekly product launches' are the norm, falling behind in rankings means losing market attention, delaying developer ecosystem growth, and potentially hindering commercialization.

More tangible challenges arise from costs. Media reports indicate that Doubao, with 200 million DAU in the first half of the year, generated daily revenue below RMB 1 million while consuming tens of millions of yuan in computing power daily.

With capital expenditures reaching RMB 200 billion or even $70 billion, balancing input and output becomes a critical challenge for ByteDance.

But in the long run, ByteDance's choices are building a moat that competitors will find difficult to replicate. First is the autonomy of its technological capabilities. Not relying on the output of external models means ByteDance's model capabilities are entirely built on its own algorithms, data, and computing power.

This autonomy will demonstrate immense value in future technological competition, especially when external models begin to tighten access permissions.

Second is the independence of its business model. ByteDance's AI product matrix forms a complete commercial closed loop, rather than being a 'parasite' dependent on an open-source ecosystem. This gives ByteDance greater initiative in business negotiations, pricing strategies, and customer acquisition.

Third is the establishment of brand trust. Amid frequent distillation controversies, the 'no-distillation' approach itself is becoming a unique brand label for ByteDance, helping it establish differentiated trust advantages in the enterprise market.

In 2026, ByteDance set its annual keyword as 'scaling new heights,' with the core goal of truly conquering large-scale model technological capabilities.

The group's top-level strategy places AI climbing as the company's top priority task, fully consolidating non-core business resources and concentrating investments in large-scale models, computing infrastructure, and Volcano Engine cloud ecosystem construction.

'No distillation' is precisely a manifestation of this strategic focus. Rather than patching up on others' foundations, it is better to concentrate resources on conquering core technologies.

Conclusion

Zhang Yiming's 'anti-distillation' declaration, on the surface, appears to be a choice of technological route, but at a deeper level, it is a strategic statement about AI long-termism.

Choosing 'slow' when everyone else is pursuing 'fast,' and opting for 'delayed gratification' when everyone is chasing rankings—this is not an easy path.

Whether this path can succeed remains to be verified by time.

But at least, it offers a different voice amid the clamor of the AI race: besides competing to see who runs faster, perhaps we should also consider where we are running to.

For ByteDance, 'no distillation' is a bold gamble—a bet that long-term technological accumulation will ultimately triumph over short-term ranking advantages, a bet that the moat of autonomous innovation is far stronger than 'standing on the shoulders of giants,' and a bet that in the AI marathon, walking steadily is more important than walking fast.

When 'distillation' has become an industry-default shortcut, do we still remember that true technological breakthroughs are never achieved through copy-and-paste?

Sometimes, slowing down is to go further.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.