08/17 2026
526

Original content by the editorial team of Keji Sishao © Youliao Business
Author | Fei Ma
In August 2026, Zhang Yiming made a notable statement during an internal meeting of ByteDance’s Seed team: ByteDance will not rely on distillation as a shortcut to enhance AI model capabilities, even if it means temporarily lagging behind domestic competitors. Internally, ByteDance strictly prohibits the distillation of open-source models and has tightened restrictions through measures such as API detection. Liang Rubo reinforced this stance by declaring, “ByteDance’s large language models will firmly pursue independent research, focus on foundational strengths, and accept short-term setbacks.”
Suddenly, “avoiding distillation” seems to have been elevated to a virtue in technical circles.
But this raises the question: What, exactly, is wrong with distillation?
If large models should not be distilled, then should reference works like the Xinhua Dictionary also be avoided? From childhood to adulthood, which piece of human knowledge hasn’t been “distilled” from others? Labeling a fundamentally normal learning behavior with a moral stigma is, in itself, a form of cognitive bias.
What is distillation? It’s not theft; it’s learning.
First, let’s clarify the concept. Model distillation involves training one’s own model using outputs generated by more advanced frontier models. In simpler terms, it’s like a less capable “student model” learning from a stronger “teacher model”—observing how it answers questions, reasons, and solves problems, and then using these insights to improve itself.
This technology is hardly new. In 2015, Geoffrey Hinton, often referred to as the “father of deep learning,” and his colleagues systematically published the foundational work on knowledge distillation. Over the following decade, distillation evolved from a tool for model compression into an almost indispensable component of large model training.
Distillation is fundamentally different from “code theft” or “database piracy.” It does not involve directly obtaining another model’s weights or complete training datasets. Instead, it learns by observing a model’s “outputs”—the answers it provides. It’s akin to sitting next to a top student and watching them solve problems; you’re not peeking into their mind but observing the process they write down on paper.
“Isn’t this just learning?”
Jensen Huang put it more bluntly: “Distillation—learning from AI, learning from others, and learning from other knowledge sources—is the foundation of intelligence.” In his view, the transmission of human knowledge is an ongoing process of “mutual distillation.”
Mark Zuckerberg also weighed in to defend distillation. In an article titled “The Future Belongs to Everyone,” he argued that the ability of models to learn from one another is a core principle for the functioning of open-source ecosystems, and that “all artificial intelligence models originate from human knowledge.” He called for protecting the principle of “learning from anything observable.”
Clearly, no one in the industry who truly understands technology views distillation as shameful.
Where does the “shame” around distillation come from?
The stigmatization of distillation is not a technical issue but a political one.
In February 2026, the U.S. AI company Anthropic released a report accusing three Chinese companies—DeepSeek, Yuezhi’anmian, and MiniMax—of launching “industrial-scale distillation attacks” against its Claude model. In June, Anthropic further alleged that Alibaba had interacted with Claude approximately 28.8 million times using 25,000 fake accounts.
However, let’s examine Anthropic’s own track record—the company was collectively sued by authors for training its model using roughly 482,000 copyrighted books downloaded from pirate websites, ultimately paying $1.5 billion in settlements. When it was illegally scraping books, it didn’t raise moral concerns, but now, merely observing a model’s outputs is labeled “theft”?
A commentary by TMTpost hit the nail on the head: Distillation cannot be equated with plagiarism or theft. “This most common technique in the industry has been thoroughly stigmatized by Anthropic.”
More intriguingly, mutual distillation among U.S. AI companies is commonplace. Developers even discovered that Anthropic’s own Claude model occasionally output answers like “I am Tongyi Qianwen” or “I am DeepSeek,” leading to industry jokes about “distilling Chinese models.”
This is not a technical controversy but a struggle for narrative control. When technological advantages become difficult to sustain, intellectual property and technical boundaries become new focal points for competition. Stigmatizing an opponent’s legitimate learning behavior allows one to claim the moral high ground in public opinion—a tactic that has been repeatedly employed, from semiconductors to AI, without ever changing the script.
The legitimacy of distillation is well-established.
How does the legal perspective view distillation?
Scholars, drawing on analytical frameworks such as the “three-step test” and “transformative use,” have demonstrated that DeepSeek-R1’s distillation does not constitute copyright infringement. The legal community argues that the characterization of AI model distillation should promote a dynamic balance between “technological innovation—rights protection—public domain” and should affirm the legitimacy of distillation.
Additionally, legal research highlights that the essence of model distillation is the creative transformation of knowledge in the public domain. Distillation technology aligns with the public domain values pursued by the intellectual property system in terms of technological progress and interest balance.
In simpler terms: Distillation itself is not illegal. What may raise legal concerns is violating service agreements (e.g., closed-source models’ terms of use explicitly prohibiting distillation) or using fraudulent means to massively call APIs. However, these are issues of “how to distill,” not “whether distillation is allowed.”
Technology is neutral. Elevating a neutral engineering method to the level of moral judgment is neither professional nor fair.
ByteDance’s decision to avoid distillation is commendable, but it should not be treated as gospel.
Returning to ByteDance’s case, Zhang Yiming’s rationale is that excessive reliance on distillation may weaken a company’s independent R&D capabilities; distillation can interfere with truly meaningful long-term technological breakthroughs; if training data largely originates from competitor models, one can only learn the opponent’s capability boundaries and fail to reach the true core of innovation.
Are these concerns valid? Yes. If a team becomes accustomed to “copying homework,” it may indeed lose the ability to solve problems independently. ByteDance’s desire to take a more difficult but autonomous path deserves respect.
But the problem is: Avoiding distillation and pursuing independent research are not mutually exclusive choices.
Multiple large model researchers agree that “completely avoiding distillation will definitely slow progress. Realistically, it seems unlikely to achieve industry-leading Coding and Agentic capabilities in the short term without relying on distillation.” One researcher made an analogy: A normal training pipeline should use distillation as a “cold start” to quickly bring the model to 90% performance before tackling the remaining 10%.
Distillation is learning. Does learning hinder innovation? Newton said, “If I have seen further, it is by standing on the shoulders of giants.” Einstein studied all previous physical theories before proposing relativity. If “starting from scratch” is the only way to be original, then human civilization would not exist.
ByteDance has chosen an extreme path—not only avoiding distillation of closed-source models but also banning open-source models. This “one-size-fits-all” approach seems less like a technical judgment and more like a performative stance.
The right to learn should not be morally coerced.
This debate over distillation reflects a deeper issue: In the rapidly evolving field of AI, what mindset should we adopt toward “learning”?
Some idolize “independent research” while disparaging “learning” as if only building from scratch demonstrates true capability, and borrowing from others is taking shortcuts or lacking backbone.
This mindset would be absurd in education—which student doesn’t first learn from textbooks, listen to teachers, and reference previous work before developing their own insights? If “learning” is shameful, then schools worldwide might as well close.
Distillation is AI’s version of “going to school”—it observes others’ answers, learns their thought processes, and understands their reasoning before forming its own interpretations. This is not plagiarism; it’s growth.
Jensen Huang put it clearly: “Humans have always learned from each other, and AI must learn from something.” Early models learned from human content on the internet; now, models learn from other models’ outputs—essentially absorbing existing knowledge to generate new insights.
If large models should not be distilled, then should the Xinhua Dictionary, encyclopedias, and textbooks not exist? Human learning is a form of “distillation”—concentrating, digesting, and recreating the wisdom of predecessors.
Don’t stigmatize distillation. Learning is never shameful; what’s shameful is treating learning as an original sin.
ByteDance can adhere to its technical path—that’s fine. But don’t package “avoiding distillation” as moral superiority, and don’t make the entire industry pay for this “distillation shame.”
Technological competition is about who runs faster and goes farther, not who is more “pure.” In this sense, remaining open to all means that advance AI is the true path to long-term success.
After all, what AI must learn is not just the answers themselves but also the openness with which humanity has “stood on the shoulders of giants” for thousands of years.
Disclaimer: This article is written based on publicly available information or information provided by interviewees. However, Keji Sishao and the author do not guarantee the completeness or accuracy of the information or references mentioned or displayed in the article and do not represent the stance of any institution or individual. If infringement is involved, please contact us for deletion. Under no circumstances shall the information or opinions expressed in this article constitute investment advice to any person.