The US Accuses China of 'Distillation.' Mathematicians Echo the Question Back at OpenAI in the Same Week

09/14 2026 490

In the second week of September, three significant events unfolded. On September 8, the US National Security Agency (NSA), Federal Bureau of Investigation (FBI), and Cybersecurity and Infrastructure Security Agency (CISA) accused six Chinese AI firms of 'industrial-scale' distillation of American cutting-edge models. Almost simultaneously, OpenAI announced it had resolved a mathematical conundrum that had baffled experts for nearly a century, only to become embroiled in a controversy over data sourcing and attribution. On September 11, mathematicians including Terence Tao and Deng Yu issued a joint statement highlighting a 'serious misalignment' between AI companies and the mathematical community. The same issue, approached from two different directions. The crux of the matter is not whether distillation constitutes copying, but rather who holds the authority to define originality, assess risk, and set the rules.

On September 8, the NSA, FBI, and CISA jointly issued a notice, accusing six Chinese AI companies of 'industrial-scale' distillation of American cutting-edge models. That same week, OpenAI announced a major breakthrough in studying singularities in the Navier-Stokes equations. However, Tristan Buckmaster, a researcher closely involved, publicly stated that he had uploaded an unpublished draft to a programming tool under OpenAI's umbrella three weeks earlier. On September 11, 25 mathematicians issued a joint statement warning of a 'serious misalignment' between AI companies and the mathematical community.

Image Source: Internet

The same issue, two directions. The US accuses China of 'distillation,' while mathematicians question OpenAI: 'Where did the data come from?'

The real question is no longer whether distillation is tantamount to copying, but rather: Who has the authority to label someone else's learning as theft and their own as innovation?

01 The US Notice: Legalizing Distillation Yet Accusing China of 'Malicious Intent'

The notice from the three US agencies, numbered AA26-251A, named six Chinese AI companies, accusing them of extracting 'billions of tokens' of data from models like Claude, GPT, Gemini, and Grok through millions of requests.

Yet, the notice itself concedes that 'knowledge distillation' is widely recognized as a legitimate and practical technique in AI research.

Image Source: Internet

The question arises: If distillation is legal, what defines 'industrial scale'? The notice provides extensive interaction data—approximately 200 million interactions, with one company accounting for 151 million—but fails to define a clear threshold for 'industrial scale.' Data alone does not constitute a standard. The final characterization remains a political judgment: Chinese companies' distillation is 'large-scale, targeted, and malicious.'

The Ministry of Commerce responded directly, stating, 'Distillation is a common practice in the AI field for models to learn from one another. It is essentially a neutral technical means used by global model companies, including those in the US.'

In January 2025, Microsoft announced the availability of an open-source model from a Chinese AI company and previewed a distilled version. This was a public commercial action, not a gray area.

If distillation poses a security threat, why can US cloud service providers use it openly while Chinese companies are accused of 'industrial-scale attacks'?

The practical effect of this notice is to shift distillation from a technical issue into a national security framework. Liu Dian, an associate researcher at the China Institute of Fudan University, said, 'The US is elevating model distillation from an issue of intellectual property and commercial competition between companies to one of national security and geopolitical technological competition.'

02 OpenAI Faces Its Own Data Provenance Issue in the Same Week

The same week the US notice was issued, OpenAI announced a major breakthrough by its model regarding singularities in the Navier-Stokes equations. This should have been a highlight for AI. However, mathematician Tristan Buckmaster from New York University released a public statement the day before, revealing a timeline that cast a shadow over this achievement:

In mid-August, while working with OpenAI's programming tool Codex, Buckmaster achieved a breakthrough on a key step of the problem and saved an unpublished draft file in his web directory for the tool. On September 3, he emailed mathematicians at OpenAI, asking whether they had used his draft. On September 6, he had two calls with OpenAI representatives, stating he 'did not receive an answer.' On September 7, he issued a public statement. On September 8, OpenAI formally announced its results.

Notably, Buckmaster explicitly stated in his announcement that he was not making any accusations against anyone nor claiming the results as his own. His question was singular: Where did the data come from?

Image Source: Internet

OpenAI's response was to acknowledge that it could not rule out the possibility that 'de-identified data participated in training,' while stating that neither the research nor the agent had seen his work in any way or accessed specific user data before its public release.

'Not directly accessed' and 'not participated in training' are two different issues. Buckmaster asked whether his draft was 'used in training,' and the response he received was 'no direct access'—a non-answer. This mirrors almost exactly the experience of another mathematician.

Image Source: Internet

03 Mathematicians' Counterattack: The Truly Powerful Part

Andreas Thom, who has long studied non-sofic group problems, accused OpenAI of providing false responses regarding the use of his conversational data.

The key issue is not the qualitative (characterization) of 'plagiarism,' but rather that Thom provided a specific, verifiable chain of evidence: In the months leading up to OpenAI's official announcement, Thom had been frequently using ChatGPT to discuss his unpublished research and had entered substantial unpublished core derivations and proof details into the chat.

He sent a query email to OpenAI researchers, posing two extremely precise questions: First, were the conversation records of our discussions about unpublished results in ChatGPT over the past few months included in the model's training data? Second, during the model's problem-solving and reasoning process, can these records be directly retrieved and accessed?

OpenAI researcher Mark Sellke replied with a single sentence: 'Regarding your conversations with ChatGPT: That did not happen.'

Image Source: Internet

This statement evaded the core issue. Sellke's response only addressed 'direct access' and provided no qualification, explanation, or evidence regarding the 'training data' question. Thom subsequently pointed out that users 'can see a setting but cannot see the backend data flow.'

The burden of proof shifts when the accuser must prove copying, while the accused need only deny it; evidence is held by the accused.

Thom only had email records, without code audit access or training data access. OpenAI needed only to say, 'That did not happen,' shifting the burden of proof onto the mathematicians. To prove their work was used, mathematicians would need access to OpenAI's training data and internal model logs—which OpenAI is not obligated to provide.

The joint statement by 25 mathematicians elevated this specific dispute to a more fundamental level. The statement noted that AI companies are using 'solving mathematical problems' as a benchmark for model capability, yet 'the goals of AI companies and the mathematical community are seriously misaligned.' The core goal of mathematical research is to form conceptual understandings and new insights, not merely to provide 'correct or incorrect' answers. Producing 'true/false' conclusions at increasing speeds 'may not nurture new ideas but could instead destroy the fertile ground where ideas are cultivated.'

Image Source: Internet

What truly carries weight in this statement is the desire to redefine 'what counts as original.'

If solving a mathematical problem requires decades of discussion, simplification, and dissemination before it becomes knowledge the academic community can understand and use—then does an 'answer' provided by an AI in a few hours truly 'solve' the problem? If a result has not undergone verification and digestion by the academic community, can it still be called 'original'?

The mathematicians' answer is: No.

04 The Politicization of Distillation: Who Defines 'Malicious Intent'?

Returning to the opening question: Who decides about distillation?

At the technical level, the US notice itself admits it is legitimate and practical, the Ministry of Commerce calls it a common practice, and Microsoft uses it openly. At the legal level, knowledge distillation is not explicitly prohibited, but its interpretive space is being narrowed.

At the political level, who defines 'malicious intent'? The US designates Chinese companies as 'industrial-scale' and 'malicious,' but US AI companies are also distilling information from all users—are they hoping to fight a double standard? Or use this as a weapon?

This is a struggle over who gets to define the rules. The US notice pulls interpretive authority away from the technical community and toward security agencies. The mathematicians' statement seeks to pull the standard for originality back from 'who published first' to 'who was first understood and verified by the community.'

Conclusion: Originality Is Not a Factual Judgment, But a Power Judgment

The fact truly established this week is not who copied whom, but a more uncomfortable reality: When the accused party of 'distillation' demands access to training data, no mechanism exists to compel the accused to produce evidence.

Its value lies not in proving OpenAI copied but in exposing the fragility of 'originality': When technical similarities cannot be legally defined as copying, when conceptual overlaps cannot be traced to their source, and when the burden of proof rests on the accuser, 'originality' ceases to be a factual judgment and becomes a power judgment.

Image Source: Internet

The mathematicians offer a way forward: Originality should not be determined by 'who published first' but by 'who was first understood and verified by the community.'

This standard is uncomfortable for both OpenAI and its accusers. But for researchers who truly respect data provenance and are willing to undergo community verification, it may be the best news.

Disclaimer: The accusations and controversies discussed in this article are based on public statements by the parties involved or media reports. The mathematicians' accusations regarding training data usage are a subject of dispute between the parties; OpenAI denies direct access to the relevant drafts and conversations but acknowledges it cannot rule out the participation of de-identified data in training. This article aims to present the sequence of events and industry impact and does not constitute legal or investment advice.

- END -

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.