AI distillation is the subject of a joint advisory published on 8 September 2026 by the National Security Agency, the Cybersecurity and Infrastructure Security Agency and the Federal Bureau of Investigation. Catalogued as AA26-251A, it names six Chinese companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI — and accuses them of extracting “billions of tokens across millions of exchanges/requests” from American frontier models since at least late 2024.
It is an unusually specific document. It lists 41 distinct American models by version. It maps the behaviour to ten MITRE ATLAS techniques. It describes four tactics it says the framework does not yet cover. And across 3,585 words of substantive text, it never once uses the words theft, stolen, illegal, unlawful, copyright, trade secret, lawsuit or sanction.
What it says instead, exactly once, is that the six companies are “violating U.S. AI companies’ terms of use.” That is the entire legal characterisation of a campaign three federal agencies describe as an industrial-scale threat to national technological leadership. This piece counts what the AI distillation advisory says, what it declines to say, and where its own two lists of accused models disagree with each other.
Every figure below comes from the AI distillation advisory as published at cisa.gov, counted this session. The technique itself is a legitimate and well-documented one — knowledge distillation has an academic literature going back a decade, and the AI distillation advisory says so in its second sentence.
Table of contents
- What the AI Distillation Advisory Actually Alleges
- The AI Distillation Advisory’s Missing Vocabulary
- Counting the AI Distillation Advisory: 3,585 Words
- The Two Model Lists That Do Not Match
- Which American Models the AI Distillation Advisory Names
- The Attribution Section With No Confidence Statement
- What the AI Distillation Advisory Asks Companies to Do
- The Tactics the AI Distillation Advisory Describes
- The Sources Behind the AI Distillation Advisory
- What the Accused Companies and Beijing Said
- What AI Distillation Means for Everyone Else
- How to Read an AI Distillation Claim
- What to Watch Next on AI Distillation
- Frequently Asked Questions About AI Distillation
- References and Further Reading
What the AI Distillation Advisory Actually Alleges
The AI distillation advisory is a joint product with three seals on it, which is itself a signal: NSA, CISA and FBI co-sealing a single document is reserved for claims the agencies want read as consensus rather than as one agency’s assessment.
The three agencies and the six companies
The six named firms are given with their full registered corporate names — DeepSeek Artificial Intelligence Technology Research Co., Ltd.; Beijing Moonshot Technology Co., Ltd.; Alibaba Group; Shanghai MiniMax Co., Ltd.; Shanghai Jieyue Xingchen Intelligence Technology Co., Ltd.; and Z.AI. Naming registered entities rather than brands is the convention for a document intended to support later action.
The six are not mentioned equally. DeepSeek and MiniMax are each named 13 times, Moonshot AI 11, StepFun 8, Z.AI 7 and Alibaba 6 — despite Alibaba being by far the largest company on the list.
What “industrial-scale” means in the AI distillation advisory
The phrase “industrial-scale” appears seven times, and the AI distillation advisory defines it by contrast rather than by measurement. Distillation, it says, is “not a supplement to these companies’ AI model development, but the critical core of it.” The scale claim rests on that structural argument, not on a published volume.
The quantities given are “billions of tokens,” “millions of exchanges/requests” and query volumes “in the thousands to millions per domain.” No company-specific number appears anywhere in the AI distillation advisory.
The one sentence that names a violation
Here is the sentence in full: “China-based AI companies route distillation requests through multiple pathways to gain unauthorized access, consequently violating U.S. AI companies’ terms of use.”
“Unauthorized” appears once in the AI distillation advisory. “Violating” appears once. Both are in that sentence. Everything else the AI distillation advisory alleges — the proxies, the bulk subscriptions, the prompt injection, the metadata scrubbing — is described as a technique, not as an offence.
The AI Distillation Advisory's Missing Vocabulary
The absence is easier to see as a list than as an argument. These are terms a reader might reasonably expect in a federal accusation of large-scale appropriation, with their actual frequency in the AI distillation advisory’s 3,585 substantive words.
Sixteen words of the law that never appear
| Term | Uses | What it would have established |
|---|---|---|
| theft / steal / stolen | 0 | That something was taken, not copied |
| illegal / unlawful | 0 | That a law was broken |
| crime / criminal | 0 | That prosecution is contemplated |
| copyright / patent | 0 | Which property right is at stake |
| trade secret | 0 | The usual US frame for this claim |
| lawsuit / court | 0 | That a forum exists for the claim |
| sanction / export control | 0 | That a policy response follows |
| penalty / fine | 0 | That there is a cost to the conduct |
| espionage / hacking | 0 | That this is an intrusion, not a purchase |
| evidence | 0 | What the findings rest on |
| terms of use | 3 | The only thing the advisory says was breached |
The four occurrences of “fine” in the document are all inside “fine-tuning,” “refinement” and “define a finite privacy budget.” None is a monetary penalty.
What it says instead: terms of use, three times
A terms-of-use breach is a contract matter between a provider and a customer. It is enforced by suspending an account, not by an agency. The three federal agencies that wrote this advisory have no role in enforcing one.
That is not a criticism of the document so much as a description of what it is. It is a threat-intelligence product written for defenders, and it stays inside that lane with real discipline.
Why the word choice is not an accident
The contrast with the political framing around it is sharp. In July, the director of the White House Office of Science and Technology Policy, Michael Kratsios, wrote on X that large-scale covert industrial distillation was “aimed at stealing proprietary U.S. technology.” The advisory that followed his April memorandum does not use that word, or any word like it, once in 3,585 words.
Counting the AI Distillation Advisory: 3,585 Words
The count below covers the advisory’s substantive body — from the executive summary through the end of the mitigations — and excludes the footnotes, reference list, disclaimer, contact block and page tags.
The term-frequency spine
The AI distillation advisory uses some form of “distill” once every 65 words. It uses the word describing what was allegedly broken once every 1,195 words.
One dollar figure, and it is disowned
There is exactly one dollar amount in the entire advisory: DeepSeek’s “publicly quoted training costs of $5.6M.” It appears so the AI distillation advisory can say the figure is “misleading as it does not include the true cost of the data acquired through extensive malicious distillation,” and it is footnoted to the DeepSeek-V3 Technical Report.
So the only number with a currency sign attached is one the authors introduce in order to reject. No replacement estimate is offered. The AI distillation advisory claims capabilities “worth billions in development costs” were extracted and never puts a figure on that either.
Every date is a season
There is no precise date anywhere in the attribution. DeepSeek’s campaign runs “between late 2024 and mid-2025.” Moonshot AI’s begins “at least mid-2025.” Alibaba and MiniMax act “in late 2025,” StepFun “between late 2025 and early 2026,” Z.AI “by mid-2026.”
Nine date references, none narrower than a third of a year. For a document that names specific model versions down to GPT-5.1 Codex Mini, the temporal precision is conspicuously coarse.
The Two Model Lists That Do Not Match
The AI distillation advisory gives its model attributions twice: once in narrative prose, company by company, and once in Table 1. The two lists are not the same list.
The prose names 51, the table names 53
Counting every model mention in each version, the narrative makes 51 and the table makes 53. Deduplicated, the prose names 38 distinct models and the table names 40. The union of both is 41 — which means three models appear in only one of the two places.
For five of the six companies, the two lists disagree. Only Z.AI’s row matches its prose exactly, and Z.AI is the shortest entry in the AI distillation advisory, with two models.
GPT-4o vanishes from Moonshot’s row
The most consequential mismatch is Moonshot AI’s. The narrative states plainly that “Moonshot AI extracted significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data to train its Kimi-K2 model.”
Table 1’s Moonshot row lists 17 models. GPT-4o is not one of them. It lists “GPT-4o mini,” which is a different model. The single American model the AI distillation advisory’s own prose identifies as the training source for Kimi K2 is missing from the table that exists to enumerate exactly that.
Four more disagreements
| Company | Prose | Table 1 | The discrepancy |
|---|---|---|---|
| DeepSeek | 12 | 14 | Table adds Gemini 2 and Grok 3 Mini; prose writes “Claude 3.7”, table “Claude Sonnet 3.7” |
| Moonshot AI | 18 | 17 | Table drops GPT-4o, the stated source of Kimi K2 |
| Alibaba | 4 | 3 | Table drops Claude Opus |
| MiniMax | 6 | 7 | Table adds GPT-5 and pins “Claude Opus” to 4.5 |
| StepFun | 9 | 10 | Table adds GPT-5.1 Codex Mini |
| Z.AI | 2 | 2 | The only exact match |
None of these is fatal to the advisory’s case, and the likeliest explanation is ordinary drafting drift between a narrative section and a summary table compiled at different times. But the table is the artefact defenders will actually load into a spreadsheet, and it is the one that is wrong about Kimi K2.
Which American Models the AI Distillation Advisory Names
The document is framed throughout as a defence of “U.S. frontier AI models” in general. Its contents are less evenly distributed than that framing suggests.
Claude 45 mentions, GPT 41, Gemini 18, Grok 6
Anthropic and OpenAI account for 86 of the 112 model-family mentions in the body. In Table 1 the two are exactly level, with 20 row entries each out of 53.
Every accused company touched Claude and GPT
All six companies have at least one Claude model and at least one GPT model in their Table 1 row. Only three of the six have a Gemini entry, and only two have anything from xAI.
Anthropic’s models are also the ones tied to the most specific allegations. Moonshot’s Kimi K3 is attributed to Claude Fable 5. MiniMax is said to have used Claude Code for its own internal software development, and to have “used prompt injections to try to trick Claude Code into believing it was a MiniMax product.”
Google and xAI appear in half as many rows
The imbalance may reflect real targeting, or it may reflect which companies supplied the most telemetry. The AI distillation advisory does not say which, and it never explains how any of the attributions were derived.
The Attribution Section With No Confidence Statement
The AI distillation advisory contains a section headed “Attribution.” It runs 1,008 words and contains no confidence language of any kind.
“Likely,” twice, carrying the whole government claim
The link to the Chinese state rests entirely on two sentences, and both use the same hedge. The executive summary says the extraction happened “Likely with Chinese government awareness.” The attribution section says “Likely with the knowledge of the Chinese government.”
Those two instances of “likely” are the only hedged qualifiers in the AI distillation advisory. There is no “we assess with high confidence,” no confidence scale, and no footnote explaining what “likely” means here. NBC News, reporting the AI distillation advisory, noted that it “did not claim that China’s intelligence played a role.”
“Evidence” appears zero times
The word “evidence” is absent. So is “assess” in the intelligence sense. “Confidence” appears twice, and neither instance is an attribution statement: one refers to “high-confidence malicious distillation requests” that a company might detect, and the other is the technical term “logits/confidences.”
For a public advisory this is defensible — sources and methods are exactly what such a document withholds. It is still worth stating plainly that readers are asked to accept a detailed set of factual claims on the agencies’ authority alone.
What an intelligence product usually says here
Joint advisories in this series more often carry an explicit basis line — incident response engagements, industry reporting, a named partner’s telemetry. This one attributes nothing to a source. The only external reporting it leans on appears in the reference list rather than in the text.
What the AI Distillation Advisory Asks Companies to Do
The remedies occupy 1,041 words. The accusation and technique sections together occupy 2,544 — a ratio of 2.44 to 1.
Ten MITRE ATLAS mitigations
The AI distillation advisory maps the alleged behaviour to ten MITRE ATLAS techniques and tactics, then lists ten distinct ATLAS mitigations by identifier, from adversarial input detection through model ensembles. It adds four “novel TTPs” it says ATLAS does not yet cover: regional restriction evasion, centralised request routing, automated metadata sanitisation and systematic quota optimisation.
Each novel technique comes with detection indicators — shared accounts across multiple IPs, sustained 24/7 usage with no idle periods, new subscriptions immediately at maximum usage rather than a gradual ramp.
The recommendation to degrade responses quietly
The most striking recommendation is that providers should silently return worse answers to suspected distillers. The AI distillation advisory suggests “reducing reasoning depth, presenting correct information with different reasoning, or stylistic inconsistencies,” varied across requests so that quality-evaluation pipelines cannot detect the change.
It then says so explicitly: “Avoid informing China-based AI company users suspected of distillation campaigns of a switch to a downgraded model.” The stated reason is that telling them “would enable them to improve their defense evasions and indicate when to roll back training.”
The carve-out for safety researchers
One group is exempted. “In contrast, AI safety researchers and third-party evaluators should be informed of model changes while still applying strong distillation mitigations.”
That single sentence concedes the obvious hazard: a policy of undisclosed, deliberately inconsistent model behaviour would otherwise corrupt every external evaluation of the model. The carve-out is the right instinct, and it is also an admission that the recommendation degrades the product for anyone caught by a false positive.
Zero actions for the government that wrote it
Every recommendation in the mitigations section is an action for a private company. The phrase “the U.S. Government” appears once in the whole advisory, in a sentence about coordination, and not in the mitigations at all. There is no commitment by NSA, CISA or FBI to do anything — no enforcement, no funding, no rule, no follow-up product.
Three agencies published a document describing a strategic threat to national technological leadership, and assigned the entire response to industry.
The Tactics the AI Distillation Advisory Describes
Set the framing aside and the technical content is genuinely useful. It is the most detailed public description of how model extraction is actually operated at scale.
Transfer stations and the gray market
The AI distillation advisory describes a “gray market of API proxies known as ‘transfer stations'” that resell frontier-model access “at a fraction of the official price,” letting buyers bypass regional restrictions and stripping the metadata that would otherwise identify them.
Three mentions of “gray market,” three of “transfer station.” The claim is sourced in the reference list to The Decoder, a trade publication, rather than to agency telemetry.
Prompt injection to pull out chain-of-thought
The most technically interesting allegation concerns reasoning traces. DeepSeek is said to have used prompts “instructing models to imagine and articulate the internal reasoning behind completed responses and write it out step by step,” defeating providers’ restrictions on chain-of-thought visibility.
The AI distillation advisory’s point is that this data teaches “not just factual knowledge, but reasoning methodologies” — which is why hidden reasoning is commercially protected in the first place.
MiniMax’s 24-hour retarget
One operational detail stands out: MiniMax “redirected exchanges to a new Claude model within 24 hours of release.” The AI distillation advisory offers it as evidence of real-time provider monitoring and pre-positioned infrastructure.
It is the single most specific number in the attribution, and the only one measured in hours rather than seasons.
The Sources Behind the AI Distillation Advisory
The reference list is short and revealing. Eight entries, of which three are White House documents and one is a trade blog.
Eight references, three from the victims
Three of the four American companies whose models are named contributed published research to the reference list: Anthropic’s write-up on detecting and preventing distillation attacks, Google’s GTIG threat tracker entry on distillation, and an OpenAI policy letter. NIST’s adversarial machine learning taxonomy supplies the technical mitigations.
That is a reasonable evidentiary base for a defensive document. It also means the AI distillation advisory’s picture of who is being targeted is shaped by which companies chose to publish.
xAI is named but contributes nothing
Grok appears six times in the body and in two of the six Table 1 rows. xAI contributes no reference, no research and no statement. It is the one named victim that is entirely silent in the AI distillation advisory written on its behalf.
A trade blog carries the gray-market claim
The “transfer stations” economics — the claim that frontier tokens are resold at a fraction of list price — is the AI distillation advisory’s most checkable market assertion, and its cited source is a post on The Decoder. It is a credible outlet, and it is not agency telemetry.
What the Accused Companies and Beijing Said
None of the six companies is quoted in the AI distillation advisory, and none had responded publicly at the time of the reporting that followed it.
China’s foreign ministry response
Chinese foreign ministry spokesperson Mao Ning, asked about the report at a regular briefing in Beijing, said she had not seen it but that China’s AI development “is a result of high-level scientific and technological self-reliance.” She added that the United States should be strengthening AI cooperation with China “rather than making groundless accusations.”
The exchange lands just over two weeks before a scheduled meeting between President Trump and Xi Jinping on 24 September.
The distillation the AI distillation advisory does not mention
The advisory is scoped to China-based companies, and that scope is legitimate. It is worth noting what falls outside it. Engadget’s report on the advisory points out that in the same period, Elon Musk acknowledged under cross-examination that xAI had used OpenAI output to train its own models — the same technique, by an American company, against an American model.
OpenAI has also said it banned accounts it suspected of distilling its technology, DeepSeek among them, as far back as early 2025. The practice the advisory describes is not confined to the six firms it names, and the document does not claim otherwise.
What AI Distillation Means for Everyone Else
For most organisations this advisory is not an action item. It is worth understanding for two reasons, and worth over-reading for none.
If you build on these APIs
The behavioural detection recommendations describe patterns some legitimate users genuinely exhibit. Sustained 24/7 automated usage, new accounts at immediate maximum throughput, high-volume similar prompts and pooled team credentials are all normal for a batch pipeline or an evaluation harness.
If providers implement the response-degradation advice, a false positive now means silently worse output rather than an error message. That is a new and largely undetectable failure mode, and the only practical defence is the ordinary hygiene of named accounts, gradual ramp-up and documented usage patterns. It belongs in the same review as the rest of your AI strategy and vendor governance.
If you deploy Chinese open-weight models
Several of the named companies publish widely used open-weight models. The advisory makes no claim about the safety, quality or licensing of those weights, and it recommends nothing about running them.
What it does establish is a documented federal position on their provenance. Organisations with procurement policies that reference government threat reporting should expect that position to surface in due-diligence questionnaires, and should read it accurately: an allegation of terms-of-use violation, not a finding of illegality. Sound data management practice means recording which models touched which data regardless.
The compliance question this raises
Because the advisory names no law, it creates no compliance obligation. Nothing in it requires any action by any organisation. Treat it as threat intelligence — useful for understanding an adversary’s tradecraft, not a control to implement. Reinforcement learning pipelines and evaluation harnesses that call third-party APIs are worth a look purely because of the detection indicators, not because of any duty created here.
How to Read an AI Distillation Claim
The same questions work on any accusation of model copying, whether it comes from a government, a vendor or a competitor.
Four questions that separate a claim from a finding
First, what is alleged to have been broken — a law, a contract, or a norm? This advisory says a contract, and says it once.
Second, what is the basis? A named engagement, partner telemetry, or authority alone. Third, what is the confidence, and is it stated? Fourth, are the document’s own artefacts internally consistent — do the narrative and the tables agree?
Where the numbers usually hide
The most useful figures in a document like this are rarely in the headline. They are in the tables, the footnotes and the counts you have to do yourself. This advisory’s most specific claim — MiniMax retargeting within 24 hours — sits in a table cell in the middle of the document.
Check the summary table against the prose
Summary tables are compiled last and reviewed least. When a document gives the same facts twice, comparing the two versions is the cheapest available audit, and here it found three models that appear in only one of the two lists.
What to Watch Next on AI Distillation
Three things would change how this document reads, and all three are observable.
Whether any provider confirms response degradation
No American AI company has said it degrades responses to suspected distillers. If one confirms the practice, every published benchmark for that model acquires an asterisk. If none does, the recommendation stays theoretical.
Whether a legal or policy action follows
The advisory names no law and threatens nothing. An export-control action, a Commerce listing or a civil suit citing AA26-251A would retroactively make this document the opening move rather than the whole move.
Whether the agencies publish a corrected table
The GPT-4o omission is the kind of error that gets quietly fixed. Whether it is, and whether the revision is flagged, is a reasonable test of how the advisory is being maintained.
Frequently Asked Questions About AI Distillation
What is AI distillation?
It is the practice of training a smaller or newer model on the outputs of a more capable one. It is a standard, published technique in machine learning research, and the advisory says so explicitly before drawing its distinction between legitimate and “malicious” use.
Does the advisory say the six companies broke the law?
No. In 3,585 words it never uses the words illegal, unlawful, theft, crime, copyright or trade secret. The single violation it names is of American AI companies’ terms of use.
Which companies are named?
DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI, each given with its full registered corporate name.
How many American models does it list?
41 distinct models across both the narrative and Table 1 — 38 in the prose, 40 in the table. Claude and GPT models appear in all six companies’ rows.
Is there proof the Chinese government was involved?
The advisory says the campaigns happened “likely with Chinese government awareness” and “likely with the knowledge of the Chinese government.” Those two hedged phrases are the entire basis given, and the word “evidence” does not appear.
Does this require my organisation to do anything?
No. It is advisory threat reporting with no regulatory force. Its detection indicators are useful if you operate high-volume API pipelines; otherwise it creates no obligation.
References and Further Reading
US accuses Chinese AI developers including DeepSeek and Alibaba of copying American AI — NBC News
Detecting and preventing distillation attacks — Anthropic
MITRE ATLAS — Adversarial Threat Landscape for Artificial-Intelligence Systems
How China’s gray market sells Claude tokens at a fraction of the price — The Decoder
DeepSeek-V3 Technical Report — arXiv
Knowledge distillation — Wikipedia
AI Safety Talks: Essential Facts on the US-China Risk — Progressive Robot
US Government Sides With OpenAI on Training LLMs on Copyrighted Material — Progressive Robot
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.