The White House just accused a Chinese lab of stealing an American AI model
Michael Kratsios, director of the White House Office of Science and Technology Policy, posted on X on July 22 that Moonshot AI distilled Anthropic’s Fable 5 model to build Kimi K3. He called it “large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research.”
This is the first time a senior US official has publicly named a specific Chinese lab and accused it of copying a specific American model.
The target is not random. Kimi K3 launched on July 16 as a 2.8-trillion-parameter open-weight model. It took first place on the Frontend Code Arena with a 76 percent win rate over Claude Fable 5. It became the largest open-weight release in history. Its weights are scheduled to go free on July 27.
An accusation that this model was built by copying the very model it beat reframes the entire narrative.
The evidence
Kratsios said Moonshot “developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection.” He did not provide details on how the US government learned this.
The technical case that has circulated publicly rests on two things. First, Kimi K3 was documented identifying itself as Claude in at least one shared conversation. Second, Ryan Greenblatt, chief scientist at Redwood Research, published a cross-entropy analysis showing K3 claims to be Claude disproportionately often across many prompts, in a pattern difficult to explain as random noise.
A single screenshot of a model misidentifying itself proves little. Models do that routinely. A systematic statistical pattern across many prompts is a different kind of claim. Cross-entropy comparison asks how surprised each model is by particular text, and consistent patterns can indicate shared training lineage.
The honest caveat: a model calling itself Claude has innocent explanations. Claude transcripts are all over the web. Any model trained on broad web data ingests some of them. Training data contamination, leftover system prompts, or copied examples from public datasets could all produce the same signal. The analysis is suggestive. It is not conclusive.
The second accusation: banned Nvidia chips through Thailand
Kratsios also said Moonshot “acquired GB300-equipped servers and has accessed GB300s in Thailand, likely to train its AI models.” Nvidia’s GB300 is among the most advanced AI accelerators available and is subject to US export restrictions limiting sales to Chinese entities.
This is the more legally concrete of the two accusations. Distillation sits in a murky area of contract law and terms of service. Export control circumvention is a specific regulatory offense with established enforcement mechanisms. Penalties can reach suppliers, intermediaries, and cloud providers, not just the end user.
It also answers a question observers raised when K3 launched. Training a 2.8-trillion-parameter model requires enormous compute. How a Chinese lab assembled that under export restrictions was never fully clear. Routing restricted hardware through a third country is the obvious workaround.
A chip can be legally sold to a permitted country, installed in a data center there, and rented to anyone with a credit card. That converts an export control question into a cloud access question that current policy handles poorly. Southeast Asia has attracted substantial data center investment precisely because it sits outside the tightest restrictions.
Why this matters beyond the accusation itself
Kratsios drew a line between legitimate and illegitimate distillation. The US supports distillation used to create smaller, more efficient models. What it objects to is large-scale covert extraction of a competitor’s full capabilities. That framing matters because it signals the administration could pursue tighter API restrictions and expanded export controls rather than just diplomatic complaints.
The timing is deliberate. Kimi K3’s open weights arrive July 27, four days after the accusation. Every organization planning to download and self-host K3 now has to weigh a contested provenance claim alongside their technical evaluation. Corporate legal teams at enterprises with government contracts or strict IP compliance requirements will flag this. Startups optimizing for cost will likely not care.
Expect adoption to bifurcate. Individual developers and startups move fast. Large enterprises wait for clarity.
This is the second distillation allegation involving Anthropic this year. The company previously accused Alibaba of using 25,000 fraudulent accounts to run 28.8 million interactions on Claude over six weeks. In April, the House Homeland Security Committee and the Select Committee on China announced a joint investigation into Chinese AI models and what they called “the large-scale theft of proprietary capabilities from American frontier AI systems through adversarial distillation.”
Moonshot AI did not respond to requests for comment from multiple outlets. The company’s own website describes Kimi K3 as the first open 2.8-trillion-parameter model and notes that “while its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite.”
Proving distillation is genuinely hard. Model weights do not carry watermarks. A distilled model’s parameters look nothing like its teacher’s. Investigators rely on behavioral forensics: does the student reproduce the teacher’s specific quirks, refusal patterns, formatting habits, or identity confusions at rates that random chance cannot explain? Those signals are circumstantial. They can be produced innocently by training on web data that contains the teacher’s outputs.
The industry has been converging on defenses rather than proofs. OpenAI, Anthropic, and Google began sharing intelligence earlier this year to detect coordinated distillation attempts. Rate limiting and query pattern detection have tightened. Terms of service now explicitly prohibit training competing models on outputs.
None of that resolves the underlying problem. The technique works. The evidence is inherently ambiguous. And the geopolitical stakes just got raised from benchmarks to accusations of theft.
Sources: CyberScoop, The New Stack, Bloomberg Law, Australian Financial Review, TipRanks/Cointelegraph, BuildFastWithAI daily digest