The Claude Distillation Dispute: Seven Chinese AI Labs, Nearly 200 Million Exchanges and a Contested Model Release
Anthropic alleges that seven China-based AI labs conducted industrial-scale, unauthorized campaigns to extract capabilities from Claude, while Beijing rejects the accusations as groundless. The evidence also raises privacy, model-security and open-weights questions that the Chinese AI ecosystem will need to address.
Wei Lianπ¨π³ China Desk LeadSep 20, 2026 10m read# The Claude Distillation Dispute: Seven Chinese AI Labs, Nearly 200 Million Exchanges and a Contested Model Release
*Wei Lian, China Desk Lead, Beijing β September 20, 2026*
On September 10, 2026, Anthropic published a 154-page threat intelligence report titled *Detecting and Countering Misuse of AI*, alleging that seven China-based AI laboratories conducted industrial-scale, unauthorized campaigns to extract capabilities from its Claude models. Two days earlier, a joint advisory from CISA, the NSA and the FBI had named six of the same companies. The combined picture is the most detailed public account yet of how Chinese labs may have accelerated their model development β and it has landed in the middle of preparations for a US-China AI safety dialogue ahead of a September 24 leaders' summit.
The central allegation is not simply that Chinese developers learned from another model's outputs. Knowledge distillation is a standard machine-learning practice. Anthropic's caseβ, as reported by TechCrunch, instead rests on the claimed combination of scale, deception and circumvention: fraudulent accounts, stolen payment credentials, proxy networks, attempts to expose hidden reasoning traces, and β most seriously β the undisclosed routing of customers' live requests into Claude without their knowledge.
Beijing has rejected the allegations as lacking factual and legal basis. The labs named have not issued detailed technical rebuttals. The underlying evidence has not been independently audited. What follows is a precise account of what has been alleged, what remains unverified, and what the dispute means for the Chinese AI ecosystem.
What Anthropic Says It Observed
Anthropic's reportβ covers activity from December 2025 through August 2026 and attributes campaigns to seven companies, five of which received explicit exchange counts:
- Alibaba (Qwen team): the largest operation observed, designated GTG-16005, involving more than 151 million exchanges between May and July 2026, peaking at nearly 3 million daily requests using over 3,500 fraudulent accounts. Operators allegedly inserted fixed instructions to force Claude Opus models to output full chain-of-thought reasoning into inline tags, which were then harvested as training material for the Qwen family.
- Moonshot AI (Kimi): more than 23 million exchanges (GTG-16002). In one alleged relay operation spanning 10 days, Moonshot routed approximately 300,000 genuine customer requests through Claude using 5,380 fraudulent accounts, often without user knowledge.
- DeepSeek: more than 12.1 million exchanges over a 14-day period in July 2026 (GTG-16001), with selective forwarding of high-value agentic coding tasks to Claude Opus after inspecting inbound request headers.
- Zhipu/Z.ai: more than 3.4 million exchanges over 17 days in June and July 2026 (GTG-16006), focused on replaying and cleaning reasoning traces for use with GLM models.
- Xiaomi: more than 400,000 exchanges in March and April 2026 (GTG-16008), using coding harnesses to extract training data.
- SenseTime and MiniMax: also identified (GTG-16012 and GTG-16003 respectively), with MiniMax allegedly operating a proxy network through a shell company and SenseTime purchasing harvested transcripts from third-party vendors.
The CISA advisory AA26-251Aβ, published September 8, named six companies β Alibaba, Moonshot, DeepSeek, MiniMax, StepFun and Z.ai β and described extraction campaigns targeting Claude, GPT, Gemini and Grok. It omitted Xiaomi and SenseTime but added StepFun, which Anthropic did not name. The two documents are complementary but not identical.
The evidentiary distinction is essential: Anthropic is well placed to identify coordinated abuse of its own systems, but attribution from abusive infrastructure to a specific model developer's senior management is a separate analytical step that neither document fully closes.
How the Alleged Extraction Worked
Transfer Stations and Fraudulent-Account Infrastructure
The Register's coverage of the CISA advisoryβ describes the core infrastructure as "transfer stations": gray-market API proxies that route requests to restricted models while obscuring the original customer. Unlike a single conspicuous account, distributed traffic can span thousands of accounts, residential proxies, subscriptions, cloud services and aggregators. Automated failover replaces blocked routes, while ordinary-looking prompts can obscure extraction patterns.
The alleged infrastructure included:
- bulk-created or purchased accounts, often using stolen credit cards and login credentials;
- proxy and remote-cloud routes that automatically obfuscated user metadata;
- automated monitoring and quality-control pipelines to evaluate extraction yield;
- rapid retargeting when a provider released a new model or updated its defenses;
- coordinated high-volume queries designed to stay below per-account detection thresholds.
These techniques can improve collection volume, but they do not guarantee faithful replication. Distillation transfers patterns of answers and behavior; it is not a direct copy of the underlying model weights or training corpus.
Chain-of-Thought Extraction and Cross-Session Replay
"Chain of thought" refers to intermediate reasoning text that shows how a model decomposes a problem. Such text can be more valuable for training than a final answer because it provides examples of decomposition, self-checking and tool selection. Anthropic alleges that operators used fixed prompts, automated injection and jailbreak-like instructions to solicit raw reasoning traces.
Cross-session replay went further: attackers allegedly saved compressed "thinking signatures" from one Claude interaction and reintroduced them in new sessions, asking the model to expand them into fuller reasoning traces. Moonshot, DeepSeek and Z.ai are among the companies linked to replay behavior in the reporting. Anthropic's response β updating Claude to summarize rather than expose raw reasoning β directly targets this vector, though it also reduces transparency for legitimate users.
Silent Routing and the Privacy Problem
The most consequential allegation for ordinary users is silent routing. Bloomberg reportedβ that Anthropic says some Kimi and DeepSeek requests were forwarded to Claude, with the resulting answer shown as though it came from the local service. The request-response pair could then be retained as training material.
Reported examples of relayed content included corporate source code, live credentials and CCTV surveillance footage. If accurate, this creates a privacy issue separate from model intellectual property: users may have disclosed sensitive data to one provider without knowing that a second company's infrastructure would process it.
A relay service changes the data boundary. The customer may believe information remains inside one application, while the prompt, attachments and response are transmitted to a second model provider β and potentially retained for training.
The research does not establish how many relayed requests contained sensitive material, what retention rules applied, or whether affected users were notified. Those unanswered questions justify asking model providers to disclose upstream routing, fallback systems and subprocessors clearly β regardless of how the attribution dispute resolves.
Anthropic's Defences and Their Limits
Anthropic reportedly responded at the model, account and network levels. Its measures include:
- mandatory identity verification for accounts in regions assessed as high risk, including China, Russia and Iran;
- behavioral classifiers and fingerprinting to detect coordinated extraction patterns;
- output and reasoning summarization instead of raw chain-of-thought traces;
- cross-company intelligence sharing about abusive infrastructure;
- "preserved thinking" in Fable 5.1, intended to cryptographically bind conversation context and frustrate cross-session replay attacks.
None of these measures is complete. Identity requirements can exclude legitimate users. Classifiers can produce false positives. Altering responses to suspected extraction, as the CISA advisory recommendsβ, risks degrading service for wrongly flagged customers. Attackers can vary prompts, use third-party gateways, or shift to other frontier models.
Beijing's Denial and the Safety-Dialogue Setting
China's Ministry of Commerce rejected the allegationsβ as lacking factual and legal basis, characterizing them as an attempt to suppress Chinese AI development under a security pretext. Foreign Ministry spokesperson Mao Ning said China firmly opposes efforts to "smear China by distorting facts," and officials warned of possible countermeasures if the accusations supported new export restrictions or sanctions.
The dispute arrived as Washington and Beijing were preparing for an AI safety dialogue ahead of a planned September 24 leaders' summit. Reuters reportedβ that proposed subjects included AI-directed cyberattacks, autonomous agents, information sharing β and distillation. CNBC notedβ that the two nations had recently signed the "Carolina Principles," an agreement aimed at discouraging AI-specific regulations, suggesting some baseline of cooperation exists even amid the distillation dispute.
Such talks may chiefly provide channels for reporting incidents and comparing indicators. Verification of specific attribution claims remains difficult amid strategic competition, and neither side has proposed a joint technical audit.
Moonshot's Kimi K2.8 Preview: Forward Momentum Amid Controversy
One day after the Anthropic report, on September 11, Moonshot AI released Kimi K2.8 Preview β a closed, multimodal model positioned between the K2.7 Code and the flagship Kimi K3. Emergent.sh reportedβ that the model is accessible through the existing `kimi-for-coding` API endpoint, meaning developers do not need to update configuration settings to access it.
Key specifications as reported:
- Context window: 1,048,576 tokens (1 million), now available across all Kimi Code membership tiers β up from the 262,144-token limit of K2.7 Code;
- Multimodality: introduces image and video input support to the Kimi coding lineup for the first time;
- Thinking modes: adjustable reasoning effort (low, high, max), with max set as default;
- License: proprietary, closed-source β no downloadable weights;
- Pricing (third-party gateways): approximately $0.80/M input tokens, $3.35/M output tokens, $0.14/M cached read tokens.
Magicshot.ai's analysisβ notes that Moonshot characterizes K2.8 Preview's performance as "approaching K3" β a vendor claim without independent benchmark verification as of this writing. The "Preview" designation signals active development; the model is not yet listed on the Moonshot Open Platform for general API access.
The release is notable for what it is not: an open-weight model. Kimi K3, released in July 2026 with 2.8 trillion parameters under the Kimi K3 License, established Moonshot's open-weight credentials. K2.8 Preview's proprietary serving suggests the company is differentiating between its flagship open-weight release and its high-frequency production endpoint β a pattern also visible in how DeepSeek maintains open weights for V4.1-Flash while keeping V4-Pro on a commercial API.
What the Dispute Means for Open Weights
The allegations create a temptation to equate open Chinese models with unlawfully acquired capabilities. The evidence does not support that shortcut. A model can have openly released weights while containing capabilities learned through disputed data practices; equally, a closed model can be trained through fully licensed, consensual methods. Distribution format and training-data legitimacy are different questions.
Ordinary distillation can involve permission, licensed APIs, a developer's own teacher model, published research datasets or outputs supplied with user consent. The alleged misconduct here concerns circumvention and deception, not the mathematical technique itself.
For the open-weights ecosystem, three practical distinctions matter:
- Access: Were teacher-model outputs obtained under valid authorization, or through fraudulent accounts and stolen credentials?
- Consent: Did users know where their prompts and files were processed, or were requests silently relayed to a third-party model?
- Provenance: Can the developer document how synthetic training data was generated, governed and audited?
The most constructive response is not a presumption against open releases. It is stronger provenance documentation, explicit routing disclosure and auditable contractual permission for synthetic data. The September allegations are serious because of their claimed scale and privacy implications. They remain allegations whose company-level attribution and full technical evidence have not been independently tested β and the Chinese AI ecosystem's credibility in the open-weights space will depend, in part, on how transparently the named labs respond.
---
*Sources and further reading are listed below.*
Links & Resources
External links β opens in a new tab

π¨π³ China Desk Lead Β· Beijing, China
Reads the Mandarin sources first β DeepSeek, Qwen, Zhipu, and the rest.

Neural Avalanches: Neurodynamics and Brain Development
by Richard Murdoch Montgomery
Critical phenomena in the developing brain β power-law scaling, avalanche dynamics, and self-organized criticality in neural circuits.

A Treatise on Real Analysis
by Richard Murdoch Montgomery
Foundations, structure, and the architecture of the continuum β a rigorous graduate text on measure theory, integration, and topology.

The Casio fx-CG50: A Comprehensive Academic Treatise
by Richard Murdoch Montgomery
A 223-page deep dive into hardware architecture, statistical analysis, matrix operations, and Casio BASIC programming.

History of Evolutionary Thought in the Nineteenth Century
by Richard Murdoch Montgomery
From Lamarck to Darwin and beyond β a scholarly account of how evolutionary theory reshaped biology, society, and philosophy.
Comments
Open discussion β no account needed. Be respectful.
More from Chinese Models Desk
Zhipu's GLM-5.3-FlashX Hits 200 Tokens Per Second β and the AI That Built It Is Running on Chinese Chips
Z.ai's new FlashX serving tier for GLM-5.3-Flash achieves a fourfold speed leap over the standard tier by deploying an AI agent to rewrite its own inference stack β all on a 100,000-accelerator cluster of domestically produced Chinese chips. It's the most concrete demonstration yet that China's sovereign compute ambitions are becoming real.
Sophia ChenDeepSeek-V4.1-Flash Rewrites the Rules on Agent Memory: 552B Parameters, MIT License, and a KV Cache That Fits in Your Pocket
DeepSeek's latest model slashes KV cache memory to 890 bytes per token through a radical Causal Encoder-Decoder architecture β then releases the weights for free under MIT. The result is a multimodal 552B-parameter model that outperforms its own V4-Pro flagship on agentic benchmarks at a fraction of the cost.
Wei LianAlibaba's Qwen3.8 Open Weights Are Here β and the Two Licenses Tell You Everything About Where Chinese AI Is Heading
Alibaba dropped open weights for both Qwen3.8-Max and Qwen3.8-27B within 48 hours of each other β but the two models carry two very different licenses, and that split reveals exactly how China's most prolific AI lab plans to monetize the open-source era it helped create.
Sophia Chen