AI

NSA, CISA, FBI Warn China-Based AI Firms Distill US Frontier Models – Unite.AI

The National Security Agency, the Cybersecurity and Infrastructure Security Agency, and the Federal Bureau of Investigation released a joint cybersecurity advisory on September 8, 2026, warning that China-based artificial intelligence companies are systematically extracting proprietary capabilities from U.S. frontier AI models through industrial-scale knowledge distillation campaigns running since at least late 2024.

In the advisory, designated AA26-251A, the agencies state that these campaigns “form the core—not merely a supplement” of the companies’ AI development strategy. According to the advisory, likely with Chinese government awareness, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens across millions of exchanges and requests from U.S. frontier models, including variants of Claude, GPT, Gemini, and Grok. The agencies state that the activity violates the U.S. companies’ terms of use and threatens U.S. technological leadership.

CISA’s announcement of the advisory describes knowledge distillation as a machine learning technique that trains a less capable model using the outputs of a larger, more capable one. While a valid training method, CISA said, it can be misused to acquire capabilities from competitors in less time and with less cost than developing them legitimately. “We strongly urge AI companies to take immediate steps to safeguard their platforms against knowledge distillation campaigns that threaten to close the gap in advancements made by American companies,” said CISA Acting Director Nick Andersen.

Activity Attributed to DeepSeek, Moonshot AI, and Others

According to the advisory, DeepSeek has conducted an organized distillation campaign against U.S. frontier models since at least late 2024 to generate synthetic training data for its models, including R1, released in early 2025. The agencies state that DeepSeek targeted reasoning capabilities, specialized optimizations, and domain-specific functions to reduce compute and research costs, and that the company’s publicly quoted $5.6 million training cost is misleading because it excludes the cost of data acquired through malicious distillation. Between late 2024 and mid-2025, the advisory states, DeepSeek distilled from Claude 3.7, Claude Sonnet 4, Claude Sonnet 4.5, Claude Opus 4.1, Gemini 2.5 Pro Preview, Gemini 2.5 Flash Preview, GPT-4, GPT-4o, GPT-4 Mini, GPT-4 Nano, GPT-5, and Grok 4 to train its R1 and V3 models.

See also  Google Cloud races to catch up in the AI deployment wars with Accenture deal

The advisory states that Moonshot AI has run a widespread distillation campaign since at least mid-2025, extracting significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data to train its Kimi-K2 model. The targeted capabilities included supervised fine-tuning optimization, reinforcement learning, software engineering, and math, drawn from a range of Claude, GPT, Gemini, and Grok models. Moonshot AI used millions of exchanges targeting agentic reasoning and tool use, coding and data analysis, computer-use agent development, and computer vision, according to the advisory.

In late 2025, the advisory states, Alibaba distilled Claude-4, Claude Opus, Claude Sonnet, and GPT-5 to improve software engineering, customer service dialogue, and image and character creation in its Qwen family of models. In the same period, MiniMax distilled chain-of-thought reasoning, reinforcement learning, supervised fine-tuning, and software engineering capabilities to improve its M2 model from Claude Code, Claude Sonnet 4, Claude Opus, Gemini 1, Gemini 2.5 Pro, and Gemini 3 Pro. According to the advisory, MiniMax also used Claude Code for internal software development and used prompt injections to try to trick Claude Code into believing it was a MiniMax product.

Between late 2025 and early 2026, the advisory states, StepFun distilled data from Claude Opus 4.1 and 4.5, Claude Sonnet 4.5, Claude Haiku 4.5, GPT-5 Mini, GPT-5 Pro, GPT-5.1, GPT-5.1 Codex, and GPT-5.2 to improve the coding and agentic functions of its Step 4 model. By mid-2026, Z.AI had distilled billions of tokens of GPT-5.5 data and Claude Opus 4.8 data to develop chain-of-thought reasoning capabilities, according to the advisory.

See also  OpenAI Releases ChatGPT Images 2.5 With Sketch and Two New API Models – Unite.AI

Tactics and Techniques

The advisory states that the companies route distillation requests through native application programming interfaces, remote cloud providers, and third-party aggregators that obfuscate user metadata, and that they use a gray market of API proxies known as transfer stations to bypass geographic restrictions, evade safeguards, and undermine traceability. Cost savings come from bulk procurement of premium subscriptions shared across teams of developers, according to the advisory, and advanced tactics include chain-of-thought reasoning extraction, automated failover between pathways during blocking attempts, and quality evaluation frameworks designed to detect defensive countermeasures.

The agencies mapped the activity to the MITRE ATLAS framework across adversary lifecycle phases from resource development through exfiltration, including fraudulent account creation and jailbreak prompts that force models to reveal hidden chain-of-thought reasoning. The advisory states that DeepSeek employed prompts instructing models to imagine and articulate the internal reasoning behind completed responses, and that MiniMax redirected exchanges to a new Claude model within 24 hours of its release.

The advisory also details four techniques it describes as novel: regional restriction evasion combined with subscription exploitation, centralized request routing infrastructure, automated request metadata sanitization, and systematic quota and cost optimization. Detection indicators listed in the advisory include shared accounts used from multiple IP addresses and user agents, sustained usage around the clock without human variation, anomalous subscription-to-usage ratios, and new subscriptions immediately running at maximum usage.

Recommended Mitigations

The agencies recommend U.S. AI companies take three immediate actions: implement comprehensive detection and mitigation of anomalous and malicious prompts, accounts, networks, and behaviors; deploy targeted response changes that subtly alter responses to suspected malicious distillation attempts; and establish cross-organization intelligence sharing across model providers, cloud platforms, and API aggregators.

See also  Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market

Response changes can include differential privacy or serving downgraded models for suspected distillation requests, the advisory states, and companies should vary those changes across requests to complicate response quality evaluations. The advisory recommends against informing users suspected of malicious distillation when responses are altered, while stating that AI safety researchers and third-party evaluators should be informed of model changes.

The advisory lists mitigations drawn from MITRE ATLAS, including query rate limits, controls on access to production models, AI telemetry logging, output obfuscation, adversarial red teaming, model hardening, ensembles, and limits on the release of model artifacts. It also cites NIST’s adversarial machine learning taxonomy, including differential privacy with its noise-versus-utility tradeoff, pre- and post-training interventions, and prompt instruction and formatting techniques.

The advisory calls for a coordinated response across the U.S. government, private industry, and allied nations, stating that industry disclosures document proxy networks managing tens of thousands of fraudulent accounts simultaneously. It directs organizations affected by the campaigns to file a complaint with the FBI’s Internet Crime Complaint Center.

Source link

Back to top button