US Agencies Say Chinese AI Firms Are Stealing From American Models at Industrial Scale

A joint NSA, CISA and FBI advisory names DeepSeek, Alibaba, Moonshot AI, MiniMax, StepFun and Z.AI as running organised campaigns to copy the inner workings of Claude, GPT, Gemini and Grok.

ThreatVectr Newsdesk· 4 min read
Aerial view of a dark server room with rows of blinking blue and amber indicator lights, dense cable bundles running between racks, cold air vents casting faint
Share

Key points

  • The NSA, CISA and FBI issued a joint advisory naming six Chinese AI companies conducting large-scale copying of US frontier AI models since late 2024.
  • DeepSeek, Alibaba, Moonshot AI, MiniMax, StepFun and Z.AI extracted billions of tokens from Claude, GPT, Gemini and Grok, likely with Chinese government awareness.
  • The campaigns route through API proxies known as "transfer stations" to bypass regional blocks and hide who is asking the questions.
  • DeepSeek's widely quoted $5.6M training bill for its R1 model excludes the data it acquired through this activity, the agencies say.
  • US AI companies are urged to watch subscription-to-usage ratios, quietly alter answers to suspected copycat queries, and share attack signals across providers.

Three US agencies have gone public with an unusually direct accusation: China's leading artificial intelligence companies are not just competing with American labs, they are systematically copying them.

The National Security Agency, the Cybersecurity and Infrastructure Security Agency and the FBI released the joint advisory this week. It names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI, and says the six firms have been running what the agencies call "industrial-scale distillation" campaigns since at least late 2024.

What is "distillation" and why does it matter?

Distillation is a technique where a smaller AI model learns by watching a bigger, smarter model answer questions. The student copies the teacher. It is a legitimate research method when you own both models. It becomes a problem when the teacher belongs to someone else and never agreed to teach.

According to the advisory, the Chinese firms sent millions of requests to US models including Claude, GPT, Gemini and Grok, harvested billions of tokens of output (a token is roughly a word fragment), and then used those answers as training material for their own systems.

The agencies say this is not a side project. It is the core of how these companies build models, and it is likely done with the Chinese government's knowledge.

Who is accused of what?

Company Models trained US models copied from
DeepSeek R1, V3 Claude 3.7 to Opus 4.1, GPT-4 to GPT-5, Gemini 2.5, Grok 4
Moonshot AI Kimi-K2, Kimi-K3 Claude Sonnet and Opus family, GPT-4o, GPT-5, Gemini 2.5, Grok Code Fast-1
Alibaba Qwen family Claude 4, Claude Opus, Claude Sonnet, GPT-5
MiniMax M2 Claude Code, Claude Sonnet 4, Claude Opus, Gemini 2.5 Pro, Gemini 3 Pro
StepFun, Z.AI various multiple US frontier models

DeepSeek gets particular attention. Its R1 model, released in early 2025, made global headlines for reportedly costing just $5.6 million to train. The agencies call that figure misleading, arguing it leaves out the value of data siphoned from US competitors.

How are they getting the data out?

The advisory describes a small industry built around evading detection. Chinese firms route requests through official APIs (the doors that let software talk to an AI model), through cloud providers, and through third-party aggregators that strip out identifying information.

A gray market of proxies called "transfer stations" resells access into regions the US firms have blocked. Teams share bulk-bought premium subscriptions to keep costs down. If one path gets shut off, automated systems failover to another.

The more advanced tactics include pulling out "chain-of-thought" reasoning, the step-by-step working a model shows before giving an answer. That reasoning is what makes frontier models valuable, and it is what the copycats want most.

What are US companies being told to do?

The advisory sets out three moves. Watch for unusual patterns like brand new accounts that immediately hit maximum usage, or subscriptions whose traffic looks nothing like a human developer's. Quietly change answers when a request looks like a distillation attempt, so the copied data becomes unreliable. Share attack signals across model providers, cloud platforms and API resellers, because the campaigns are deliberately spread thin to avoid tripping any single alarm.

Should ordinary users be worried?

Not directly. If you use ChatGPT, Claude, Gemini or a Chinese chatbot, this dispute does not put your personal data at extra risk today. The concern is longer-term and commercial: US labs spend hundreds of millions training a model, and the agencies say Chinese rivals are getting the benefit for a fraction of the cost, which affects who leads the field in five years.

For businesses building on top of these APIs, the practical takeaway is to expect stricter account verification, tighter rate limits, and more aggressive anomaly detection from US providers in the months ahead.

© 2026 Threat Vectr