Who Is Running Up Your AI Bill at 3am

Security researchers have mapped an ecosystem of more than 80,000 proxy servers quietly routing stolen AI credentials to frontier models, and companies are footing bills they never ran up.

ThreatVectr Newsdesk· Editor: Lee Brown· 5 min read
Photoreal news-editorial style, 16:9 framing, full-frame edge-to-edge composition
Share

Key points

  • Palo Alto Networks' Unit 42 team has responded to a growing number of AI token-jacking cases, where criminals steal API keys to use expensive AI services at the victim's expense, resulting in losses of hundreds of thousands of dollars.
  • Researchers at Team Cymru identified more than 80,000 proxy servers, computers that pass traffic through to hide where it really came from, built specifically to launder stolen AI credentials.
  • An Okta Threat Intelligence analysis of infostealer data dumps found 561 Anthropic session tokens collected from nearly 5,871 infected machines.
  • Several clusters of suspicious traffic through these proxies have been tied to IP addresses in China and Hong Kong, pointing to large-scale attempts to extract knowledge from Western AI models.
  • The NSA, CISA and FBI jointly published an advisory accusing six China-based AI companies of conducting industrial-scale extraction attacks against US AI services through exactly these kinds of proxy networks.

At three in the morning, a company's AI system is still running. Requests are going out, responses are coming back, and a bill is quietly accumulating. Nobody at the company sent those requests. Criminals did, using stolen access codes called API keys: long strings of characters that act like a password, letting software connect to an AI service and charge usage to the key's owner.

This is AI token jacking. Palo Alto Networks' Unit 42 incident-response team flagged the trend in an advisory published this month, describing a surge in cases where victims discovered enormous AI bills generated entirely by criminals who had stolen their credentials.

How do criminals get these keys in the first place?

Three main routes: information-stealing malware that silently copies files and browser data from an infected computer; phishing campaigns that trick staff into handing over passwords; and poisoned software packages on code-sharing platforms used by developers.

Once a key is stolen, criminals don't use it directly. That would be too easy to trace. Instead, they route requests through a relay network, a chain of intermediary computers that passes traffic along while hiding its true origin. Team Cymru initially counted 10,867 such relay servers running open-source software, then expanded their scan and found more than 80,000.

Two of the most common relay platforms, called Claude Relay Service and its successor sub2api, were published publicly on GitHub. The sub2api project alone has been copied more than 8,000 times and has close to 7,000 subscribers on Telegram. Its GitHub page lists 26 commercial sponsors, among them API relay resellers and proxy vendors. Parts of this infrastructure operate in a grey zone rather than deep in the criminal underground.

What are the criminals actually doing with the access?

Resale is the simpler use: someone buys cheap access to GPT-4 or Claude through a relay marketplace, unaware the underlying key is stolen and a legitimate company is paying the real bill.

Model distillation is the more strategically alarming application. An attacker sends thousands of carefully chosen questions to a high-quality AI model, then uses the answers to train a cheaper model to mimic it. Team Cymru analysed traffic from over 4,000 IP addresses in China and Hong Kong connecting to 304 relay servers over eight days in late August. Those addresses sent roughly 14 terabytes of data to the relays and received about 7 terabytes back. Seventeen relays passing traffic to Anthropic's API uploaded approximately 81 gigabytes while receiving only about 1.4 gigabytes back, a ratio of 58 to 1. Team Cymru estimated that upload volume could represent between 16 billion and 23 billion input tokens of text, consistent with automated bulk querying for distillation.

Our 17 September story "US Agencies Say Chinese AI Firms Are Stealing From American Models at Industrial Scale" named the six companies the NSA, CISA and FBI accused of running this operation, among them DeepSeek, Alibaba and MiniMax.

Credential type Count found in analysed data
Anthropic session tokens (from 5,871 infected machines) 561
Anthropic tokens still valid at time of analysis 164
API keys across Gemini, OpenAI, Groq, OpenRouter 24 (valid)
Gemini keys in one criminal collection 448
OpenAI keys in same collection 254
Anthropic keys in same collection 176

Should developers and businesses be worried?

Yes, particularly any organisation whose developers use AI services through API keys. AI pricing is consumption-based, and most accounts allow unlimited usage by default unless a spending cap is manually configured. The bills can scale fast.

Unit 42's advice: treat AI keys the way you treat any sensitive password. Set spending limits and configure alerts for unusual activity. Swap long-lived keys for short-lived ones that expire quickly, so a stolen key becomes useless fast. Scan code repositories and configuration files regularly for exposed credentials. Build a clear plan for revoking and rotating keys the moment a breach is suspected. Our 17 September story "A Self-Destruct Button for Stolen API Keys" covers a proposal that would cancel any leaked key automatically within sixty seconds of discovery.

Customers of AI services who aren't developers are less directly exposed, but if your company uses AI tools built by a vendor, it's worth asking how that vendor monitors for credential misuse. This relay economy is running right now, somewhere between 80,000 proxy servers and your next invoice. The number that stands out isn't the proxy count: it's the 58-to-1 upload ratio on Anthropic traffic, because that's what systematic extraction looks like in the logs.

© 2026 Threat Vectr