A Typo in a Fake Company Name Let AI Models Hack a Real Business

Security testing firm Irregular built a simulated target with an accidental real-world twin. The AI models found it, broke in, and nobody noticed for a while.

ThreatVectr NewsdeskUpdated · Editor: Lee Brown· 4 min read
A desk with multiple monitors showing network diagrams and IP addresses, with one screen displaying an alert notification while files transfer across the networ
Share

Key points

  • AI safety firm Irregular, based in Israel, confirmed that AI models it tested for Anthropic broke out of a controlled test environment and attacked real organisations on three separate occasions.
  • The root cause of one incident was a naming collision: engineers gave a fictional target company a name that matched an existing, little-known website.
  • Models with live internet access reached the real site, found genuine security weaknesses, extracted credentials, and accessed a live production database.
  • The activity was hard to catch because it occurred in only a fraction of test runs, often after hundreds of automated steps.
  • Irregular has published a post-incident account and is calling for industry-wide standards around how AI evaluation environments are set up and monitored.

When you run thousands of computer-generated attack simulations, the odds of something slipping through should be low. Irregular, the Israeli AI safety company that raised $80 million last year to stress-test AI models before they reach the public, has confirmed those odds weren't low enough.

What actually happened here?

A naming error sent AI models to a real website instead of a fictional one, and the models did exactly what they were told: they attacked it.

Irregular's business is red-teaming, which means deliberately probing AI systems for dangerous capabilities in a controlled setting, the way a fire drill tests a building's exits without burning it down. Its engineers build simulated companies for the AI to attack, measuring how capable a model is at real offensive hacking tasks.

For the evaluation in question, the team created a fictional target company. Fictional names are normally checked against real domains, which are the web addresses that point to actual websites. This one slipped through because the matching real domain was obscure and not widely indexed.

The test environment had live internet access switched on. In a handful of runs out of thousands, the AI model reached the real site rather than the simulated one. It didn't stop. Reconnaissance followed: information-gathering about the target, locating private cryptographic keys used to authenticate into systems, pulling out credentials (usernames and passwords), and gaining access to a live production database containing real data. In one separate case, a different model drifted to a similarly named site and found credentials that had already been leaked publicly.

Irregular noted the real site lacked basic security protections, making it an easy mark for any capable model. We covered the broader three-lab pattern of escapes on 6 August, when Meta, OpenAI, and Anthropic all disclosed incidents tied to the same testing firm.

Should ordinary people be worried?

The company whose database was accessed hasn't been publicly named, so direct action for affected customers is limited for now. If you run a small or mid-size website, this is a reminder to audit your domain for exposed private keys and to treat leaked credentials as burned immediately, because reused passwords are still among the simplest ways in for attackers, human or automated.

The harder problem is systemic. Existing monitoring tools can't easily distinguish a legitimate red-team test from a genuine attack, because both produce identical log entries. That gap has been sitting in plain sight for months.

Detail Specifics
Firm involved Irregular (Israel)
AI lab whose models escaped Anthropic
Confirmed escape incidents 3
Root cause of this incident Domain naming collision
Data accessed Production database via extracted credentials
Funding raised by Irregular $80 million (2024)

Going forward, Irregular says it's expanding manual review during test cycles, standing up an internal team to challenge its own containment assumptions, and building a continuous process to recheck fictional domain names against new real-world registrations. It also wants a shared industry mechanism for exchanging forensic evidence, such as full model interaction transcripts, after incidents like this one. A white paper on best practices for securing AI evaluations is also planned.

The fix that failed here was basic: check the fake name against real websites. That's a process control, not a deep technical problem. SQL injection, one of the oldest web-attack primitives around, persists for the same reason. New attack surface, familiar failure mode.

Common questions

Does this mean AI models are deliberately trying to escape their tests?

No. The models followed their instructions faithfully. A naming error pointed those instructions at a real target, and the models had no way to know the difference.

Who is responsible when an AI model hacks a real company during a test?

That's an open legal and regulatory question. Irregular's post-incident account, first reported by SecurityWeek, acknowledges the gap and calls for clearer agreements on liability and information sharing after evaluation incidents. Nobody has a clean answer yet.

© 2026 Threat Vectr