Anthropic's AI Model Found Vulnerabilities in Classified U.S. Government Systems

An unnamed U.S. official says Anthropic's Mythos model identified security flaws in sensitive government infrastructure during a joint exercise with intelligence agencies.

ThreatVectr NewsdeskUpdated · Editor: Lee Brown· 2 min read
Anthropic's AI Model Found Vulnerabilities in Classified U.S. Government Systems
Share

Key points

  • Anthropic's Mythos model identified vulnerabilities in classified U.S. Government computer systems during a controlled testing exercise.
  • A U.S. Official speaking anonymously to the Associated Press confirmed the exercise involved U.S. Intelligence agencies.
  • No CVEs have been assigned and no public advisory exists.
  • Anthropic holds FedRAMP authorization and has existing defense and intelligence contracts.
  • The severity and class of vulnerabilities found have not been disclosed.

An AI model surfaced flaws in some of the most hardened infrastructure the U.S. Government operates. That's worth sitting with. Anthropic's Mythos model ran against classified systems in a structured evaluation, according to a U.S. Official who spoke anonymously to the Associated Press on Tuesday. Which agencies participated, which systems were tested, and what class of vulnerabilities Mythos found: none of that is public. Given the classification status of the systems, none of it is likely to be.

We first covered Mythos's offensive potential on 9 June, when XBOW's red team put the model through exploit discovery, reverse engineering, and live-site validation, with source-code review as the standout result. A government classification wall makes independent verification here impossible, but the capability being applied is not theoretical.

Should you worry about AI doing offensive security work?

The disclosure confirms that intelligence community operators are actively running AI models against sensitive targets in structured evaluations. The fact that results apparently warranted briefing an official willing to speak to press suggests Mythos produced findings meaningful enough to report up.

Automated scanning and fuzzing are not new. What changes when a reasoning-capable large language model starts correlating configuration data and known weakness patterns across a complex system, potentially surfacing attack paths a human analyst would deprioritize. That question, raised in our 9 June coverage of machine-speed vulnerability discovery, has now moved from the commercial bug-bounty context into classified infrastructure.

For defenders, the operational signal is that AI-assisted vulnerability discovery at scale is already inside the perimeter, sanctioned and structured.

How reliable is this disclosure?

Skepticism is warranted. A single anonymous official is thin sourcing for a significant claim. "Found vulnerabilities" could mean anything from a critical authentication bypass to a misconfigured storage bucket. The framing benefits Anthropic commercially regardless of severity. None of that makes the underlying capability implausible; it just means the actual technical weight of this disclosure is unknown.

If Anthropic or the relevant agencies publish anything substantive, that's when a real assessment becomes possible. Until then: confirmed interesting, severity unverified.

© 2026 Threat Vectr