Most AI Models Fail New Malware-Investigation Test Inspired by Nuclear Sabotage

A benchmark built around a real industrial-attack case shows that today's leading AI systems struggle to follow a malware investigation all the way to a conclusion.

ThreatVectr Newsdesk· 3 min read
Full-frame photoreal editorial shot of a developer workstation at night, multiple monitors showing abstract code and an error-tracking dashboard with blurred un
Share

Key points

  • SentinelOne, a cybersecurity company, built a new benchmark to test whether AI models can fully investigate a piece of malware from first clue to final verdict.
  • The benchmark is built on the "Fast16" case, a reference to malware linked to industrial sabotage, including attacks on nuclear facilities.
  • Most leading AI models tested could not sustain a complete investigation, failing partway through the analytical chain.
  • No AI model is currently reliable enough to run a malware hunt without a human analyst checking its work.

SentinelOne has built a test that most of the world's best-known AI systems quietly fail.

The test is called the Fast16 benchmark. It asks an AI to do what a senior malware analyst does: take a suspicious piece of software apart, piece by piece, until the whole picture is clear. Not just spot that something looks bad. Actually finish the investigation.

Most models cannot.

What is the Fast16 case, and why does it matter?

Fast16 refers to a set of malware samples connected to attacks on industrial control systems, the kind of software that runs power grids, water treatment plants, and nuclear facilities. The name has circulated in threat-intelligence circles as a shorthand for a particularly complex class of attack.

SentinelOne chose it precisely because of that complexity. Investigating this malware is not a one-step task. It requires an analyst to hold many threads at once, pivot from one clue to the next, and reason about what the attacker was ultimately trying to do. That chain of reasoning is where AI models break down.

Why do AI models struggle with this?

They lose the thread. An AI model works well on individual, self-contained questions. Ask it to explain one function in a piece of code, and many models do a decent job. Ask it to connect that function to a command-and-control server, link that server to a known criminal group, and then decide whether the malware's payload could disable industrial equipment, and the model starts to drift.

Security researchers call this "sustained reasoning." Humans do it naturally during an investigation. Current AI systems, it turns out, do not.

This matters because a growing number of security vendors are marketing AI as a tool that can handle investigations with minimal human oversight. SecurityWeek, which first reported on SentinelOne's research, noted the benchmark lands at a moment when those claims are facing real scrutiny.

Should security teams be worried about relying on AI?

Yes, if they are relying on it unsupervised. No tool that cannot finish a thought is safe to leave alone.

That is not a reason to discard AI from security workflows. Analysts already use it to speed up early triage, the first rough sort of which alerts are worth a closer look. Where it earns its place is as a first-pass assistant, not a replacement for the human who connects the final dots.

For ordinary people, the takeaway is simpler. Critical infrastructure, hospitals, utilities, the systems that keep daily life running, depends on skilled human investigators. A tool that fails this benchmark, without a person checking its work, leaves a gap that a sophisticated attacker can walk through.

SentinelOne has not published a full list of which models passed or failed. That detail, when it arrives, will tell a clearer story about which AI vendors' confidence is warranted.

© 2026 Threat Vectr