A 15-Minute Framework for Spotting What AI Systems Can Do Wrong

Security expert Adam Shostack built PHANTOM-B to help organisations find the risks hiding inside AI-powered software before those risks find them.

ThreatVectr NewsdeskUpdated · Editor: Lee Brown· 4 min read
A whiteboard or digital canvas displaying the PHANTOM-B framework with interconnected boxes showing different categories of AI system risks, with a security exp
Share

Key points

  • Adam Shostack presented PHANTOM-B, a new AI threat modelling framework, at Black Hat USA 2026.
  • A client's quickly built AI app, pointed at real customer data with no formal security review, prompted the first real-world test, surfacing meaningful risks in 15 minutes.
  • PHANTOM-B covers eight AI risk categories: Prompt injection, Hallucination, Anthropomorphization, Non-explainability, Training issues, Overreliance, Missing security engineering, and Bias.
  • OWASP, NIST and Microsoft have each published AI security guidance, but no single agreed standard exists.
  • Companies are deploying AI faster than their security teams can review it.

A few weeks ago Adam Shostack opened an email from a client. One of their staff had assembled an AI-powered app using a coding tool, with minimal human review, and pointed it at real customer data. The client wanted to know what could go wrong. Fast.

Shostack gave himself 15 minutes.

By the end he had a list of genuine risks, including hallucination (where an AI system confidently produces false information) and bias (where outputs systematically favour or disadvantage certain groups). Neither would have appeared on a standard security checklist.

What is PHANTOM-B and how does it work?

PHANTOM-B is a threat modelling framework, a structured checklist that security teams use to find weaknesses before attackers do. Shostack built it specifically for software that uses large language models (LLMs), the AI engines behind tools like ChatGPT.

Each letter names a distinct risk:

Letter Risk Plain meaning
P Prompt injection Criminals hide instructions inside text the AI reads, hijacking its behaviour
H Hallucination The AI produces confident but false outputs
A Anthropomorphization Teams trust the AI as if it thinks like a person, granting it more authority than it deserves
N Non-explainability The AI cannot reliably explain how it reached a decision
T Training issues Flawed or manipulated training data shapes bad outputs
O Overreliance Staff or systems depend on AI output without adequate checks
M Missing security engineering Conventional software vulnerabilities that AI deployments still carry
B Bias Systematic unfairness baked into the model's responses

Shostack presented the framework at Black Hat USA 2026. PHANTOM-B sits alongside existing tools, not above them. The older STRIDE framework, a widely used checklist for conventional software, still handles the rest of an application. PHANTOM-B handles the AI parts.

We first covered Shostack's work on AI security guardrails on 17 August 2026, when he reviewed OpenAI's Black Hat presentation on AI models passing covert messages during training.

Why does ordinary security checking fall short for AI?

Conventional software follows fixed rules. Feed it the same input twice and you get the same output. AI systems don't work that way. They interpret plain-language instructions and generate responses that can vary each time, making their behaviour genuinely hard to predict.

That unpredictability creates gaps that existing tools were never designed to catch.

Jeff Williams, founder of OWASP (the Open Worldwide Application Security Project, a non-profit that publishes free security guidance) and CTO of Contrast Security, told CSO Online: "AI didn't break threat modelling. It exposed weaknesses that were already there."

Brian Glas, vice president at security consultancy CODIFIC and a project lead for the OWASP Top 10 for LLM Applications, points to a second problem: weaknesses can chain across AI systems. A poisoned document fed into an AI assistant could alter its response, which then causes a connected tool to take a harmful action elsewhere. "A chaining of threats or weaknesses results in the exploitation of the application," Glas told CSO Online.

Should businesses worry about apps built without security checks?

Yes, and the pattern is spreading. Shostack's 15-minute test was triggered by exactly the kind of software security teams rarely see: a quickly assembled AI app handling real customer data, built by someone who wasn't a professional developer, deployed before anyone asked hard questions.

NIST, the US National Institute of Standards and Technology, recommends in its AI Risk Management Framework that threat modelling should be a continuous habit, not a one-off audit. Microsoft's AI security guidance makes a similar point: when AI systems can access data or trigger other tools, one weakness can ripple outward in ways that are hard to trace.

Threat modelling itself isn't new, but as our August story on AI governance failures showed, unclear rules for how staff use AI tools are already breaking projects from the inside. Adding unreviewed AI apps to that mix raises the stakes further.

Shostack's practical advice is to start small. A 15-minute session won't catch everything, but it should answer whether the risk is acceptable or whether someone needs to look harder.

"You make the experiments cheap," he told CSO Online, "and when the experiment is cheap, you can run it repeatedly."

For anyone using an AI-powered product at work, the immediate question is whether your organisation has reviewed what the tool does with your data and formally checked what happens when it produces wrong or biased results. The honest answer at most companies right now is no.

© 2026 Threat Vectr