Poisoned Tool Descriptions Turn Helpful AI Agents Into Quiet Exfiltration Channels
Microsoft Incident Response shows how a single malicious MCP tool description can coax an agent into leaking corporate data without tripping a single policy check.

Key points
- Microsoft Incident Response research shows attackers can hijack tool-using AI agents through poisoned tool descriptions alone.
- The agent breaks no rules: every step looks routine, so no alert fires in a default configuration.
- Defenders rarely inspect tool manifests, treating them as configuration rather than untrusted content.
- Detection engineering for agent runtimes lags far behind the threat, with logging that records success rather than intent.
- The technique is capability research, not an observed intrusion, but the gap between published method and in-the-wild use keeps narrowing.
An AI agent that obeys every rule can still betray you.
That's the uncomfortable finding from new research by Microsoft Incident Response, which shows how attackers can hijack tool-using agents through nothing more exotic than a poisoned description string. The agent never violates policy, never escalates privileges. It just follows instructions buried in a place defenders rarely inspect.
The technique sits in a category researchers have been circling for a while, sometimes tracked as indirect prompt injection and sometimes, in the Model Context Protocol ecosystem, as tool-description poisoning. The shape is consistent: untrusted text reaches the model through a channel the developer treated as configuration, not content.
Here, the bait is the tool manifest itself. When an agent enumerates available tools, it reads each tool's natural-language description to decide when to call it. An attacker who controls or tampers with that description can embed instructions the model treats as authoritative. Ask the agent to pull a ticket summary, and the poisoned tool quietly redirects conversation data or credentials to an attacker-controlled endpoint.
No exploit. No CVE.
Microsoft's researchers note that in a default configuration, each step of the abuse chain looks routine. The agent calls a tool it's allowed to call, passes data it's allowed to read, and the destination is whatever the tool's implementation forwards to. Logging captures a successful tool invocation rather than a policy violation.
That's the operational problem. Detection engineering for agent ecosystems is immature, and most telemetry today answers "did the agent do something forbidden?" rather than "did the agent do something a reasonable human would have refused?"
Should you worry?
Yes, if your organisation runs tool-using agents against internal data. The technique is cheap and survives model upgrades because it doesn't target the model; it targets the trust the model places in its own toolbelt. When we covered MCP's enterprise security posture on 26 June, the pattern was already clear: the protocol hands the hard security work to whoever builds on top of it, and tool manifests were not on most teams' review lists.
For defenders, four things matter most. Treat tool descriptions and metadata as untrusted input on the same tier as user prompts: they are content. Pin and review third-party MCP servers and agent plugins the way you'd review a software dependency, including diffs on description fields between versions. Constrain egress from agent runtimes; an agent that can reach arbitrary outbound destinations is an exfiltration primitive waiting for a prompt. Instrument for semantic anomalies: unusual data volumes, unfamiliar tool chains, or calls that mix sensitive readers with external writers in a single session.
Attribution-wise, this is capability research rather than an observed intrusion set, and Microsoft frames it that way. But the shorter version matters: agents don't need to be jailbroken if their tools can lie to them.



