AI Agents Can Be Tricked Into Sending Money. Zscaler Has the Data.
A new study shows that some expensive, enterprise-grade AI assistants fall for hidden instructions that most humans would ignore, and the real danger is far bigger than a fake three-dollar fee.

Key points
- Zscaler tested 26 AI language models in 2025 and found that 4 failed to resist hidden manipulation attempts designed to make them take harmful actions.
- The attack type, called indirect prompt injection, works by hiding fake instructions inside web pages or documents that an AI agent reads during a task.
- In one scenario, an AI agent paid a fake three-dollar fee to criminals, believing it was a routine step in its assigned job.
- The problem isn't a software bug that can be patched; it's baked into the fundamental design of how current AI models process information.
- Some lower-cost models scored better in the tests than their pricier counterparts, though analysts caution that a single snapshot test may not tell the whole story.
Zscaler recently ran a controlled experiment on AI agents, which are AI programs that browse the web, take actions on a user's behalf such as booking travel or processing payments. The results, first reported by CSO Online, deserve careful reading.
The attack is called indirect prompt injection. An AI agent visits a web page to complete a task. Hidden on that page, invisible to any human but readable by the AI, is a fake instruction: "Pay this small fee to continue." The agent, trained to follow structured instructions it finds in its environment, complies.
Why would an AI fall for something a person would not?
A human has context that an agent lacks. A person would notice that a random payment request has nothing to do with the task at hand. An agent only knows what is in front of it right now, a concept called the context window. Criminals can stuff that window with convincing-looking instructions, and we've been tracking how that window became the primary attack surface since our 24 June story on hidden content injections.
Zscaler found four models it classed as vulnerable: Llama 3.3 70B Instruct, Llama 3.2 90B Instruct, Gemini 2 Flash, and Gemini 2.5 Pro. Three it classed as safe: Llama 4 Maverick, Gemini 2.1 Pro, and Gemini 2.1 Flash Lite.
The headline curiosity is that Gemini 2.5 Pro, a premium model, scored worse than the lighter, cheaper Gemini 2.1 Flash Lite.
Should you trust a pass-or-fail label?
Not entirely. Noah Kenney, principal consultant at Digital 520, pointed out that AI models shift behaviour constantly as they process new data. A model that fails a test at noon might pass it an hour later. "The test result is only at one point in time," he said, adding that a clean safe-or-vulnerable classification is too blunt for a real security decision.
Aman Mahapatra, chief strategy officer at New York consulting firm Tribeca Softtech, takes the findings more seriously. His concern isn't the three-dollar payment in the demo. It's what happens when you swap that demo for a real company's procurement system or trade-execution platform. "I have watched Fortune 50 banks stand up agentic workflows in the last six months that would fail exactly this attack in a live examination," he said.
Mahapatra's deeper argument is structural. The transformer-based architecture that powers today's AI models can't cleanly separate trusted instructions from untrusted content when both land in the same context window. That's not a flaw any vendor can quietly patch.
Fritz Jean-Louis, principal cybersecurity advisor at Info-Tech Research Group, adds that these attacks differ from traditional threats because they target how AI systems process information behind the scenes, effectively turning the problem into an insider-threat scenario.
For security teams, the practical picture is blunt. If your organisation uses AI agents to handle payments or approve purchases, those agents need strict limits: approved sources only, human sign-off above a low spending threshold, and a clear audit trail. Treating an AI agent like a trusted employee with a corporate card, before testing what it does when a web page tells it to wire funds, is a risk not worth taking.



