GuardFall: A 1970s Shell Trick Walks Past AI Coding Agent Safety Checks

Adversa AI says ten of eleven open-source coding agents fall to a command-substitution bypass that any sysadmin would recognize on sight.

ThreatVectr NewsdeskUpdated · Editor: Lee Brown· 3 min read
GuardFall: A 1970s Shell Trick Walks Past AI Coding Agent Safety Checks
Share

Key points

  • Adversa AI's GuardFall research found that ten of eleven open-source coding and computer-use agents fail to catch shell command-substitution syntax at the point of execution.
  • The one agent that held up, Continue, evaluates the resolved command rather than the raw string the model emits.
  • Agents typically run with the developer's own shell privileges, putting SSH keys, cloud credentials and session tokens in reach.
  • The failure is an authorization problem: the agent knows who is asking but cannot correctly constrain what gets executed.
  • No CVEs have been assigned; this is a design-pattern problem across an ecosystem, not a single product flaw.

The guardrails meant to keep AI coding agents from running destructive shell commands can be sidestepped with a syntax trick older than most of the people writing the agents.

Researchers at Adversa AI call the technique GuardFall. It abuses the gap between what a safety filter sees in a command string and what the underlying shell executes once command substitution kicks in. Wrap the dangerous part in $(...) or backticks and the surface-level pattern match reads something benign. The shell expands it before running. This is not a new class of bug: it is the same lesson sudo configuration guides have been repeating for decades, dressed up in an LLM wrapper.

When we reported on command injection reaching CISA's Known Exploited Vulnerabilities catalogue on 9 June, the underlying mechanic was nearly identical: a filter operating on the wrong representation of the command.

Should you worry?

Adversa tested eleven popular open-source coding and computer-use agents. Ten fell. The only holdout was Continue, which the firm says was built to evaluate the resolved command rather than the literal string the model emitted. Filtering the prompt-side artifact is authentication-adjacent theater; filtering the post-expansion command is what actually gates execution.

The blast radius depends on what the agent can touch. Most of these tools run with the developer's own shell privileges, which typically means access to source trees, SSH keys in ~/.ssh, cloud credentials cached by the AWS or gcloud CLIs, and any session tokens in the environment. An attacker who can influence the agent's input through a poisoned README, a malicious MCP (Model Context Protocol) tool description, or an issue comment the agent is asked to triage can turn a nominally safe command request into arbitrary code execution.

This is an authorization failure, not an authentication one. The agent correctly identifies who is asking; it simply cannot correctly decide what the resulting action is allowed to do. MFA would not have helped. Neither would a stronger model. The control plane is in the wrong place.

What would actually fix this

Four concrete controls are worth implementing now. Evaluate commands after shell expansion in dry-run or abstract syntax tree form, never as raw strings the model produced. Run agents inside a sandbox with an explicit allowlist of binaries and filesystem paths, treating the host shell as hostile. Strip or refuse command-substitution syntax in tool-call arguments where it serves no purpose. Log the resolved command rather than the requested one, so post-incident review reflects what actually ran.

Microsoft's containment work for agentic workloads, which we covered on 3 June, is one of the few architectural responses moving in the right direction, though GuardFall shows the gap between governance frameworks and the shell still needs closing manually.

Adversa's writeup is available from the firm directly at adversa.ai. No CVEs have been assigned at time of writing, which tracks: this is a design-pattern problem across an ecosystem, not a single product flaw.

A generation of agent frameworks shipped guardrails written as if bash were a sandbox. It is not, and it never was. The fix is older than the bypass, which should embarrass nobody in particular and worry everyone in general.

© 2026 Threat Vectr