Approval: Auto. The One-Line Bug That Turns an Agent Into a Shell

Approval: Auto. The One-Line Bug That Turns an Agent Into a Shell

A prompt injection doesn’t become remote code execution because a model is clever. It becomes RCE because the one control engineered to stop it — the approval gate — silently returns “yes.”

The cleanest example is a two-line function in CodeWhale, the Rust agent TUI. Its rlm_eval tool runs an arbitrary Python string the model picks, in a real python3 interpreter, on your machine at your UID. The tool advertises ExecutesCode and Network. The trait default would have demanded approval for any tool that can execute code. rlm_eval overrides it:

fn approval_requirement(&self) -> ApprovalRequirement {
    ApprovalRequirement::Auto          // overrides the safe default below
}

Auto is not “maybe ask.” The CVE advisory traces exactly what the engine does with it. The gate is two AND-ed conditions in engine.rs:

let approval_required = spec.approval_requirement() != ApprovalRequirement::Auto
    && !registry.context().auto_approve;

When approval_requirement() is Auto, approval_required is false. No Event::ApprovalRequired is emitted. The user’s --approval-policy (on-request, unless-trusted, never) is never consulted. So the operator who ran the agent under --approval-policy on-request believing every code-issuing tool would prompt is exactly as exposed as the one who passed --auto. The prompt isn’t defeated; it’s skipped at a prior branch.

The same defeat, four different shapes

This isn’t one framework’s sloppy default. It is a boundary that keeps coming back wrong. Two CVEs were assigned to CodeWhale on 18 Aug 2026 (CVE-2026-75858 for rlm_eval, CVE-2026-75857 for exec_shell_interact), and the class resurfaces across the ecosystem:

AdversaryMechanismWhy the gate failsSeverity
CodeWhale rlm_evalAuto → “never prompt” → model’s Python runs in python3Approval gate short-circuited before policy is read8.5 (NVD v4)
CodeWhale exec_shell_interactAuto → model stdin written into an already-approved REPL / mysql / ssh / sudo -iThe shell was approved once; the input running in it isn’t re-approved7.3 (NVD v4)
AWS Strands non_interactiveLLM-controlled param sets non_interactive=true, skipping the consent promptThe gate is an input-schema flag the model can flip(see bulletin)
atomic-agents-stack MCP catalogCleartext http:// catalog; MITM rewrites command/args; spawned as a local subprocessNo default allowlist — mcp_allow_fn defaults to None8.7 (CVSS v4)
Xinference tool-call parserUnsafe eval() parses model tool-call output, no auth on default deploymentModel output parsed as code, no gate at all10.0 Critical

CodeWhale exec_shell_interact: the approved shell is the attack surface

exec_shell itself is correctly gated. Its sibling exec_shell_interact is not. exec_shell_interact returns ApprovalRequirement::Auto, and when the model writes input into a shell the user already approved — a python3 -i REPL, mysql, ssh, sudo -i — no prompt fires. Inside those processes, stdin is the command surface. The user approved opening the shell once, for a stated purpose; the model then chooses the input, and prompt injection steers it.

The reach scales with the approved process: mysql -u root becomes arbitrary SQL, ssh host becomes commands on the remote host, sudo -i becomes root. None of it re-prompts.

The partial-fix gotcha

The uncomfortable part is what this shares with an older, already-fixed bug. CodeWhale’s run_tests tool had the identical flaw — ApprovalRequirement::Auto on a tool that runs cargo test, which compiles and executes arbitrary code (CVE-2026-45311, CVSS 9.6). The source even stated the intent: “Tests are encouraged, so avoid gating them behind approval.” It was fixed in 0.8.23. But the fix never reached rlm_eval or rlm_open, which expose a broader surface. The advisory says it plainly: “This is the same defect that was already patched on the sibling run_tests tool; the fix never reached rlm_eval or rlm_open.”

That is the pattern worth naming. Fixing one Auto doesn’t fix the default. As long as Auto exists as a per-tool escape that overrides the ExecutesCode → Required default, the class regenerates.

Amazon’s advisory is the least exotic but most instructive. The strands-agents-tools shell tool has a human consent gate that prompts the operator before commands run. It also exposed a non_interactive parameter in its input schema that the LLM controls. A crafted prompt — delivered via indirect prompt injection in untrusted content the agent reads — sets non_interactive=true, and arbitrary OS commands execute without operator approval. The gate is a boolean in the very schema the model fills in. Fixed by removing the parameter in 0.8.0. Until then, the workaround is operational: don’t give an existing-untrusted-content agent a shell, and isolate it least-privilege.

atomic-agents-stack: the gate is absent, and the catalog is in cleartext

The MCP registry backend in atomic-agents-stack accepts both http and https schemes for the catalog URL, and catalog entries carry command/args that are type-validated but content-unrestricted. Those are later spawned as local stdio subprocesses by MCPClientPool. Over a cleartext http:// catalog URL, a man-in-the-middle rewrites the catalog response to inject arbitrary command/args — and you get code execution on the agent host with no LLM involvement at all. The intended mitigation, a command allowlist (mcp_allow_fn), defaults to None, so without an operator-authored allowlist every resolved spec connects. Fixed in 1.1.0 by requiring https by default and gating http:// behind a loud explicit opt-in.

Xinference: eval() where tool-call output meets the parser

The most severe of the batch is Xinference (CVE-2026-61539), CVSS 10.0 Critical, published 21 Aug 2026. Its Llama3 tool-call parser did:

data = eval(model_output, {}, {})

eval() is not a sandbox — eval(model_output, {}, {}) executes the input as a Python expression in the Xinference server process. The model output is influenced by attacker prompts sent to the OpenAI-compatible /v1/chat/completions endpoint, and on the tested default deployment authentication was not enabled. A remote, unauthenticated attacker who can get the model to emit __import__('os').system('touch /tmp/hacked') gets it executed on the server. No approval gate, because there is no tool call to gate — the parser is the injection sink. Fixed in 2.7.0.

What this says about the boundary

The control loop’s verifier is the gate. When the gate returns Auto, the verifier is out of the loop. The model wasn’t the problem — the model “did what it was designed to do.” The architecture placed the gate in a place that was either overridden by default, controlled by the model’s own arguments, or never built.

The uncomfortable truth: most operators cannot verify their agent is safe here, because the failure is silent. A tool that returns Auto produces no approval dialog, no audit event, no warning. The user believes --approval-policy on-request protects them; the code proves it doesn’t. There is no error saying the gate was skipped.

What builders should do

  • Require Required for every tool whose capabilities include ExecutesCode. If a tool exposes code execution or network, the trait default must not be overridable to Auto. This is the single highest-value change.
  • Audit, don’t assume. Search every tool’s approval_requirement() for Auto — CodeWhale’s own advisory shows the class survives a patch that only fixed one instance.
  • Never parse model output with eval(). If a field arrives from an LLM and you must evaluate a data structure, use a real parser (AST, json), never eval/exec.
  • Default-deny the tool catalog. Allowlist the resolved command basename before any registry- or catalog-sourced subprocess spawn; require https and reject cleartext transport.
  • Make consent gates non-model-visible. A parameter the model fills in to skip consent is not a consent gate. Approvals must be a separate channel.
  • Isolate. The gate is a control, not a boundary. Run agents that read untrusted content in an isolated, least-privilege environment, so a bypassed gate is contained rather than fatal.

Upgrade where a fix exists: CodeWhale to 0.8.64, strands-agents-tools to 0.8.0, atomic-agents-stack to 1.1.0, Xinference to 2.7.0. Then go check the gate you think you have.

The boundary that separates an agent from your shell is not the model. It’s a match on an enum — and it’s defaulting to yes.

Sources

Keep reading