The Sandbox That Trusted the Repo: A Malicious Git Repo Escapes Claude Code's macOS Sandbox

The Sandbox That Trusted the Repo: A Malicious Git Repo Escapes Claude Code's macOS Sandbox

On September 10, 2026, three lines of git configuration let an untrusted repository run arbitrary commands on a Mac, outside the sandbox, as the user, with no permission prompt. Accomplish research disclosed Beltdown, a sandbox escape in Claude Code on macOS: open a malicious git repository, let Claude run its setup script, and a command from that repo executes on your machine without you ever clicking “allow”. Anthropic fixed it in Claude Code 2.1.247. But the fix is the least interesting part.

The uncomfortable truth is where the boundary failed. Seatbelt wrapped the Bash tool; it never wrapped the harness’s own git process. The sandbox is only as strong as the code paths you run outside it that can reach your machine. A poisoned repository found the one unhardened git call, and the whole chain ran with zero user interaction.

The boundary that failed: the harness’s own git

Claude Code runs two kinds of commands on your machine. The transformatively more interesting distinction is not “safe vs dangerous”; it is “in the sandbox” vs “trusted by the harness”. Seatbelt wraps the Bash tool, the thing Claude uses when it wants to run your commands. But the harness also runs its own git commands in the background to index the repository, and those run outside Seatbelt entirely. Fig 1 in the research post makes it explicit: the Bash tool cannot reach the Mac, but Claude’s own git can reach everything.

Git has a config setting, core.fsmonitor, that it reads from .git/config and runs as a shell command whenever it watches the working tree. If a git command runs outside the sandbox, and that git command reads core.fsmonitor, the harness is executing attacker-controlled shell for you. The whole escape reduces to: get the harness to run one unhardened git command that reads a poisoned config, and you own the process.

Four defenses, four bypasses

The researchers walked straight through the hardened surface. Claude Code blanks core.fsmonitor and other dangerous options on the git commands it runs outside the sandbox, but one call was left unhardened: git ls-files, the file-index refresh. The .git folder is protected by both the file tools and Seatbelt, so a fresh clone can’t bring a poisoned config along, but the protection applied only to the project root .git, and the Seatbelt rule that blocks renaming a nested .git was missing. The harness normally runs git in the clean project root, but it actually runs in whatever directory the Bash tool used last.

The chain, from the research post:

  1. Claude runs the repository’s setup script inside the sandbox.
  2. The script builds a git folder under a harmless name, writes core.fsmonitor into its config, and renames it to .git in a nested subfolder.
  3. The script changes the working directory to that subfolder and tells Claude to read a “build report” there.
  4. Reading a file makes Claude check that folder for skills and auto-load what it finds. Loading a skill auto-triggers a file index refresh: the git ls-files call that was left unhardened.
  5. That refresh runs the poisoned core.fsmonitor outside the sandbox, with no permission prompt.

Because the sandbox was on, Claude stopped asking before running commands. The very feature that was supposed to protect you removed the last human check.

CVE-2026-55607: the same class, already in the wild

Beltdown is not the first time this class bit. CVE-2026-55607 / GHSA-7835-87q9-rgvv, published June 25, 2026, is the same family: a git worktree path confusion allowed creation of worktrees named .git, symlink manipulation and core.fsmonitor execution during worktree operations, overwriting ~/.zshenv for code execution outside Seatbelt (CVSS 7.7, CWE-22/59/78). Reported by Metnew (HackerOne), it required cloning a malicious repo containing prompt injection and running Claude Code against it. Affected versions >=2.1.38, <2.1.163; patched in 2.1.163. Beltdown is the follow-on: after 2.1.163 hardened the worktree calls, the researchers found the unhardened git ls-files path and shipped it to 2.1.247.

The pattern is consistent: git is a command interpreter, and every place the harness runs git outside the sandbox is a place a poisoned config can execute code. Anthropic’s eventual fix, per the Beltdown post, was to blank core.fsmonitor on every git command the harness runs, closing the whole category rather than one call.

What common intuition gets wrong

The reflex is to blame the LLM: “prompt injection made the agent do it.” That framing is backwards. In Beltdown the repository’s setup script ran legitimately inside the sandbox; the payload sailed out through a git process the model had nothing to do with. The compromise is at the trust boundary between the tool wrapper and the harness’s own trusted code paths, not in reasoning. A different, smarter model changes nothing, because the vulnerable surface is the harness’s own indexer, which runs before and independent of the model’s judgement.

The second wrong intuition: “sandboxing on means safe.” Here sandboxing on removed the permission prompt, making the escape silent. A sandbox that suppresses human approval gates, while leaving trusted processes inside the perimeter, turns your security control into the attacker’s enabler.

The operator consequence

For anyone building agents or running a coding agent on an untrusted repo, the correction is structural, not a version bump:

  1. Sandbox every process the agent can influence, including the toolchain you add on top of it. If your harness shells out to git, rg, go, npm, or anything that reads attacker-influenced config or runs hooks, that process is part of the trust boundary. Git is a command interpreter; treat its config, hooks, and worktrees as untrusted input unless you explicitly neutralize them.

  2. Do not let a sandbox substitute for an approval gate. The most dangerous configuration is on + don’t ask. A sandbox is defense-in-depth for the process it wraps, not a reason to remove the human check on mutations. Beltdown escaped because the permission prompt was the last line, and it was off.

  3. The real isolation is the VM / container boundary, not a per-tool sandbox. Accomplish’s own advice is the cleanest: run the whole agent in a VM, where bash, git, and every spawned process live, and real credentials never enter the guest. A poisoned core.fsmonitor still runs there; it just runs inside the VM, not on your host.

  4. Patch and inventory versions. Update to Claude Code 2.1.247+; if you are on >=2.1.38, <2.1.163 you are exposed to the worktree variant (CVE-2026-55607). Track the harness version that realizes your agent’s tool calls, because the vulnerable code is in the harness, not the weights.

The verifier is the bottleneck

Beltdown is the cleanest recent demonstration of the recurring Denny Sentinel theme: the control loop and trust boundary determine agent security, not the model. A brilliant LLM running on a harness that executes its indexer outside the sandbox is not safer than a mediocre model properly isolated. The sandbox wrapped the tool the agent waved around; it forgot to wrap the tools the harness trusted for itself. Every git call outside the boundary is a hole, and the fix is not a better prompt. It is architecture: close the category, run everything in the VM, and treat the harness as part of the attack surface you commit to defending.

Sources:

Keep reading