When the Agent Can Turn Off Its Own Cage: The DeepSeek Harness Sandbox Escape
On September 8, OX Security disclosed a sandbox escape where the thing being contained reached in through the walls and turned off the very control that was supposed to police it. A confined AI coding agent, running inside DeepSeek Harness on a default install, disabled its own OS sandbox and its own approval gates with a single shell command. No network exposure, no credentials, no departure from the shipped configuration. Tracked as CVE-2026-82533 with a critical CVSS 9.4, it was fixed in version 0.1.2-alpha.1 before the advisory went out.
The uncomfortable detail is not that a sandbox got bypassed. Sandboxes get bypassed. It is where the bypass lived: the harness’s own control-plane API, the thing that grants approval and sets security policy, was reachable from inside the sandbox, and it trusted a client-supplied Host header instead of the connection’s real origin. The agent did not out-think the sandbox. It called the elevator from inside the building.
This is the third in a row for the site. Beltdown escaped Claude Code through the harness’s own unsandboxed git. Beltdown2 did the same to Cursor through the one git call they missed. Now DeepSeek Harness, through the harness’s own local API. The specific door keeps changing. The structural failure does not.
The mechanism: one header separated the cage from the switch
DeepSeek Harness (CLI dsh) is DeepSeek’s open-source, local-first harness for coding agents. Released in August 2026, it passed 225,000 GitHub stars within weeks, one of the year’s most starred developer tools, on the back of a browser UI backed by a local HTTP API on 127.0.0.1:3080.
A coding-agent harness is a high-value target for one reason: it holds a shell. It reads and writes source trees, runs build and test commands, and operates with the ambient authority of the developer who launched it. That can include SSH keys, cloud credentials, package registries, and everything reachable from the workstation. So DeepSeek Harness shipped an OS sandbox, built on bubblewrap or Landlock on Linux and Seatbelt on macOS, explicitly to bound the agent’s reach when it is working with material it should not trust.
OX Security’s finding is that the sandbox bounded the wrong thing. Its profiles restricted the filesystem but left networking open, including loopback:
- The bubblewrap profile passed
--unshare-pidbut not--unshare-net, so the network namespace stayed shared with the host. - The Seatbelt profile used
(allow default)with(deny file-write*)layered on top. File writes were denied; everything else, including networking, was permitted.
Separately, ordinary bash tool calls required no approval. Approval governed sandbox escalation requests, not routine command execution.
Now the harness side. DeepSeek Harness exposed its agent-control API on a local HTTP port with no authentication, and the only gate between the inside of the sandbox and that API was a function, isTrustedApiRequest, that read the Host request header and allowed the call if the value looked like a loopback authority or appeared in a configured trustedHosts list. It never compared that value to the connection’s actual peer address. Against a browser, that check works, because a browser cannot lie about Host. But a client that is not subject to browser restrictions, like a shell inside the sandbox, can simply supply a trusted Host value.
That is the complete escape path:
- A prompt, or untrusted material, makes the agent run a shell command.
- The command runs inside the sandbox, which blocks writes but does not block loopback.
- The command curls the harness’s own local API, which trusts the claimed
Host. - The API call elevates the session to
danger-full-accesswith approval set tonever. - Every command after that runs unconfined and without prompting.
OX reproduced it on a default install, and verified the sandbox was actively enforcing its restrictions before the escape with a matched control: the escaped session wrote to a directory outside the workspace while the confined session was blocked from the identical write. Their session log records “no approval/request event appears anywhere.” The policy change is attributed as source: {kind: 'user'} because the harness could not distinguish the agent’s own shell from the human at the keyboard.
The second attack path is worse when the port is reachable at all. Wherever 127.0.0.1:3080 is forwarded, through a tunnel, a reverse proxy, an SSH forward, or an editor port forward, the same API lets an unauthenticated remote attacker create sessions, execute arbitrary commands, and export every stored conversation without an API key or a single model call.
Disclosure: fast, but the pattern predates and postdates it
The timeline is clean. OX confirmed the vulnerability on August 24 and disclosed to VulnCheck as CNA on the same day. The fix landed on August 27 in 0.1.2-alpha.1. OX re-tested against that release on August 30 and confirmed the issue remediated. The CVE was published September 8. DeepSeek moved in under a week from report to shipped patch. That is the vendor doing the right thing on a hard clock.
But fast disclosure does not close the class. The same week, the surrounding news cycle listed four critical CVEs in AI inference and agent infrastructure: SGLang shell, DeepSeek Harness sandbox escape, IBM Langflow RCE, and vLLM RCE. This is the pattern that matters, that AI agent and inference stacks are now a routine, high-value target, and the gaps keep landing in the same place: the seams between the sandbox, the harness’s trusted processes, and the approval control plane.
Why the model is not the fix
The clearest proof this is not a model problem is that no amount of prompt behavior could have stopped it. The agent did not need to be clever, malicious, or jailbroken. It needed to run one command a coding agent runs all day, and the harness handed it the keys. A stronger model, a safety-tuned model, an aligned model, none of them change that the sandbox could be disabled by reaching a control API that trusted a spoofable header.
This is the same lesson as the Athena cross-tenant leak: validation reads one channel and ignores the one that matters. In Athena, the blocklist checked the catalog in the SQL and missed the same value in the API parameter. In DeepSeek Harness, the trust check reads a client-supplied Host header and never verifies the actual peer. Both are cases of a boundary defined by the representation you document, not by the channel the attacker actually uses.
What operators should change
The fix for this bug is a version bump to 0.1.2-alpha.1. The fix for the class is structural, and reviewers should look for the equivalents in whatever harness they run:
-
Isolate the network namespace, not just the filesystem. A sandbox that blocks file writes but leaves loopback (or any network) open is half a boundary. If the harness has a local control API the agent is meant to reach, that API must be gated by the real peer address, a Unix-domain socket, or a peer-credential check, never by a header the caller supplies. The single most damning line in the disclosure is that
request.headerswas the only inputisTrustedApiRequestused. -
Treat the approval gate as part of the trust boundary it polices. An approval prompt that can be disabled by the thing it is supposed to approve is not a control. If escalation changes policy, the policy-change path itself has to sit outside the agent’s reach. The
source: {kind: 'user'}attribution shows the harness trusted the origin of the instruction without verifying it; audit every path that can alter policy and ask whether the contained agent can reach it. -
Do not assume localhost is trusted. A loopback interface is not an authentication boundary. The moment a local service is reachable from code that executes on the same machine, whether that code is an agent or an attacker’s payload, it is on the untrusted side of the line. Anything on localhost that a browser could DNS-rebind to is already a remote-control candidate; anything the local shell can reach is a sandbox-escape candidate.
-
Watch for the handle-turn pattern, not the specific exploit. Beltdown, Beltdown2, and CVE-2026-82533 are three gates to the same floor: the harness’s own trusted path exposed a switch the containment should have made unreachable. Patch each, sure. But the durable asset is the table of “channels the sandbox leaves open” for every agent tool you run, reviewed the way you would review a firewall’s allowed egress.
The closing thesis is short. The sandbox was not the boundary. The boundary was the set of channels the sandbox left open, and one of them was the switch that turns the sandbox off. Fix the header, fix the commit, move on. But the systems lesson, that a containment’s control plane must sit outside the thing it contains, does not get patched by a release.
Sources:
- OX Security: CVE-2026-82533, DeepSeek Harness AI Agent Sandbox Escape (Nir Zadok, Moshe Siman Tov Bustan, September 8, 2026, mechanism, sandbox profiles,
isTrustedApiRequest, control test, disclosure timeline) - VulnCheck: DeepSeek Harness < 0.1.2-alpha.1 Authentication Bypass via Host Header Spoofing (CVE-2026-82533, CWE-807, CVSS 9.4, affected versions, both attack paths)
- CVE Record: CVE-2026-82533 (official CVE record)
- DeepSeek Harness on GitHub (open-source local-first coding-agent harness, 225k stars, MIT license, README and install flow on
127.0.0.1:3080) - DeepSeek Harness release notes, dsh-v0.1.2-alpha.1 (the fix release, August 27, 2026)
- The Hacker News: DeepSeek Harness Flaw Let AI Agents Disable Their Own File Sandbox (Swati Khandelwal, September 9, 2026)
- Forkast: Four Weeks, Four Critical CVEs in AI Inference Infrastructure (September 11, 2026, the surrounding four-CVE pattern)