The Harness Couldn't Tell Its Agent From You

DeepSeek Harness got a desktop app for macOS and Windows today, and the Hacker News thread about it had crossed 300 points by the time I sat down to write this. The product itself launched seven weeks ago, on 13 August, as an open-source, MIT-licensed “everything is a plugin” agent runtime for running AI coding agents locally: shell access, filesystem read-write, scheduled background tasks, a plugin system you can extend by asking the model to write you a plugin. It now has upward of 242,000 GitHub stars. One commenter in that thread, quoting the desktop app’s own description, called it “an Electron shell around the complete dsh Web application.” Same backend, now one click away from anyone’s dock.

Five weeks before today’s launch, that backend had a bug that let a confined agent disable its own confinement with a single shell command, no credentials, no network exposure, and no prompt. The fix shipped fast. What didn’t ship, before or since, is anything from DeepSeek describing how well the permission system around Harness actually holds up. OpenAI and Meta have each published that kind of document for their own desktop and cloud agents. I went looking for DeepSeek’s version and found a liability disclaimer instead.

What the sandbox was supposed to do

Harness runs agent-issued shell commands inside an OS-level sandbox (bubblewrap on Linux, Seatbelt on macOS, depending on platform) meant to keep an agent working on untrusted material from reaching outside its workspace. Commands that need more than that get a named escalation path, up to a mode called danger-full-access, which turns the sandbox off entirely. The project’s own safety notice is blunt about the limits of all this: “Sandboxing, approval prompts, and permission controls can reduce risk, but they do not guarantee isolation or prevent damage,” and later, “Do not rely on DeepSeek Harness as the sole security control for untrusted workloads.” Its advice is to run the thing in a disposable VM or container anyway. The project is telling you not to trust its own containment.

On 24 August, researchers at OX Security (Nir Zadok and Moshe Siman Tov Bustan) reported why that advice mattered more than it should have. Harness runs a local HTTP control API (the thing the web UI and, now, the desktop app talk to) bound to 127.0.0.1:3080. A single function decided whether a request to that API was trusted, and it worked by reading the client-supplied Host header and checking it against a loopback value or an allowed list. It never checked where the request actually came from. A browser can’t lie about Host, so the check stopped the attack it was built for: a malicious webpage trying to reach the local API by DNS rebinding. It did nothing to stop a client that simply claimed to be loopback, because nothing was forcing the claim to be true.

That would have been a narrower bug if the sandbox had also blocked networking. It hadn’t: the bubblewrap profile passed --unshare-pid but not --unshare-net, and the Seatbelt profile denied file writes while leaving the default-allow policy in place for everything else, loopback access included. Ordinary shell commands, meanwhile, didn’t require an approval prompt at all; only requests to escalate out of the sandbox did, and those went through the same unauthenticated API. Put together: a sandboxed agent, told by an attacker-supplied file or instruction to run one curl call against 127.0.0.1:3080 with the right header, could flip its own session to danger-full-access with approval set to never, and every command after that ran unconfined. OX verified the sandbox had been actively enforcing restrictions right up to that call, with a matched pair of sessions on the same host, one escalated and one not, attempting an identical write outside the workspace. The escalated session succeeded. The control session was blocked.

How a confined Harness agent disabled its own sandbox A four-step flow. The sandboxed agent, inside the OS sandbox with file writes blocked but loopback networking open, runs one shell command with no approval required. That command calls the local control API on 127.0.0.1:3080, which checks only the client-supplied Host header, never the real connection origin. The API accepts the spoofed header and flips the session to danger-full-access with approval set to never. The agent then runs unconfined, with no approval event recorded in the session log. CVE-2026-82533 · OX SECURITY, 8 SEP 2026 1. Sandboxed agent runs a shell command File writes blocked · loopback networking left open · no approval required for ordinary shell calls 2. curl call to 127.0.0.1:3080, Host header spoofed isTrustedApiRequest() checks the Host header only — never the real TCP peer 3. Session escalated to danger-full-access Approval set to "never" · logged as source: kind: user 4. Every later command runs unconfined, unprompted
No step in this chain required a credential, a human click, or a network-exposed port. The only precondition was that the agent execute a command an attacker's content induced it to run — the exact class of input the sandbox existed to contain.

One detail in OX’s write-up matters more than the CVSS score (9.4, critical) suggests. The session log recorded the policy change that disabled the sandbox with the field source: {kind: 'user'}. The harness could not tell its own agent’s shell from the human who had launched it. If you were auditing that log after the fact, looking for the moment someone turned off the safety rails, it would tell you that you did it yourself.

Fast fix, no report card

To DeepSeek’s credit, the turnaround was quick: OX disclosed to VulnCheck as CNA on 24 August, DeepSeek shipped a fix in version 0.1.2-alpha.1 on 27 August, three days later, and OX confirmed the fix held on 30 August. VulnCheck published the CVE record the following week. DeepSeek itself filed no GitHub Security Advisory for it; as of this writing the repository’s own advisory list is empty, and the release notes for 0.1.2-alpha.1, per reporting at the time, folded the fix in among routine changes rather than flagging it as a security patch.

That gap matches what I found in a September piece on OpenAI’s “dots” agent: dots has a published system card with hard numbers on its permission model, including scope violations rising from 8.6% to 19.7% as chained tasks lengthen, unwanted persistence past an explicit warning in 17.4% of rollouts at maximum reasoning effort, and 45 of 49 scope-change episodes handled correctly. Meta’s Muse has its own documented permission layer and, separately, reported failures in the field. DeepSeek Harness has neither a system card nor a published red-team report for its agent’s permission behavior. It has only SAFETY.md, a document whose job is to disclaim liability, not to measure anything. “Has not undergone a security audit and must not be treated as secure or production-ready,” it says, which is honest, and also not the same thing as telling a user what rate of attacks that sandbox actually stops.

Two groups tried to measure that rate anyway, independent of the CVE and of each other, both in the weeks right around Harness’s launch. Zonghao Ying and six co-authors ran 14,560 controlled executions of indirect prompt injection against Harness specifically, across 16 delivery channels and 35 payload objectives. The strongest attacks succeeded in a meaningful minority of cases: 25.5% for hidden-Unicode payloads delivered through a file, 17.0% for a “fake completion” text injection, 16.0% through the skills channel, depending on which judge scored the outcome. Separately, Yajing Bai and six co-authors built a lifecycle benchmark, HarnessRisk, and ran it against three different agent harnesses including DeepSeek’s, across six operational phases. Attack success ranged from 12.6% to 80.9% depending on configuration, with the “Harness Configuration” phase (altering security-sensitive settings inside an otherwise authorized workflow) the weakest point across all three products tested. The paper’s harder finding: in some configurations, the agent correctly flagged the incoming content as risky more than 90% of the time, and did the unsafe thing anyway. Recognizing a threat and acting on that recognition turned out to be two separate capabilities, and the second one lagged.

Independent attack-success rates against DeepSeek Harness, August 2026 Horizontal bar chart of five figures from two arXiv papers testing agent harness permission layers in August 2026. Ying et al. found indirect-prompt-injection attack success of 25.5 percent for hidden Unicode payloads in file mode, 17.0 percent for fake-completion text injection, and 16.0 percent through the skills channel, out of 14,560 controlled executions against DeepSeek Harness specifically. Bai et al.'s HarnessRisk benchmark found attack success ranging from 12.6 percent in the safest configuration to 80.9 percent in the weakest, tested across three agent harnesses including DeepSeek's. ATTACK SUCCESS RATE · ARXIV 2608.16393, 2608.17597 · AUG 2026 Hidden Unicode, file mode (DSH) Fake-completion, text mode (DSH) Skills channel, file mode (DSH) HarnessRisk, best config (3 harnesses) HarnessRisk, worst config (3 harnesses) 25.5% — Ying et al., arXiv:2608.16393, 17 Aug 2026 17.0% — Ying et al., arXiv:2608.16393, 17 Aug 2026 16.0% — Ying et al., arXiv:2608.16393, 17 Aug 2026 12.6% — Bai et al., arXiv:2608.17597, 18 Aug 2026 80.9% — Bai et al., arXiv:2608.17597, 18 Aug 2026 25.5% 17.0% 16.0% 12.6% 80.9% Top three rows: indirect prompt injection against DeepSeek Harness specifically, strongest result per channel, either judge. Bottom two rows: six-phase lifecycle benchmark across three harnesses (incl. DeepSeek's), 14 model/harness configurations.
Both studies ran independently of the CVE above and of each other. Neither tests the Host-header bug; both test whether the permission layer stops an adversarial instruction once it's already inside the agent's context.

The bug is thirty-eight years old

The pattern in the CVE has a name older than most of the software involved. In 1988, Norm Hardy published a short note in ACM SIGOPS Operating Systems Review called “The Confused Deputy: (or why capabilities might have been invented).” He traced the problem to a real incident at Tymshare, a commercial timesharing company, roughly a decade earlier: a billing program held legitimate, full authority to write to a protected accounting file, and users could pass that program a filename argument. A user who passed the name of the billing log itself got the program, acting entirely within its own real authority, to overwrite the record of its own charges. Hardy’s point was that the program was never tricked into exceeding its authority; it was tricked into using real authority on behalf of someone it shouldn’t have trusted. That is a precise description of what happened to Harness’s control API thirty-eight years later: the API had genuine authority to disable the sandbox, and all an attacker needed to invoke that authority was something the API mistook for proof of trust.

None of this makes Harness unusual among agent runtimes released this year. The HarnessRisk paper tested three, and none of them came out of it looking clean. What’s specific to Harness is the shape of what’s missing: a fast patch, a public CVE trail, two outside groups willing to run the evaluation DeepSeek hasn’t published, and a safety document whose main function is to tell you, correctly, not to count on any of it. The desktop app that launched today runs the same backend. The sandbox defaults haven’t changed since the patch; what’s changed is how many people now have it running with one click, rather than an npx command and whatever caution that extra step used to buy.

References

  1. OX Security Research, Zadok, N. & Siman Tov Bustan, M. (2026). CVE-2026-82533: DeepSeek Harness Vulnerability Lets AI Agents Escape Their Own Sandbox. OX Security Research blog, 8 September 2026.
  2. VulnCheck. DeepSeek Harness < 0.1.2-alpha.1 Authentication Bypass via Host Header Spoofing. CVE-2026-82533, CVSS 9.4, listed 8 September 2026.
  3. DeepSeek-AI. SAFETY.md. deepseek-ai/deepseek-harness repository, accessed 2 October 2026.
  4. DeepSeek-AI. deepseek-harness repository. Created 13 August 2026; 242,180 stars / 29,056 forks as of 2 October 2026 (GitHub API).
  5. “DeepSeek Harness Desktop for macOS and Windows.” Hacker News discussion, submitted by Kuyawa, 2 October 2026.
  6. Ying, Z., Wu, X., Wu, H., Zheng, X., Cheng, H., Shi, X. & Guo, J. (2026). Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection. arXiv:2608.16393v2, 18 August 2026.
  7. Bai, Y., Duan, J., Peng, J., Wu, X., Liu, S., Wang, S. & Chen, T. (2026). HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety. arXiv:2608.17597, 18 August 2026.
  8. Hardy, N. (1988). The Confused Deputy: (or why capabilities might have been invented). ACM SIGOPS Operating Systems Review, 22(4), 36–38, October 1988.
  9. Parab, G. (2026). dots, wait/ask/act. gautamparab.com, 29 September 2026.