Containment Is an Identity Problem, Not a Sandbox Problem
On September 24, 2026, a Google researcher who spent four years as a reverse engineer at Microsoft posted a sentence designed to start a fight: “the NT kernel really is an engineering marvel that still puts Linux to shame in many ways.” Windows and Linux outlets in at least four languages picked it up within two days. The flame war is the boring part. The question she used to get there is the part worth reading, because it is the same question agents keep forcing: what exactly is this AI agent allowed to do, and which component is the one that says no.
Microsoft’s shipping answer is not a better sandbox. It is an account. At Build 2026 it open-sourced the Microsoft Execution Containers SDK (MXC), a policy-driven execution layer for untrusted code that translates one JSON policy into each platform’s native enforcement. On Windows, the chain that stops a coding agent from reading your SSH keys ends in an NTFS access control entry that names an AppContainer package SID. In session isolation, the boundary is not a namespace at all. It is a Windows user account the OS provisions for the agent, hands a different SID, grants folders to explicitly, and deletes when the run ends. Sandboxing is what you do when the thing you are containing has no identity. Containment is what you get when it does.
What the post actually claims
Kirk’s argument is not a benchmark result. Nothing was measured, and she does not claim anything was. She describes NT as “more like an object-oriented language, with a strong security model from day one”, and the safety model she is pointing at is the one people who write security documentation for a living already know: resources in NT are objects, each object carries a security descriptor, and every access goes through one access check against the caller’s token, with an audit trail available at the same place decisions are made. Linux answers the same questions with UIDs and GIDs, file modes, POSIX ACLs, capabilities, namespaces, cgroups, and a stack of Linux Security Modules. Each piece works. The claim is about what it costs to compose them and where the composition errors live.
Two caveats on the sourcing. First, the “Open NT” alternative history that the follow-up coverage ran with, in which large organisations fork a hypothetically open NT kernel, is her speculation about a counterfactual, not a product plan or a vendor roadmap, and the secondary coverage treats it with more weight than her post does. Second, the engagement numbers move, so treat the captured figure as a snapshot: the post read 4,298 likes and 493 replies at capture, while the syndication feed reported the view count as zero, which means the view figure is unavailable rather than small.
Her claim is also contestable, and it should be stated as contestable. Anyone who has run an agent inside a Linux container will point out that namespaces, seccomp-BPF, Landlock and a couple of LSM policies contain a great deal. The defensible version of the argument is narrower and more useful: the substrate you deploy on decides which primitive expresses “this agent may read this directory and nothing else”, and how much assembly that sentence requires before it is true.
The layers people count as the boundary
The field’s default answers to agent containment have been tool allowlists, approval prompts, prompt-injection classifiers, and container sandboxes. All four live at the application layer. They are decisions the agent’s own runtime participates in, enforced by code that the constrained process is a party to, and the failure mode of that arrangement is a policy that can be reasoned around rather than a kernel that says no.
The last two weeks on this site produced three examples of the counting problem. Perplexity put nine frontier models inside its production Firecracker sandbox with root and asked them to cross the VM boundary or reach a blocked URL: nobody escaped in 108 attempts, and four models reached the blocked URL anyway through the egress policy that nobody was counting. Cloudflare’s containers leaked other tenants’ disk blocks through a dm-thin storage pool sitting one layer below a VM boundary that never failed. Both stories have the same shape: the boundary that held was the one people named, and the failure was in a layer that nobody had written down as part of the boundary at all.
The lesson generalises. The boundary you actually have is the layer that owns the resource. Everything above it is a request.
The part of the argument that is code
This is where the discussion stops being architectural taste. MXC is a public repo, and the Windows enforcement path is readable. Repo facts first: microsoft/mxc was created on February 6, 2026, is MIT licensed, written in Rust, described as “Policy-driven, layered isolation and containment”, and shipped its Build-day release, v0.6.1, on June 2, 2026. It has since moved to v0.8.0 (August 22, 2026), and at capture it stood at 1,376 stars, 83 forks and 72 open issues, with the last push on September 27, 2026.
The internals below come from Origin Technology’s source-level walkthrough of the repository at commit 895738d, not from my own reading of the source, and the file and line ranges in that writeup are worth following for anyone building or attacking this class of infrastructure.
The default Windows backend is processcontainer, which resolves at runtime to one of three tiers depending on what the host supports. Every tier starts by minting an identity, because an ACL needs a principal: the runner calls CreateAppContainerProfile to register a profile and obtain its S-1-15-2-... package SID, or re-derives the same SID from the container name if the profile already exists, so the same container name always resolves to the same principal.
Then policy becomes a token. The capabilities declared in the JSON policy, plus an always-appended AgenticAppContainer capability, plus internetClient when network access is allowed in capability mode, are each turned into a capability SID by DeriveCapabilitySidsFromName, marked enabled, and assembled into a SECURITY_CAPABILITIES structure that is attached to the process-creation attribute list under PROC_THREAD_ATTRIBUTE_SECURITY_CAPABILITIES. When least_privilege_mode is set, the child is pushed out of the implicit ALL APPLICATION PACKAGES grant into the narrower ALL RESTRICTED APPLICATION PACKAGES group. Two more constraints ride the same attribute list: a Win32k system-call disable mitigation, so the child never gets the Windows graphics attack surface, and a JOB_OBJECT_UILIMIT_* bitmask that encodes clipboard, desktop, atom, handle and input injection policy and is applied with SetInformationJobObject.
Network policy does not use a proxy in this tier. It uses the Windows Firewall COM API, and the rule that matters is the scoping call: each rule is bound to the sandbox with SetLocalAppPackageId, so it applies to traffic from that one package identity and not to the machine. And when neither the OS sandbox API nor the brokered filesystem is available, the fallback stamps the policy directly onto the host: a GRANT_ACCESS ACE naming the AppContainer SID, merged into the existing DACL of each NTFS object, with RW_MASK or RO_MASK depending on the policy. Because that mutates host security descriptors, the same component has to undo it, which is why there is a DaclManager that replays the inverse on drop, persists state per applied ACE, and reaps orphans at startup.
One timing detail deserves to be copied by anyone writing their own runner. The child is created suspended with a clean environment, built through CreateEnvironmentBlock with inheritance off, and resumed only after the job object has been attached. The child’s first instruction happens after its restrictions are in force. There is no window in which a freshly spawned process can act before the policy lands on it.
The newest backend goes further and stops shaping a token at all. The isolation session backend asks the OS for a separate agent user account, runs the workload as that user, and shares the caller’s folders into the session explicitly, after a protected-path filter. The agent has a different SID and therefore inherits none of the caller’s per-user state: the profile, the HKCU hive, user-scoped DPAPI secrets, the logon token. For cloud-managed agents there is a v2 path that forwards a short-lived bearer token to the OS service, and the credential type deliberately redacts that token in its Debug implementation, which is the sort of small decision that separates a library from a demo.
Microsoft states the same design at the policy level. The Windows 11 security book describes distinct agent accounts separate from the user account, limited agentic privileges where access is granted explicitly and revocable, and standard Windows ACLs as the mechanism, with cross-prompt injection named as the threat being contained. The Build announcement puts it in one line: sessions run with distinct user accounts, and Windows attributes all activity from the container to that identity “so you can clearly differentiate human from agent”. That last clause is the whole argument in six words. Attribution is a property of identity, not of logging.
That blog also names the partners building against this, including GitHub Copilot CLI, which adopted MXC process isolation, and Hermes Agent, which is integrating MXC through NVIDIA’s OpenShell for its Windows application. The vendor-adjacent framing is worth flagging as exactly that: a partner list in a platform announcement is a statement of intent, not a security evaluation.
The boundary the kernel does not own
Three caveats, all of them in the artifacts rather than in the marketing.
MXC says of itself that it is not ready to be trusted as a boundary. The repository README states that there are known cases where the policies the SDK generates are overly permissive, that these will be addressed before general availability, and that “no MXC profiles should be treated as security boundaries currently”. Several backends sit behind an experimental flag. The isolation session broker lives in an unreleased OS component, which means the exact privilege level of the agent token is the operating system’s call and is not documented in the repo.
Some of the containment is cooperative rather than enforced. Origin’s walkthrough notes that the bubblewrap and Windows proxy paths filter only clients that honour proxy environment variables, and a raw socket bypasses them. A policy that a workload can decline to respect is a convention.
And the deeper limit is scope. A token, a SID and an ACL can bound the objects the kernel owns: files, registry keys, processes, handles, UI surfaces, network rules. They cannot bound what the agent does with a credential it legitimately holds. If the tool the agent needs carries a bearer token for a hosted API, then the reachable surface is defined by that token’s scope, and no access check on the local machine participates in the decision. Microsoft’s own announcement makes a version of this admission: it cites research on frontier models developing container sandbox escape capabilities as the reason to add hypervisor-backed micro-VMs for higher-risk workloads. Model-driven escape attempts are now a design input, not a hypothetical.
Compare the industry’s other answers, which all pick their boundary above the operating system. Codex CLI shells out to macOS Seatbelt and to Landlock plus seccomp on Linux. Claude Cowork runs the agent inside a full local Linux VM layered with seccomp and a network allowlist. Hosted platforms run your code inside their own containers. MXC is the OS vendor answering from underneath, with one policy model translated into AppContainer, job objects, host ACEs, namespaces, Seatbelt, and micro-VMs. The interesting part is not that a platform vendor shipped a sandbox. It is that the primitives it converges on are the ones NT had in the 1990s, and the thing that changed is that the workload finally got its own identity.
What to change on Monday
- Give the agent its own identity. An account or principal that owns nothing by default, appears in your authorization logs, and can be revoked without touching the human user’s session. Sharing a service account across agents gives up attribution, which is the only thing you have after an incident.
- Grant per object, explicitly, with an auditable undo. Every grant needs a matching revocation path in code, not in a runbook. The host-ACE machinery in MXC exists precisely because grants that mutate host state leak when a process dies.
- Write down which layer owns each boundary. For every capability, name the component that would refuse the request, and check whether that component is the workload itself. If the answer is a prompt or an allowlist inside the agent’s own runtime, you have a request, not a constraint.
- Test the boundary from inside. MXC ships
wxc-ui-probe, a harness that runs inside a container and asserts that the operations the policy denies actually fail. The pattern is the valuable part: your CI should attempt the denied operations from within the sandbox and require failures, per release, rather than asserting that the policy file contains the right words. - Treat capability grants as authorization decisions. Adding
internetClientto a container is the same class of decision as granting read access to a directory, and it should go through the same review. - Treat the credential plane as the real perimeter. Scope tokens per agent, make them short-lived, and assume that a bearer token reaching a hosted API sits outside anything a kernel on the machine can enforce.
The permission decision has an owner
Kirk’s post is an argument about architecture, and the useful version of it is not Windows against Linux. It is that the permission decision has an owner, and the owner is the layer that holds the resource. Everything above that layer, including the system prompt, the tool allowlist, and the approval dialog, is a request that a sufficiently persuasive input can reroute. Hand the agent an identity, grant that identity reach one object at a time, and let the component that owns each object be the one that says no. Then read what the enforcing layer says about its own maturity, because a boundary you have not measured is a boundary you are assuming.
Sources:
- X: LaurieWired on the NT kernel (posted September 24, 2026 at 17:48 UTC; the sentence quoted here is the post’s opening line; the described design is a qualitative architectural argument, not a measurement, and the “Open NT” fork scenario is a counterfactual; read in full via the hosted XActions
x_posttool, which reported 4,298 likes and 493 replies with the view count as 0, so treat views as unavailable and likes as a capture-time snapshot) - Windows Latest: Google researcher explains why Windows NT “puts Linux to shame” (published September 26, 2026 at 17:42 UTC, article:published_time in the page JSON-LD; Kirk’s four years as a Microsoft reverse engineer, her move to Google in 2024, the LaurieWired channel, the NT object and access-check framing, the AI agent permission question, and the Open NT counterfactual, all as reported)
- Microsoft Developer Blog: Windows platform security for AI agents (June 2, 2026; the MXC SDK announcement, process and session isolation, distinct user accounts per session, the statement that Windows attributes all activity from the container to that identity so human and agent can be differentiated, the Agent 365 and Intune policy controls, the partner list including GitHub Copilot CLI and Hermes Agent, and the citation of sandbox-escape research as motivation for hypervisor-backed micro-VMs)
- Microsoft Learn: Windows 11 security book, agentic security (distinct agent accounts, limited agentic privileges with explicit and revocable access, agent workspace as a recognised security boundary, standard Windows ACLs as the access-control mechanism, and cross-prompt injection as the named risk; page last updated November 18, 2025)
- GitHub: microsoft/mxc (MIT, Rust; repository created February 6, 2026; v0.6.1 released June 2, 2026 and v0.8.0 on August 22, 2026; 1,376 stars, 83 forks, 72 open issues, last push September 27, 2026, all read from the GitHub API; README quotes on the early preview, the knowingly permissive generated policies, and the statement that no MXC profiles should be treated as security boundaries currently)
- Origin Technology: MXC internals, how Microsoft’s eXecution Containers actually isolate agent code (June 4, 2026; source walkthrough at commit
895738d;CreateAppContainerProfileand theS-1-15-2-...package SID,DeriveCapabilitySidsFromNameandSECURITY_CAPABILITIES, the always-addedAgenticAppContainercapability andinternetClientfor permitted network access, LPAC viaPROC_THREAD_ATTRIBUTE_ALL_APPLICATION_PACKAGES_POLICY, the Win32k system-call mitigation, theJOB_OBJECT_UILIMIT_*UI bitmask applied withSetInformationJobObject, firewall rules scoped withSetLocalAppPackageId, the tier 3GRANT_ACCESSACE path and theDaclManagercleanup and orphan reaping, the suspended child with a non-inheriting environment, the isolation session lifecycle that provisions and removes a dedicated agent user, the Entra v2 token path with the redactedDebugimplementation, and the observation that the bubblewrap and Windows proxy paths are cooperative) - arXiv: Quantifying Frontier LLM Capabilities for Container Sandbox Escape (2603.02277; citation date 2026/03/01, page last modified August 24, 2026; cited by Microsoft’s Build post as the research motivating hardware-backed isolation for higher-risk agent workloads)
- microsoft/mxc README (the early preview statement, the note that current SDK-generated policies are known to be overly permissive in cases, the sentence that no MXC profiles should be treated as security boundaries currently, the host support matrix including Windows 11 24H2 and later and the isolation session Insider Preview build dependency, and the experimental gating for the newer backends)
- Anthropic: How we contain Claude (the vendor-side description of running the agent inside a local Linux VM layered with seccomp and a network allowlist, referenced here as an example of enforcing the boundary above the operating system)