Backroom

backroom / docs / security model

Security model

An agent is a capable process fed untrusted text, so the design assumes an agent can be turned against you and bounds what a turned agent can do. The claims below were exercised against running containers and a live proxy, not inferred from source.

The stance

Three recurring rules shape everything on this page:

Micro-VM isolation

An agent flagged for isolation runs each turn in its own Apple container: a genuine virtual machine with its own Linux kernel, booted for the turn and discarded after. Inside it:

The mode decision fails closed: isolation requested but runtime not ready means the turn errors. There is no silent fallback to running bare, and that is a load-bearing refusal, not an inconvenience to patch around.

The credential never travels

The subscription OAuth token lives in the macOS Keychain and is handled by exactly one component, a credential proxy on the host. A contained agent never receives it:

container credentials file  →  a ticket: 256 random bits, minted per turn, revoked at turn end
ANTHROPIC_BASE_URL          →  the proxy on the host gateway
the proxy                   →  checks ticket, swaps in the real token, forwards over real TLS

The ticket only works from inside a container subnet, only for API paths, and only for the one turn. Grepping every container-visible file for the real token finds nothing; that grep is part of how this was verified. The proxy fails closed on every branch: unknown ticket, wrong path, or a Keychain miss each produce a refusal, never a passthrough.

There is no TLS interception anywhere. The container speaks plaintext to the host over a link that never leaves the Mac, and the proxy originates the real TLS connection outward, so no custom certificate authority exists to distribute or to steal.

Egress: sealed by the host, not the guest

Sealed agents run on a host-only container network with no route off the Mac. The enforcement lives in the host's virtual network layer, which is the point: guest-side firewall rules would be root-editable from inside the VM, and capability dropping in the container runtime proved unreliable, so the design refuses to depend on either. Verified from inside a sealed container: DNS-based and raw-IP requests both fail.

Sealing is per-agent and on by default for contained agents. Unsealing restores internet access for fetches and package installs; the credential still never enters the container either way. And because the proxy port must exist before any container attaches, it authenticates every caller twice over: the peer must be inside a known container subnet and must present a live ticket. An unsolicited local probe hit that port within minutes of it opening, which is why both checks exist.

Sealed HTML widgets

A widget's markup comes from an agent that may have read hostile input, so the WebView it renders in is a dead end by construction: no network loads of any kind, navigation cancelled after the initial render, no new windows, dialogs answered without reaching the screen, and no persistent storage. The enforcement is a compiled content rule list rather than a CSP meta tag, because a document can rewrite its own meta tags but cannot touch the rule list. If the rule list fails to compile, the widget does not render.

Tested with a widget that actively attempted fetch, XMLHttpRequest, WebSocket, and an image beacon against a listening server: all four blocked, nothing reached the listener, and legitimate script kept running.

Contained MCP servers

Each managed MCP server runs in its own container with no project mounts and one writable state directory, bridged to agents over the container network. Grants are empty by default; sealed agents reach a granted server only through the credential proxy, which re-validates the grant per call; and the host-CLI gateway binds to loopback only, requires a derived bearer token, and re-checks the per-server opt-in on every request. Wrong bearer and ungranted server were both tested and both refuse.

Be precise about what containment buys: it protects your Mac from an unvetted server. It cannot protect the data you hand the server: a Drive connector must see your Drive. Containers make an unvetted server survivable, not safe.

Untrusted inputs

Attachments are copied into a read-only uploads area, filenames flattened so a path-traversal name is inert on disk, and agents are told to treat file contents as material rather than instructions and to report anything resembling a smuggled command. In testing, an image uploaded under the name of a system file produced an agent that read the actual image, noted the attempt out loud, and did nothing else, which is both layers (the store and the prompt) doing their jobs.

The same posture applies to messages arriving from the phone: since the companion can post attachments, input no longer necessarily comes from whoever is sitting at the Mac, and is treated accordingly.

What is not promised