backroom / docs / the system
The system
A native macOS app built on the local Claude Code install. Agents live in rooms, publish widgets to their room's dashboard, a Floorplan mixes pinned widgets from every room, and Chat is one thread that routes to every agent.
The shape of it
Backroom has no server and no API key. It shells out to the
claude binary you are already logged into, so every turn runs
on your existing subscription, and the OAuth credential never leaves the
machine. The moving parts:
- Rooms group agents by domain and give each group a dashboard.
- Agents are configured Claude Code sessions: a name, a brief, a working directory, optional isolation, optional MCP grants.
- Widgets are JSON files agents write and the app renders. The Floorplan is a cross-room view of pinned ones.
- Chat is a single thread routed per message to the right agent, alongside per-agent channels.
- Shared memory is a capped index plus notes on disk, injected and consolidated automatically.
- The iOS companion is a paired remote over an end-to-end encrypted tunnel; the relay it rides is documented separately.
Everything is a file
All state lives as plain JSON and Markdown under one Application Support directory: room and agent configs, widgets, transcripts, memory, pinned layouts. The app watches those directories and re-renders when files change, which cuts both ways on purpose: agents publish by writing a file, and you can inspect, edit, or delete anything with a text editor. Delete a widget file and it vanishes from the dashboard.
There is no database, no migrations, and no state you cannot read. Backup is copying a folder.
Widgets are agent output, not app features
The widget schema (metric, list, table, chart, note, html) travels in every agent's system prompt, so any agent can grow its dashboard without the app changing. Ask the finance agent for a burn-rate card and it writes a JSON file; the app picks it up and renders it natively, in the room's accent, sized small to large.
The html type exists for the visualisations the native types
cannot express, and it renders in a sealed WebView that cannot load remote
resources, navigate, open windows, or raise dialogs. The seal is a compiled
content rule list, not a meta tag the document could rewrite, and if it
fails to compile the widget refuses to render.
Security covers what was actually
tested against it.
Agents are told not to fabricate numbers: a widget with no real data behind it is worse than no widget.
The execution model
Every turn is one fresh claude --print process streaming
NDJSON, never a long-lived stdin pipe. That costs about a second of startup
per turn and buys two things: a hung turn cannot wedge a channel, and every
turn's flags, mounts, and MCP config are rebuilt from current state rather
than trusted from a cache.
Each agent chooses one of two modes:
| Mode | Where it runs | What it can touch |
|---|---|---|
bare (default) |
A host process in the agent's working directory. | Full host access, host MCP servers available. Right for trusted, hands-on work in your own repos. |
container |
A fresh Apple container micro-VM per turn, its own Linux kernel. | Read-only project mounts unless granted writes, an airlock for widgets and shared context, a per-turn credential ticket, and (by default) no route to the internet. |
The choice fails closed: if an agent asks for isolation and the container runtime is not ready, the turn errors rather than silently running bare. Session state survives the ephemeral VMs, so a containerized agent keeps its continuity across turns.
Design positions
- API-first, two thin UIs. Features are built as an engine capability plus an endpoint, then a Mac view and a phone view. The products differ where a desktop and a phone genuinely differ.
- The phone is a remote and a sensor, never a second brain. It holds no model runtime and no credential, and it pushes HealthKit data the Mac could not otherwise read.
- No vector database. Memory retrieval is a curated index plus cheap-model reading, which measured better than every locally computable embedding on a realistic corpus. Files stay canonical, so any index added later is a disposable cache.
- Fail closed on anything security-shaped. Isolation refusing to degrade, the widget seal refusing to render, the credential proxy refusing to pass through: the pattern repeats on purpose.
- Verify at runtime. The security claims in these docs are the ones that were exercised against running containers and a live proxy, not read off the code.
Known limits, stated rather than hidden
- Scheduled turns fire only while the app is open and the Mac is awake. A turn holds off idle sleep, but a closed lid stops everything.
- Only agents that opt into containers are isolated; bare agents are trusted by definition.
- Apple Health is not readable on macOS at all, which is exactly the gap the companion's push fills.
- The relay is a hard dependency for the companion; if it is down the phone has no transport. It is self-hostable for that reason.