backroom / relay / docs / how it works
How it works
A room holds at most two peers. Both dial out, meet by presenting the same secret room id, and from then on every binary frame one sends is delivered to the other, byte for byte, unread.
Rooms of two
A client opens one WebSocket:
GET /ws?room=<id>
The <id> is a URL-safe token ([A-Za-z0-9_-]),
by default between 16 and 256 characters. It carries all the entropy, so it
must be a real secret: 32 random bytes base64url-encoded is the intended
shape. Both peers derive it from the secret they already share, so nothing
new is ever exchanged through the relay.
The first peer to arrive waits alone. When the second arrives the pair is
complete, the waiting peer receives a peer_joined signal, and
forwarding begins. A third connection presenting the same room id is
refused with WebSocket close code 1008: the pair is complete,
and nobody can displace an established member.
Verbatim forwarding
Once two peers are present, every binary frame one sends is delivered to the other unmodified. The relay does not frame, chunk, reorder within a peer, or transform anything. It also does not buffer for an absent partner: a frame sent while you are alone in the room is dropped, not queued. Reconnect-and-resync is the endpoints' responsibility, and the endpoints are equipped for it because they already treat the network as unreliable.
Binary is mandatory in both directions of the peer traffic. A text frame from a client is a protocol violation and closes the connection. This is what makes control signalling unambiguous: any text frame you receive is from the relay itself and can never be your partner's ciphertext.
Backpressure is handled by refusing to become a mailbox. Each peer has a small send queue (32 frames by default); a peer that cannot drain its queue is treated as dead and the pair is torn down, at which point both ends reconnect and resync. The relay is a pipe, not a store.
Presence signals
The relay sends small text frames to announce presence. They are unencrypted because they carry nothing the relay does not already know from routing. Handle them out of band and never feed them to your decryptor.
| Text frame | Meaning |
|---|---|
{"type":"peer_joined"} |
A partner has arrived, or rearrived, in your room. Sent to the peer that was already waiting. This is the cue to run, or rerun, your end-to-end handshake. |
{"type":"peer_left"} |
Your partner disconnected. Your own connection stays open. |
{"type":"keepalive"} |
Exists purely to be traffic on the wire. Ignore it: do not treat it as a peer event and do not reply. |
Ignore unknown type values rather than failing on them, so a
future control frame does not break an installed client.
On peer_left you are not disconnected. You
become a single-occupant room with no one to route to until the next
peer_joined. The stable end (the Mac) should not churn a
reconnect every time the flaky end (the phone) backgrounds or drops off
cellular. Keep the connection, stop sending, and let the other end
redial. The room id is reusable immediately.
Staying alive: pings and keepalives are different jobs
Two mechanisms run at once, and they are not redundant. Establishing that cost a production debugging session, so it is written down everywhere it matters:
-
Pings prove the peer is alive. The relay pings each peer
every 30 seconds (
RELAY_PING_INTERVAL) and reaps one that misses the 10-second round trip (RELAY_PING_TIMEOUT). A phone that drops off cellular vanishes without a clean close; the missed ping is what removes it. -
Keepalives keep the path open. WebSocket pings are
control frames, and the middleboxes that matter do not count control
frames as activity. Cloudflare drops an idle proxied WebSocket at around
100 seconds even with both ends pinging faithfully; measured, a fully
compliant connection died at 93 seconds of data silence. So the relay
sends a small keepalive data frame every 45 seconds
(
RELAY_KEEPALIVE_INTERVAL) to every peer, including one waiting alone, which is the case that cannot solve itself.
A client must keep a read in flight at all times. Most
WebSocket clients, including coder/websocket and
URLSessionWebSocketTask, only process control frames while a
read is pending. An implementation that writes a request and then waits
without reading gets reaped in about 40 seconds, and the failure looks
like a broken pipe on the next write rather than anything to do with
pings.
When you are refused
| Situation | Response |
|---|---|
| Missing, malformed, or wrong-length room id | 400 Bad Request |
| Third connection to an occupied room | close 1008 (policy violation) |
| Relay at its global room cap | close 1013 (try again later) |
| Per-IP connection rate exceeded | 429 Too Many Requests |
| Frame larger than the size limit | close 1009 (message too big) |
| Text frame sent by a client | connection closed (protocol violation) |
There is also GET /healthz, a liveness probe returning
200 ok plus an X-Relay-Rev header naming the
deployed commit. It reveals nothing about rooms or traffic, and there is
deliberately no endpoint that lists either.