Skip to content

Kerbside Proxy Architecture

This document provides a detailed technical description of Kerbside's proxy architecture: the process model, the connection state machine, traffic handling, and the SPICE firewall.

The SPICE proxy is the Rust kerbside-proxy binary (rust/kerbside-proxy/). The pure-Python kerbside package provides the REST API, the console-source drivers, the data model, and the daemon that supervises the proxy; the proxy itself terminates TLS, drives the SPICE handshake, and relays traffic, consulting the Python side over a gRPC control socket for authorization and bookkeeping. It reuses the ryll shakenfist-spice-protocol crate for the SPICE wire format.

Process Architecture

kerbside daemon run (kerbside/main.py) supervises the proxy as a single child process (subprocess.Popen, via kerbside/proxy_supervisor.py), binding the gRPC control socket first. The proxy is an async tokio program: it accepts connections and handles each on its own task, rather than forking a worker process per connection. Shared state (tokens, channel bookkeeping, audit) lives in the database, reached indirectly through the control-plane gRPC service — the proxy never touches MariaDB directly.

+---------------------------+
|    kerbside-daemon        |  binds the gRPC UDS, then supervises the child
+-------------+-------------+
              | subprocess.Popen
              v
+---------------------------+   gRPC over UDS    +----------------------+
|    kerbside-proxy (Rust)  | <----------------> | KerbsideProxy service |
|    one tokio task / conn  |  authorize, audit, | (in the daemon)       |
+-------------+-------------+  channel records   +----------------------+
              | TLS relay
              v
         hypervisor SPICE

Connection Listener

listen.rs binds two ports (defaults):

  • Secure (5900): TLS. The proxy terminates client TLS with its host certificate, then drives the SPICE link handshake.
  • Insecure (5901): a minimal responder that issues the SPICE need_secured redirect, steering clients to the TLS port. No plaintext SPICE traffic is proxied.

Accepted client sockets get TCP keepalive (mirroring the backend leg), so a silently dead peer is detected rather than pinning a task forever.

TLS Configuration

The proxy holds a host certificate/key (PROXY_HOST_CERT_PATH / PROXY_HOST_CERT_KEY_PATH) and a CA bundle (CACERT_PATH). Client-facing TLS is terminated with the host certificate. The backend leg (backend.rs, via the ryll SpiceClient) originates TLS to the hypervisor. When the console carries a host_subject, the backend leg verifies the server certificate's subject against it using spice-common matching semantics, substituting for hostname verification, so a mis-issued or substituted backend certificate is rejected. When the console has no host_subject, the CA chain is the only identity check. The rustls ring crypto provider is installed at startup.

For Shaken Fist consoles the enforced host_subject is not configured on the proxy: it is pinned at scrape time from the hosting node's published SPICE server certificate subject (spice_server_cert_subject), and carried through to the proxy in the AuthorizeConnection Target. See Console Sources.

Connection State Machine

Each accepted connection runs this sequence (session.rs, backend.rs, relay.rs):

  1. TLS terminate the client connection (secure port).
  2. SPICE link handshake using the ryll server-role handshake drivers: parse the client SpiceLinkMess, reply, and negotiate.
  3. Ticket decryption / authorization. The client's RSA-encrypted ticket is decrypted and sent to the control plane via the AuthorizeConnection RPC, which resolves it to a hypervisor Target (host, port, backend ticket, host_subject, session_id, and the firewall policy) or returns Denied. A denial closes the connection.
  4. Backend connect. The proxy opens the backend leg to the hypervisor's SPICE port (honouring a need_secured retry to the TLS port) and completes the server-side handshake with the backend ticket.
  5. Relay. Traffic is relayed in both directions until either side closes, the idle-read timeout fires, the firewall issues a terminating verdict, or the control plane terminates the session.

The connection registers in the session registry (session.rs) under its session_id after authorization and deregisters on teardown, so the control plane can drop it (see "Session termination" below).

Packet Framing and the Relay

Post-handshake SPICE traffic uses a 6-byte mini-header (MessageHeader: a u16 message type and a u32 size) ahead of each body. The relay (relay.rs) frames every message by that header rather than copying bytes opaquely, which is what makes inspection and the firewall possible. Each direction is a pump; tokio::select! drives both plus a cancellation arm, so whichever future finishes (a closed half, a timeout, a terminating verdict, or a control-plane cancellation) tears the whole connection down cleanly.

Before a body is buffered, an L0 check bounds it (per-channel size caps and an optional rate ceiling), so an attacker-controlled length in the header cannot drive an unbounded allocation. Framed messages are then checked at L1 against the per-channel/direction message-type allowlist. See "the SPICE firewall" below.

Channels

The proxy models the standard SPICE channels: main, display, inputs, cursor, playback/record, usbredir, and the port channel. The L1 allowlist (rust/kerbside-proxy/src/allowlist.rs) is derived per channel and direction from the ryll shakenfist-spice-protocol message-type tables.

Metrics

The proxy exposes a Rust-native Prometheus /metrics endpoint (metrics.rs), bound to --metrics-address (default loopback, config PROMETHEUS_METRICS_ADDRESS) rather than the public VDI address, since the endpoint is unauthenticated. Metrics cover accepted/authorized connections, active connections/channels, bytes relayed per direction, and firewall verdicts.

Error Handling

Handshake, TLS, authorization, and backend-connect failures close the connection and are logged with a per-connection correlation id (a tracing span around the connection lifecycle). A denied authorization, a failed backend TLS verification, an oversized/dis-allowed message, and an idle-read timeout each end the connection deterministically rather than leaking a task.

Security Considerations

Token Validation

Tickets are validated by the control plane, not the proxy: AuthorizeConnection resolves the decrypted ticket against the database and returns a Target or Denied. The proxy enforces the decision and never reaches the database on its own.

Two different tokens exist for a Shaken Fist console, and they live on different surfaces:

  • The HTTP JWT bearer is the Ed25519-signed token Shaken Fist mints and a viewer presents at the /sf-console.vv exchange endpoint on the Python API. The API verifies it offline against the cluster's cached signing public keys (signature, aud, exp) and enforces single use through a jti replay table (sf_token_jtis) — a replayed token is rejected. The Rust proxy never sees this JWT; the exchange happens entirely on the Python side, which then issues an ordinary Kerbside console token and .vv file.
  • The on-wire SPICE ticket is what the proxy actually authorizes. It arrives RSA-encrypted inside the SPICE link handshake, is decrypted, and is resolved by the control plane via AuthorizeConnection. The proxy enforces only this decision; the JWT exchange has already completed before any SPICE connection is made.

Audit Logging

State-changing events (channel created/removed, firewall verdicts, session termination) are recorded as audit events through the control-plane RPCs (RecordAuditEvent, and the channel register/deregister calls), so the audit trail is owned by the Python side and consistent across proxy nodes.

Process and Blast-Radius

The proxy runs as a single unprivileged process consulting a filesystem-guarded UDS (directory 0700, socket 0600); it holds no database credentials and no long-lived secrets. Tickets are short-lived and are never logged. The firewall (below) bounds what an untrusted client can send through to a hypervisor.

The SPICE Firewall

Rather than relaying post-authentication traffic opaquely, the relay (rust/kerbside-proxy/src/relay.rs) is inspection-first: it frames every message by its 6-byte MessageHeader and passes it through a Policy before forwarding. That policy is EnforcingPolicy (policy.rs): a real, enforcing application-level SPICE firewall, on by default.

Per-message pipeline

For every framed message the relay's pump runs:

  1. Header parse. The 6-byte header (message_type: u16, message_size: u32) is parsed first; nothing below inspects the body until this succeeds.
  2. Policy::check_header (L0, pre-body). Called exactly once per message, before its body is buffered, so an over-cap or over-rate message is refused without ever accumulating it:
  3. a per-(channel, direction) size cap — tight (4 KiB) on the inputs-client and cursor-client directions (key/mouse events are small and fixed), generous (16 MiB) everywhere else (the display server bulk direction and any channel with no modeled grammar);
  4. a coarse per-direction rate/throughput ceiling (fixed window), disabled by default — a deployment can opt in once real traffic patterns justify a value.

These policy caps sit below the relay's own unconditional MAX_MESSAGE_SIZE (16 MiB) absolute frame guard, which the relay enforces unconditionally regardless of policy — a peer that claims an absurd message_size is refused before the relay would buffer gigabytes waiting for a body that never arrives. 3. Policy::inspect (L1). Once the full body has been buffered, the message type is classified against a compiled-in per-channel, per-direction allowlist (allowlist.rs), derived from the ryll shakenfist-spice-protocol name tables unioned with the SPICE common-base opcodes (server SPICE_MSG_* 1..=7, client SPICE_MSGC_* 1..=6, valid on every channel). Main, display, inputs, cursor, and playback have their own tables; usbredir/port/webdav ride the spicevmc table. Record, smartcard, and tunnel have no modeled grammar — they classify as ChannelUnmodeled and get L0-only enforcement plus observe-only handling for their message types (never a type-based terminate), rather than treating every type on those channels as a violation. 4. Verdict. Forward (write the original framed bytes unchanged) or Terminate (flush and end the whole relay — a single SPICE channel is one duplex TCP connection, so either direction ending ends the session). Drop is a variant on the Verdict enum reserved for future L2/L3 use (e.g. defanging); no v1 rule emits it, because silently dropping a mid-stream SPICE message would desynchronise the channel.

Warn-only mode

FirewallPolicy.mode (EnforcementMode::Enforce default, or WarnOnly) is the single place the enforcement decision is made (EnforcingPolicy::apply). In Enforce, a rule's blocking verdict is applied and recorded action=enforced. In WarnOnly, the SAME blocking verdict is downgraded to Forward — the session is never actually blocked — but still recorded action=observed plus a tracing::warn!. This lets an operator run real traffic through a deployment and see exactly what Enforce would have tripped before switching to it; it is a mode, not a grace period, since the compiled default ships Enforce.

A forbidden channel type (a whole channel the deployment's FirewallPolicy.permitted_channels excludes) is denied earlier, in session.rs, before any relay is set up — a protocol-correct PermissionDenied to the client plus an audit event, the same shape as the existing unknown-channel-type rejection.

Audit and metrics

Firewall violations must not become one audit event per message — a hostile or broken peer could otherwise flood the audit log. Instead, both relay directions share one lock-free per-connection VerdictTally (atomic counters keyed by (rule, action)); relay::run reads it once after the two pumps finish and, if anything was recorded, emits a single coalesced summary RecordAuditEvent for the whole connection (e.g. "Firewall verdicts this connection: disallowed_type (enforced=2, observed=0)"). Every verdict is also exported live as the Prometheus counter kerbside_proxy_firewall_verdicts_total{channel,direction,rule,action}, where rule is one of disallowed_type, unmodeled_type, size_cap, or rate_cap, and action is enforced or observed.

Policy delivery

Python continues to own policy; the Rust proxy only enforces it ("Python decides policy; Rust enforces it"). The tunable knobs — mode and permitted channels — are delivered per-connection in the AuthorizeConnection gRPC reply as a FirewallPolicy protobuf message, built from Python's FIREWALL_MODE/FIREWALL_PERMITTED_CHANNELS config (see docs/configuration.md). The L1 allowlist tables themselves are not delivered over gRPC: which message types are structurally valid on a channel is a fact about the SPICE protocol, not a deployment policy, so they are compiled into the proxy. Size caps, the rate ceiling, and per-verdict severities keep their compiled defaults in v1 (no gRPC config surface yet).

Out of scope for phase 4, and not yet implemented anywhere in the Rust proxy: L2 body validation (scancode ranges, clipboard/file-transfer/ usbredir device-class filtering), session recording, and L3 rewriting/injection.

Process Supervision and Session Termination

Phase 5 (docs/plans/PLAN-rust-proxy-phase-05-daemon-integration.md) makes the Python daemon able to run the Rust proxy, and makes API-driven session termination actually drop in-flight connections rather than only blocking new ones.

Process model: the daemon supervises the Rust proxy as a child

kerbside daemon run supervises the Rust proxy binary as a child:

+---------------------------+
|    kerbside-daemon        |  main.py:daemon_run
|    (main.py)               |  - binds the gRPC UDS server FIRST
+-------------+-------------+
              |
              | subprocess.Popen
              v
+---------------------------+
|    kerbside-proxy         |  the Rust binary (rust/kerbside-proxy/)
|    (async tokio tasks)    |  - dials the gRPC UDS at startup
+---------------------------+  (ClearNodeChannels) and lazily thereafter
  • The gRPC UDS server is bound before the child is launched, because the Rust proxy dials it at startup and would fail if the socket did not yet exist. (A subprocess child, unlike a fork, does not inherit the gRPC server's fds/threads.)
  • kerbside/proxy_supervisor.py's find_proxy_bin() resolves the binary (env KERBSIDE_PROXY_BINPATH → the in-repo target/{release,debug} dev build dir) and build_proxy_argv() maps config to the proxy's CLI flags (--vdi-address, --secure-port, --cert, --metrics-address, etc.) — the firewall knobs are excluded, since those are delivered per-connection over gRPC (see the firewall section above), not on the command line.
  • The daemon polls the child's liveness every second and forwards SIGTERM with a 15-second deadline before SIGKILLing it; exits non-zero if the child dies unexpectedly. The 15-second deadline is sized above the proxy's own 10-second graceful-drain window (below), so the daemon gives the proxy a chance to drain before forcing it.

How the binary gets there: packaging (phase 6)

find_proxy_bin()'s middle leg — shutil.which('kerbside-proxy') — is what resolves the binary in a real deployment, and phase 6 is what makes that leg succeed. The crate is published to PyPI as a separate kerbside-proxy package: a maturin bindings = "bin" wheel (rust/kerbside-proxy/pyproject.toml) whose compiled binary is laid into the wheel's *.data/scripts/ directory, which pip installs onto PATH. kerbside exact-pins kerbside-proxy at the same version, so pip install kerbside transitively installs a matching proxy and the gRPC contract matches by construction.

The two packages are released in lockstep from a single v* tag: setuptools_scm gives kerbside its version, and tools/stamp-proxy-version.sh stamps that same version into the crate (for maturin) and into the kerbside dependency pin. tools/build-proxy-wheel.sh builds prebuilt manylinux_2_28 wheels for x86_64 and aarch64 (the latter cross-compiled with maturin --zig, so no aarch64 build host is required); no source distribution is published, so an unsupported platform gets a clean pip error rather than a doomed source build. In development you bypass all of this: find_proxy_bin() falls through to the in-repo cargo build output, or you set KERBSIDE_PROXY_BIN explicitly. See docs/plans/PLAN-rust-proxy-phase-06-packaging.md and RELEASE-SETUP.md.

Session termination: dropping in-flight connections

Before phase 5, ConsolesTerminate/SessionTerminate only removed or expired the DB token, which blocks a new connection attempt; a client already connected kept its channels open until it disconnected itself — there was no path from "the API terminated this session" to "the proxy holding that session's sockets finds out".

The distributed-deployment constraint. Kerbside can run with the REST API and the proxy processes on different machines (the proxy is often sized separately from the API), and proxies can sit behind a load balancer so that different channels of the same session land on different proxy nodes. The only component every one of these can reach is the shared MariaDB — there is no API→proxy or proxy→proxy RPC across machines. The ProxyControl gRPC stream, by contrast, is strictly local: one stream between a daemon and the single proxy it supervises on the same host, over the same UDS as every other RPC in the table above. Termination therefore has to be a DB-mediated intent that each node acts on independently for the channels it happens to hold:

REST API                         Proxy node A              Proxy node B
(may be elsewhere)                (holds channels           (holds other
                                    1,2 of session S)         channels of S)
    |                                   |                          |
    | INSERT session_terminations(S)    |                          |
    v                                   |                          |
+----------+                            |                          |
| MariaDB  | <--- polls get_terminations_for_node("A") ------------+
| (only    | <--- polls get_terminations_for_node("B") -------------------+
|  shared  |                            |                          |
|  bus)    |                            |                          |
+----------+                            |                          |
                                         v                          v
                              ProxyControl: TerminateSession(S)   ...same...
                                         |
                                         v
                          Rust proxy: SessionRegistry.terminate(S)
                                         |
                                         v
                    relay::run's select! sees token.cancelled() -> teardown
                    (once per channel this node holds -- here, 2 relays end)

Concretely:

  1. API (api.py, ConsolesTerminate/SessionTerminate): keeps the existing token expire/remove and additionally calls db.request_session_termination(session_id, reason), which upserts a row into the new session_terminations table (session_id primary key, requested_at, optional reason; migration c4e7a1b9d2f3). Idempotent — re-terminating an already-terminated session just refreshes requested_at.
  2. Daemon (servicer.py::ProxyControl, one stream per local proxy): each poll interval, calls db.get_terminations_for_node(NODE_NAME) — an intersection of session_terminations with this node's live proxychannels rows — and yields a TerminateSession(session_id) for every id not already sent on this stream, interleaved with Heartbeats so the stream stays alive between events. A DB error is logged and the loop continues rather than killing the stream. Natural token expiry (no explicit termination) is never pushed — this matches the Python proxy's existing behaviour, where expiry only blocks new connections.
  3. Rust proxy (session.rs::SessionRegistry, rpc.rs:: run_proxy_control): a Mutex<HashMap<session_id, CancellationToken>> refcounted across the session's live channels (one entry per session, incremented on each channel's registration, decremented and removed on the last channel's teardown). Receiving TerminateSession cancels the token; relay.rs's select! gained a third arm on token.cancelled() alongside the two pumps, so every channel of the session this node holds tears down cleanly (no error, same shape as a protocol-level Terminate verdict) and the client is disconnected. Cancelling an unknown or already-torn-down session is a no-op, so a late or duplicate push (another node's poll cycle, a retry) is safe.
  4. Reaper (daemon maintenance loop, reap_session_terminations): deletes session_terminations rows older than 5 minutes — long past the point every node has polled and acted — keeping the table small.

A merely-expired (not explicitly terminated) token is never pushed as a TerminateSession, so natural token expiry does not tear a live session down mid-use — only an explicit API terminate does.

Graceful drain on shutdown

On SIGTERM, the proxy's shutdown_signal future resolving triggers the same teardown path used for termination: it stops the listeners from accepting new connections, then calls SessionRegistry::terminate_all() (cancelling every live session's token at once) and waits up to a 10-second deadline for the active-connection count (the same gauge behind /metrics) to reach zero before returning. The 10-second drain deadline is deliberately shorter than the daemon's 15-second SIGTERM-to-SIGKILL window above it, so a supervised restart lets in-flight sessions finish tearing down on their own terms rather than being cut off mid-drain by the daemon's SIGKILL.

Verifying it live

tools/direct-qemu/VERIFY-TERMINATION.md drives this end to end against a real client: the mock gRPC server's ProxyControl stream emits a one-shot TerminateSession a configurable number of seconds after the first authorization (MOCK_GRPC_TERMINATE_AFTER), standing in for the API→DB→daemon-poll leg so the harness can exercise just the proxy side live. A recorded run against remote-viewer shows all four of its channels (main, display, inputs, cursor) torn down by a single TerminateSession event.

📝 Report an issue with this page