Every decision is a query you can read, with ClickHouse's own timing
Measured on: ClickHouse Cloud (recorded 13:38 PDT). Every request the hospital's edge sees lands in ClickHouse. Breakglass acts only when a query says a risky route is being hit, and closes only when the next query says nothing gets through.
From data to action · INC-0015
Recorded run on ClickHouse Cloud/portal/messages· 23 requests in 5 min, 23 unblocked→ decision query read 8,207 rows in 9.5 ms server time→ BG-CTL-PORTAL-REQUIRE-SESSION-FOR-ACTIONS→ contained in 1:26 (0 hostile requests through; a scripted patient booked)Scale context: the table holds 3.57M rows, 3.5M of them replayed real public logs (labelled; never read by a decision).
Benchmark · the same decision query on a bigger table
Real public logs, replayed- Decision query
- 8.1 ms · read 16,397 rows
- Full-table scan
- 31.4 ms · read 3.56M rows
Measured on the ClickHouse Cloud. The decision query filters on the hospital's site, and the table is sorted by (site, route_id, ts), so ClickHouse reads only the hospital's own rows however many other rows the table holds.
Replayed rows are real logged requests with their real timestamps, under their own replay: site; no Breakglass query reads them. Client addresses aren't stored. Sources: NASA-HTTP (Jul + Aug 1995), Internet Traffic Archive · Security Repo by Mike Sconzo, CC BY 4.0.
The benchmarked decision query
SELECT route_id, client_kind, count() AS n_total, countIf(blocked = 0) AS n_unblocked FROM edge_requests WHERE site = 'managed' AND ts >= '2026-10-09 19:59:17.000' AND has(['portal.messages', 'portal.book', 'portal.signin', 'portal.slots'], route_id) GROUP BY route_id, client_kind ORDER BY route_id, client_kind
Queries that decided an incident
All incidents →portal.messages → BG-CTL-PORTAL-REQUIRE-SESSION-FOR-ACTIONSThe query
SELECT route_id, client_kind, count() AS n_total, countIf(blocked = 0) AS n_unblocked, countIf(blocked != 0) AS n_blocked, min(ts) AS first_ts, max(ts) AS last_ts FROM bg.edge_requests WHERE site = 'managed' AND ts >= '2026-10-09 19:56:59.985' AND has(['portal.book', 'portal.messages', 'portal.billpay'], route_id) GROUP BY route_id, client_kind ORDER BY route_id, client_kind
What each query drives
Detection → remediation → monitoringdecision.route_traffic Which exposed routes are being hit right now; decides whether Breakglass may act and which route the control must cover.agent.custom_sql The agent's own read-only query (guarded: single SELECT, readonly=2, 10 s, 10k rows).verify.exposure After a control: did any hostile request still get through? A single one reverts the control.verify.critical_routes After a control: any 5xx on sign-in, slots or booking? Patients first; a failure reverts the control.monitor.watch Every armed watch, every few seconds: hostile traffic back on a contained route reopens the incident.timeseries The per-10-second risk line on each incident page.journey_stats Synthetic patient journeys passing over time.analytics.table_sizes Rows and bytes stored per table (system.parts).analytics.ingest_rate Edge requests ingested in the last minute and hour.route_traffic Traffic on a route.critical_stats Errors on critical routes.Latest reads
At recording| Query | Rows read | In ClickHouse | Trip |
|---|---|---|---|
| monitor.watchMonitoring | 8,201 | 7.6 ms | 197 ms |
| monitor.watchMonitoring | 8,192 | 16.9 ms | 159 ms |
Rows read and server time come from ClickHouse itself (the X-ClickHouse-Summary header); round trip includes the network from this laptop to ClickHouse Cloud (us-west-2).