Breakglass
Incidents
Request analytics · ClickHouse Recorded Connecting…

Every decision is a query you can read, with ClickHouse's own timing

Measured on: ClickHouse Cloud (recorded 13:38 PDT). Every request the hospital's edge sees lands in ClickHouse. Breakglass acts only when a query says a risky route is being hit, and closes only when the next query says nothing gets through.

ProblemA rule that blocks the wrong thing takes patient booking down; one that blocks too little leaves the flaw open. Guessing either way is unsafe.
SolutionEach step is a ClickHouse query over the live request log: which exposed route is being hit, did the control stop it, did critical routes stay up.
Why it mattersThe answer has to arrive in milliseconds, at any table size, or the agent acts on stale traffic.
How it's usedRead the decision query, its rows read and its server time; the scale test runs the same query on a table holding millions of real, replayed web requests.

From data to action · INC-0015

Recorded run on ClickHouse Cloud
/portal/messages· 23 requests in 5 min, 23 unblocked→ decision query read 8,207 rows in 9.5 ms server time→ BG-CTL-PORTAL-REQUIRE-SESSION-FOR-ACTIONS→ contained in 1:26 (0 hostile requests through; a scripted patient booked)

Scale context: the table holds 3.57M rows, 3.5M of them replayed real public logs (labelled; never read by a decision).

Edge requests stored
3.57M
73,124 from the hospital's edge · 3.5M replayed public logs
Ingest
250 /min
11,658 in the last hour
Decision query
9.5 ms
in ClickHouse · read 8,207 rows · INC-0015
Reads this session
5
p50 8.57 ms · p95 22.16 ms in ClickHouse

Benchmark · the same decision query on a bigger table

Real public logs, replayed
Now
3.56M rows in the table
Decision query
8.1 ms · read 16,397 rows
Full-table scan
31.4 ms · read 3.56M rows
median of 15 runs, server time · 13:04 PDT

Measured on the ClickHouse Cloud. The decision query filters on the hospital's site, and the table is sorted by (site, route_id, ts), so ClickHouse reads only the hospital's own rows however many other rows the table holds.

replay:nasa-http-1995 · 3.46Mreplay:secrepo-2026 · 39,660

Replayed rows are real logged requests with their real timestamps, under their own replay: site; no Breakglass query reads them. Client addresses aren't stored. Sources: NASA-HTTP (Jul + Aug 1995), Internet Traffic Archive · Security Repo by Mike Sconzo, CC BY 4.0.

The benchmarked decision query
SELECT route_id, client_kind, count() AS n_total, countIf(blocked = 0) AS n_unblocked
FROM edge_requests
WHERE site = 'managed' AND ts >= '2026-10-09 19:59:17.000' AND has(['portal.messages', 'portal.book', 'portal.signin', 'portal.slots'], route_id)
GROUP BY route_id, client_kind ORDER BY route_id, client_kind

Queries that decided an incident

All incidents →
INC-0015✓ CONTAINED9.5 ms in ClickHouse · read 8,207 rows
Hostile traffic on portal.messages → BG-CTL-PORTAL-REQUIRE-SESSION-FOR-ACTIONS
The query
SELECT route_id, client_kind,
count() AS n_total,
countIf(blocked = 0) AS n_unblocked,
countIf(blocked != 0) AS n_blocked,
min(ts) AS first_ts,
max(ts) AS last_ts
FROM bg.edge_requests
WHERE site = 'managed' AND ts >= '2026-10-09 19:56:59.985' AND has(['portal.book', 'portal.messages', 'portal.billpay'], route_id)
GROUP BY route_id, client_kind
ORDER BY route_id, client_kind

What each query drives

Detection → remediation → monitoring
Detection
decision.route_traffic Which exposed routes are being hit right now; decides whether Breakglass may act and which route the control must cover.
—
agent.custom_sql The agent's own read-only query (guarded: single SELECT, readonly=2, 10 s, 10k rows).
—
Remediation
verify.exposure After a control: did any hostile request still get through? A single one reverts the control.
—
verify.critical_routes After a control: any 5xx on sign-in, slots or booking? Patients first; a failure reverts the control.
—
Monitoring
monitor.watch Every armed watch, every few seconds: hostile traffic back on a contained route reopens the incident.
2× · 7.6 ms
timeseries The per-10-second risk line on each incident page.
—
journey_stats Synthetic patient journeys passing over time.
—
Analytics
analytics.table_sizes Rows and bytes stored per table (system.parts).
1× · 3.73 ms
analytics.ingest_rate Edge requests ingested in the last minute and hour.
1× · 8.57 ms
route_traffic Traffic on a route.
—
critical_stats Errors on critical routes.
—

Latest reads

At recording
QueryRows readIn ClickHouseTrip
monitor.watchMonitoring8,2017.6 ms197 ms
monitor.watchMonitoring8,19216.9 ms159 ms

Rows read and server time come from ClickHouse itself (the X-ClickHouse-Summary header); round trip includes the network from this laptop to ClickHouse Cloud (us-west-2).