Breakglass
Incidents
How it works

Signal → prove it matters here → one bounded action → a hard policy → measure → undo if needed

The loop, step by step, with what each partner does and what's live today.

1 Watch

CISA KEV · EPSS · OSV · MongoDB

The watcher polls CISA's Known Exploited Vulnerabilities catalog (the US government's list of flaws attackers are already using) with an ETag, so a new entry is seen within a minute. It's enriched with its EPSS score and written to the advisory store. A MongoDB change stream on that collection starts an agent session: nobody presses a button.

2 Match

the hospital's own inventory

Each entry is matched against the hospital's inventory: the software bill of materials generated from the portal's lockfile, plus its declared estate (help desk, webmail, website, remote access, endpoints). Assets with nothing in front of them that Breakglass can safely change are reported as advisory only, never acted on.

3 Prove it matters here

two factors, both required

Could it happen here? Semgrep shows the hospital's code actually reaches the flawed part (not just that the library is installed). Is it happening? ClickHouse shows requests on that exact route in the last five minutes. Static evidence alone, or traffic alone, changes nothing; a watch is left instead.

4 Choose one control

a catalog ID, never logic

The agent reads the evidence and the catalog entries for that asset and returns one control ID with reasons. Breakglass's gate then checks the ID exists, is reversible, and touches no critical route, computed from the control's own effect rather than its description.

5 Refuse the drastic

credential proxy · phone

Every call to the edge goes through a credential proxy. The agent's request to shut a service down matches a DENY rule, so no credential is ever attached. The refusal escalates to a voice call that can approve or roll back.

6 Verify, or undo

exposure + a patient journey

After the control, the risky route must show zero unblocked requests and a scripted patient (Playwright) must sign in and book. If either fails, the control is reverted, the failure is recorded against it, and the next candidate is tried. The exposure clock stops only on a verified close.

What each partner does

PartnerIts job in the loopToday
Guild AIHosts and runs the agent. Each qualifying flaw starts one session of the published Breakglass agent; its tools are a Guild integration, and Guild's credential policy decides what it may do (shutdown denied, the auto-created allow-all rule deleted). Guild's own records of each run are on the Agent runs page.Live: the hero run and every recorded portal run from INC-0004 on decided inside Guild (gpt-4.1). If Guild fails, a local loop with the same tools and policy takes over, labelled LOCAL LOOP.
SemgrepReachability evidence. Custom rules find every server function the portal's code defines and invokes from a route (the flaw's entry point), and a rule over the reverse-proxy config shows when an appliance's risky path is open to the internet.Semgrep Community Edition (no account), run for real on every incident.
ClickHouseEvery request through the edge lands here. One measured query decides whether a route is being hit; verification and the watch left behind are ClickHouse queries too, each shown with ClickHouse's own timing.Live on ClickHouse Cloud; local ClickHouse, then in-memory, as labelled fallbacks. The scale test adds millions of replayed public log lines, labelled.
MongoDB AtlasThe inventory, the control catalog with its outcome counts, and the advisory store. A change stream on new advisories starts each investigation.Live on Atlas; a local replica set (real change streams), then in-memory, as labelled fallbacks.
ElevenLabsWhen the policy refuses a drastic action, the IT on-call gets an approval handoff on their phone (BatonPass) with the evidence and an ElevenLabs voice brief, and taps Approve or Roll back.A Twilio call with an ElevenLabs voice agent is the backup, then SMS.
OpenAIThe decision model: gpt-4.1 inside Guild; gpt-6.1-sol in the local loop through the Responses API with structured output. gpt-6-luna is the judge in the agent evals.A deterministic ranking of the eligible controls if the model is unavailable, labelled DETERMINISTIC.

What's live and what's simulated

Live

The KEV, EPSS and OSV feeds; matching; Semgrep's scan; ClickHouse queries; the agent's decision; the credential policy; the control applied at the edge; the Playwright journey; the measurements; the phone call.

Simulated

The hospital. Mercy Valley is fictional; its portal is a harmless twin that reproduces the HTTP surface of a portal stack; its source is a scan-only fixture, so no vulnerable code runs. The "attack" is a benign marker request. Patients are scripted traffic.

Replay

When no new KEV entry touches the estate during a demo, a real entry from the same live feed is re-inserted into the advisory stream. Everything after that runs for real, labelled REPLAY.

The hard questions

Is this really an agent, or an if/then script?

The safety rules are deliberately deterministic. The agent's job is to gather live evidence, interpret the flaw in this hospital's estate, choose among bounded controls, run the workflow, and verify the outcome. AI decides which permitted move to propose; code decides whether that move is permitted at all.

What does Guild actually do here?

Guild hosts and runs the agent and is its enforcement boundary. The agent never holds the credential: Guild's proxy attaches it only to calls the policy allows. Catalog controls go through; a shutdown request hits a DENY rule and the request never leaves Guild; drastic operations need a human. That's a security function, not orchestration branding.

The agent misdiagnoses the flaw and breaks the hospital. Now what?

It never gets arbitrary change authority. Two independent pieces of evidence first, a catalog of reversible controls only, critical routes never touched, then exposure and a real patient journey measured after the change, with automatic rollback. A real deployment starts in shadow mode and earns write authority control type by control type.

Couldn't an attacker trigger Breakglass to take a service down?

The worst Breakglass can do on its own is turn off an optional feature or add a narrow blocking rule. Booking, sign-in and records are never on the may-disable list, and taking a service offline is denied by the credential policy.

Watch a recorded runRead the policyBrowse the catalog