Signal → prove it matters here → one bounded action → a hard policy → measure → undo if needed
The loop, step by step, with what each partner does and what's live today.
1 Watch
CISA KEV · EPSS · OSV · MongoDBThe watcher polls CISA's Known Exploited Vulnerabilities catalog (the US government's list of flaws attackers are already using) with an ETag, so a new entry is seen within a minute. It's enriched with its EPSS score and written to the advisory store. A MongoDB change stream on that collection starts an agent session: nobody presses a button.
2 Match
the hospital's own inventoryEach entry is matched against the hospital's inventory: the software bill of materials generated from the portal's lockfile, plus its declared estate (help desk, webmail, website, remote access, endpoints). Assets with nothing in front of them that Breakglass can safely change are reported as advisory only, never acted on.
3 Prove it matters here
two factors, both requiredCould it happen here? Semgrep shows the hospital's code actually reaches the flawed part (not just that the library is installed). Is it happening? ClickHouse shows requests on that exact route in the last five minutes. Static evidence alone, or traffic alone, changes nothing; a watch is left instead.
4 Choose one control
a catalog ID, never logicThe agent reads the evidence and the catalog entries for that asset and returns one control ID with reasons. Breakglass's gate then checks the ID exists, is reversible, and touches no critical route, computed from the control's own effect rather than its description.
5 Refuse the drastic
credential proxy · phoneEvery call to the edge goes through a credential proxy. The agent's request to shut a service down matches a DENY rule, so no credential is ever attached. The refusal escalates to a voice call that can approve or roll back.
6 Verify, or undo
exposure + a patient journeyAfter the control, the risky route must show zero unblocked requests and a scripted patient (Playwright) must sign in and book. If either fails, the control is reverted, the failure is recorded against it, and the next candidate is tried. The exposure clock stops only on a verified close.
What each partner does
| Partner | Its job in the loop | Today |
|---|---|---|
| Guild AI | Hosts and runs the agent. Each qualifying flaw starts one session of the published Breakglass agent; its tools are a Guild integration, and Guild's credential policy decides what it may do (shutdown denied, the auto-created allow-all rule deleted). Guild's own records of each run are on the Agent runs page. | Live: the hero run and every recorded portal run from INC-0004 on decided inside Guild (gpt-4.1). If Guild fails, a local loop with the same tools and policy takes over, labelled LOCAL LOOP. |
| Semgrep | Reachability evidence. Custom rules find every server function the portal's code defines and invokes from a route (the flaw's entry point), and a rule over the reverse-proxy config shows when an appliance's risky path is open to the internet. | Semgrep Community Edition (no account), run for real on every incident. |
| ClickHouse | Every request through the edge lands here. One measured query decides whether a route is being hit; verification and the watch left behind are ClickHouse queries too, each shown with ClickHouse's own timing. | Live on ClickHouse Cloud; local ClickHouse, then in-memory, as labelled fallbacks. The scale test adds millions of replayed public log lines, labelled. |
| MongoDB Atlas | The inventory, the control catalog with its outcome counts, and the advisory store. A change stream on new advisories starts each investigation. | Live on Atlas; a local replica set (real change streams), then in-memory, as labelled fallbacks. |
| ElevenLabs | When the policy refuses a drastic action, the IT on-call gets an approval handoff on their phone (BatonPass) with the evidence and an ElevenLabs voice brief, and taps Approve or Roll back. | A Twilio call with an ElevenLabs voice agent is the backup, then SMS. |
| OpenAI | The decision model: gpt-4.1 inside Guild; gpt-6.1-sol in the local loop through the Responses API with structured output. gpt-6-luna is the judge in the agent evals. | A deterministic ranking of the eligible controls if the model is unavailable, labelled DETERMINISTIC. |
What's live and what's simulated
The KEV, EPSS and OSV feeds; matching; Semgrep's scan; ClickHouse queries; the agent's decision; the credential policy; the control applied at the edge; the Playwright journey; the measurements; the phone call.
The hospital. Mercy Valley is fictional; its portal is a harmless twin that reproduces the HTTP surface of a portal stack; its source is a scan-only fixture, so no vulnerable code runs. The "attack" is a benign marker request. Patients are scripted traffic.
When no new KEV entry touches the estate during a demo, a real entry from the same live feed is re-inserted into the advisory stream. Everything after that runs for real, labelled REPLAY.
The hard questions
The safety rules are deliberately deterministic. The agent's job is to gather live evidence, interpret the flaw in this hospital's estate, choose among bounded controls, run the workflow, and verify the outcome. AI decides which permitted move to propose; code decides whether that move is permitted at all.
Guild hosts and runs the agent and is its enforcement boundary. The agent never holds the credential: Guild's proxy attaches it only to calls the policy allows. Catalog controls go through; a shutdown request hits a DENY rule and the request never leaves Guild; drastic operations need a human. That's a security function, not orchestration branding.
It never gets arbitrary change authority. Two independent pieces of evidence first, a catalog of reversible controls only, critical routes never touched, then exposure and a real patient journey measured after the change, with automatic rollback. A real deployment starts in shadow mode and earns write authority control type by control type.
The worst Breakglass can do on its own is turn off an optional feature or add a narrow blocking rule. Booking, sign-in and records are never on the may-disable list, and taking a service offline is denied by the credential policy.