Checklist: YSD-22058 (recovery) · drill kind recovery Decision owner
role: the on-call engineer for svc-study-workspace Machinery:
apps/yemaya/svc-study-workspace/src/persistence/chaos.integration.spec.ts
There was a real plan for this — detection, three steps with owners and durations, two audiences and a 60-minute objective — written as a fixture inside a spec, to exercise the plan assessor. It and a key-compromise plan in the same file were the only continuity plans in the workspace, and a plan that exists to be graded is not a plan anybody can follow at 3am.
Running this document IS the drill. Record it in
procedure-exercise-register.json.
What the system does on its own — read this first#
Most of this outage is already handled, and knowing that changes what you should do. Driven on every change against real PostgreSQL and object storage:
- The saga waits durably on the dead store and heals when it returns. It does not fail, and it does not lose the work. Do not "clear the queue".
- A redelivered outbox row re-processes idempotently, so a store that comes back and replays does not double-apply.
- A step never acks while a stale copy is still served — a failed cache invalidation holds the step rather than declaring success.
- A partial delete resumes exactly once after a crash: bytes gone, records intact, and the resume finishes the job rather than restarting it.
⚠️ The main way to turn this outage into data loss is to intervene. Restarting workers, draining outboxes by hand, or re-running a deletion because it "looked stuck" all fight machinery designed to wait. The default action is to restore the store and let the saga heal.
Detection#
- Readiness reports the store unavailable (
readiness.ts). - Outbox depth climbing without an error rate to match — the signature of durable waiting, not of failure.
- Not "a customer tells us": by the time a learner notices, the queue has been waiting for some time.
The steps#
| # | Action | Owner role | Expected |
|---|---|---|---|
| 1 | Confirm which store is gone and that the queue is waiting rather than erroring | on-call engineer | 5 min |
| 2 | Restore the store, or fail over to its replica | on-call engineer | 30 min |
| 3 | Watch the outbox drain and confirm the depth returns to baseline | on-call engineer | 20 min |
| 4 | Reconcile: confirm no step acked while the store was gone | on-call engineer | 10 min |
Recovery time objective: 75 minutes. Recovery point objective: 0 — the saga waits rather than dropping, so an outage inside the window should lose nothing.
Why the number changed on 2026-08-15#
It read sixty minutes until then, and the four steps above are budgeted at 5 +
30 + 20 + 10 = 65. A run that went exactly to plan would have missed it, so
the number was not a commitment anybody could keep by following this document —
it was a commitment to depart from it. The programme already refuses that shape
(assessDisasterRecoveryPlan: "the objective is contradicted by the plan meant
to meet it") and had never put this runbook through the rule. Seventy-five is
the step budget plus ten minutes of headroom, and a number with no headroom over
the plan meant to meet it is missed by any run that hits one snag. Owner
decision: raise it rather than trim a step, because the budgets are the
considered estimate of the work.
Communication#
| Audience | Within |
|---|---|
| The on-call channel | immediately |
| Affected tenants, if the outage passes 30 minutes | 30 min |
Done means#
- The store answers, readiness passes, outbox depth is back to baseline.
- No step acked during the outage window.
- Recorded in the procedure exercise register with the real elapsed time — which is the number the 60-minute objective is a claim about.
Known shortfall#
An outage of a store is exercised. A regional outage is not, and the
scenario list names it (regional-outage in continuity/disaster-recovery.ts)
with recovery as its drill kind. Promoting a secondary region to primary
appears in this workspace only as a fixture string. Do not read this document as
covering that case.