Owner: delivery pipeline on-call (rota:delivery-pipeline). Last reviewed:
2026-08-14. Version: 1.
The queue and everything resting on it stop. Work already accepted is the question: the interrupted paths finish nothing, and whether accepted work is replayed or lost is decided by the durable record rather than by the queue.
What dep:job-queue actually is on this estate, because it decides the whole
recovery. There is no broker. apps/oshun/bff/src/generation/jobs-route.ts
holds jobs in an in-process map and persists the whole set as a snapshot
through createSnapshotSink
(libs/oshun/persistence/src/durable-snapshot-store.ts) into the
admin_store_snapshot table — one row per (storeKey, scopeId), the key being
generation-jobs.
So the durable record of accepted work is in dep:primary-store, and a
queue loss on this estate is a process losing its map. Two consequences:
- The recovery is a reload, not a replay. Restarting the BFF re-loads the snapshot. Nothing has to be reconstructed from an event log because there is no event log in this path.
- A queue loss with a healthy relational store loses only in-flight
transitions — the work between the last successful
persistJobs()and the stop.requireDurableGenerationJobs()exists so that a mutation is not acknowledged before it is durable; where it is called, the acknowledged set and the durable set agree.
1. Detection and declaration#
Detection signal. Jobs stop reaching terminal states while submission keeps
succeeding: dep:job-worker-pool rests on dep:job-queue in the S10.5
substrate, so the queue's failure is observed as work that is accepted and never
finishes rather than as a submission error. The objectives that move first are
slo:job-success-rate and slo:gate-run-completion.
- Detection authority: delivery pipeline on-call, contact
rota:delivery-pipeline. - Declaration authority: production operations duty officer, contact
rota:prod-ops. - Gap, stated.
docs/operations/on-call.mddescribes a weekly primary/secondary rotation and names no person and no rotation config; none of the register'srota:identifiers appears in it or indocs/operations/alert-routing.md. The name above is registered; the route from it to whoever is carrying the pager is not.
2. Communications#
| Audience | Within | Content |
|---|---|---|
| Actors | 15 min | that submitted work is not running, whether it was accepted, and that it need not be resubmitted |
| Operators | 5 min | that this runbook is running, and by whom |
Both travel over dep:notification-channel, which this scenario leaves
standing.
Tell actors not to resubmit. A resubmission during a queue outage is the
duplicate this estate is least able to collapse: the snapshot is keyed by a
freshly minted randomUUID() job id, so two submissions of the same work are
two jobs.
3. Dependency assumptions#
dep:primary-storeis healthy. It holdsadmin_store_snapshot, which is the queue's durable record — if it is not healthy, this is not a queue incident, it is v1-dr-database-loss.md.dep:blob-storeis healthy: job outputs are written there and a replayed job writes there again.
4. Recovery infrastructure#
In place. The workers and the queue are the same process tree as the rest of the BFF; there is nowhere else for them to run and nothing to provision.
5. Secrets#
- Provider credentials the workers need to finish the work they reload
(
OSHUN_MESHY_API_KEYand its siblings, read at the worker and failing closed when absent).
Recoverable out of band: the configuration-and-secrets state class is in the
off-box backup (infra/hetzner/backup/offbox-backup.sh).
6. Restore order#
dep:job-queue— bring the queue back first by restarting the process that loads thegeneration-jobssnapshot. The workers have nothing to take from an empty queue, and a worker pool started first will report itself healthy while doing nothing.dep:job-worker-pool— the pool rests on the queue in the substrate, so it follows it.
No reference edge touches either, so the ordering constraint here is the substrate one rather than the referential one.
7. Validation#
- The reloaded snapshot's job count equals the count recorded before the stop. Write the count down before restarting — this is the check the estate cannot perform for you, and without it a snapshot that loaded nothing looks exactly like one that loaded everything.
- One synthetic job submitted after recovery reaches a terminal state.
- The queued backlog drains rather than sitting at a constant depth: a worker pool that reconnects to a queue it cannot read holds a steady non-zero depth and reports healthy.
Gap, stated.
box-loss-restore-drill.shassertsredisRestoreVerifiedfrom aPING, and an empty Redis answersPING. Redis is captured byoffbox-backup.sh(a forcedBGSAVE, then thedump.rdb), but nothing on this estate compares its restored contents to a count. If the queue's state ever moves fromadmin_store_snapshotinto Redis, this section is wrong and must be rewritten before that move ships.
8. Fencing — NOT APPLICABLE#
Nothing is promoted, so there is no old primary to reject. The adjacent risk is
a real one and is not covered by anything here: two workers reloading the
same snapshot. The snapshot is a whole-set write (persistJobs serialises
through a write tail), so a second process that loads and saves concurrently
overwrites rather than merges. Bring back exactly one.
Worth knowing while you do: the estate already fences leases elsewhere.
live_media_pipeline_job_attempts.lease_token
(libs/shared/live-media/src/sql-media-pipeline-job-store.ts) is checked in the
UPDATE ... WHERE of every heartbeat, lease expiry and completion, so a worker
whose lease was reassigned updates zero rows. The generation-job snapshot has no
such token, which is why "bring back exactly one" is an instruction to a person
rather than a property of the store.
9. Failover#
There is no standby queue and no second worker pool. The estate has one of each.
10. Reconciliation#
scripts/operations/v1-restore-reconcile.mjs compares restored queues against
what the outside world observed, and it is the control that answers the question
this scenario turns on: was an accepted job's effect already delivered? A
job that reached a provider before the stop and is reloaded as pending will be
run again.
Give it expected counts. Every collection it is handed can be empty, and empty reconciles clean.
The replay procedure is not written. There is no queue drain or replay script on this estate —
ls scripts/matches nothing forqueue,drainorreplay. S10.12.c is where those procedures belong, and this runbook will not invent one. Until then, replay is the reconciler's blocker list read by a person.
11. Failback#
The estate never left its normal topology; failback is the confirmation that it is back in it — backlog draining, one synthetic job terminal, and the reloaded count matching what was recorded. Record the instant: recovery time runs from the fault, not from the restart.
12. What this runbook does not cover#
- The relational store that holds the snapshot — see v1-dr-database-loss.md.
- Provider failures the workers hit while draining — see v1-dr-model-provider-outage.md.
- Dead-letter triage.
libs/oshun/workbench-kit/src/job-queue-views.tsmodels thedead-letteredstate; the operator procedure for draining it is not written.
13. Machine-checkable statement#
{
"id": "runbook:v1-queue-loss",
"scenarioId": "gd:queue-loss",
"version": "1",
"detectionAuthority": {
"role": "service on-call",
"name": "delivery pipeline on-call",
"reachableOver": ["dep:notification-channel"]
},
"declarationAuthority": {
"role": "operations duty officer",
"name": "production operations duty officer",
"reachableOver": ["dep:notification-channel"]
},
"communications": [
{
"audience": "actors",
"channelDependencyId": "dep:notification-channel",
"withinMs": 900000
},
{
"audience": "operators",
"channelDependencyId": "dep:notification-channel",
"withinMs": 300000
}
],
"assumes": ["dep:primary-store", "dep:blob-store"],
"recoveryInfrastructureDependencyId": null,
"secrets": [
{
"id": "provider worker credentials",
"heldInDependencyId": "dep:job-queue",
"recoverableOutOfBand": true
}
],
"restoreOrder": ["dep:job-queue", "dep:job-worker-pool"],
"fencing": null,
"promotesAStandby": false,
"validation": [
"the reloaded generation-jobs snapshot holds the job count recorded before the stop",
"one synthetic job submitted after recovery reaches a terminal state",
"the backlog depth falls rather than holding steady"
],
"failback": "confirm the backlog drains, one synthetic job is terminal and the reloaded count matches; the estate never left its normal topology"
}