Disciplines · Audits

Eve per-plane reliability SLI/SLO and error-budget contract

UNKNOWN and cannot satisfy an SLO, release gate, or recovery claim.

5sections9 minread

On this page

Status: defined-uninstrumented

Task 13.1 fixes the reliability objectives and incident response metadata for every plane in the source-verified Eve threat model. It is a policy contract, not evidence that telemetry or alerts are deployed.

Global rules#

  • Window: rolling-30-days.
  • Missing data: Missing, late, unreconciled, or below-minimum telemetry is UNKNOWN and cannot satisfy an SLO, release gate, or recovery claim.
  • Denominators: Failures, retries, timeouts, cancellations, partial outcomes, failed-attempt cost, and fallback attempts remain attributable and cannot be dropped.
  • Amendments: Any threshold, applicability, exclusion, owner, or degraded-mode change requires a new version, rationale, SRE plus domain-owner approval, and commit before affected observations are viewed.

Plane coverage#

Plane Owner Applicable SLOs Explicit N/A cells
session-http assistant_product 11 4
operator-http reliability 9 6
prompt-assembly assistant_product 8 7
member-domain-tools platform_engineering 10 5
client-tool-bridge reliability 10 5
retrieval platform_engineering 8 7
operator-tools assistant_product 10 5
memory-and-session compliance 7 8
workbench-intent platform_engineering 11 4
model-provider reliability 12 3
voice assistant_product 12 3
audit-and-evidence compliance 8 7

Every plane has exactly one disposition for every catalog metric. “N/A” is a checked boundary statement, not a missing cell.

Objectives#

SLO SLI target Error budget Alerts
session-http.availability >= 0.999 ratio bad-event-fraction: 0.001 3
session-http.ttft >= 0.99 ratio; p50 <= 1000 ms bad-event-fraction: 0.01 2
session-http.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
session-http.task-completion >= 0.99 ratio bad-event-fraction: 0.01 2
session-http.cancel-latency >= 0.99 ratio bad-event-fraction: 0.01 3
session-http.duplicate-action = 0 events hard-zero: 0 1
session-http.late-action = 0 events hard-zero: 0 1
session-http.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
session-http.grounding >= 0.98 precision; factual-claim coverage >= 0.95 bad-event-fraction: 0.02 3
session-http.cost >= 0.99 ratio; p95 priced cost-to-budget ratio <= 0.80 bad-event-fraction: 0.01 3
session-http.data-boundary-violations = 0 events hard-zero: 0 1
operator-http.availability >= 0.999 ratio bad-event-fraction: 0.001 3
operator-http.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
operator-http.task-completion >= 0.99 ratio bad-event-fraction: 0.01 2
operator-http.cancel-latency >= 0.99 ratio bad-event-fraction: 0.01 3
operator-http.duplicate-action = 0 events hard-zero: 0 1
operator-http.late-action = 0 events hard-zero: 0 1
operator-http.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
operator-http.cost >= 0.99 ratio; p95 priced cost-to-budget ratio <= 0.80 bad-event-fraction: 0.01 3
operator-http.data-boundary-violations = 0 events hard-zero: 0 1
prompt-assembly.availability >= 0.999 ratio bad-event-fraction: 0.001 3
prompt-assembly.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
prompt-assembly.task-completion >= 0.99 ratio bad-event-fraction: 0.01 2
prompt-assembly.cancel-latency >= 0.99 ratio bad-event-fraction: 0.01 3
prompt-assembly.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
prompt-assembly.grounding >= 0.98 precision; factual-claim coverage >= 0.95 bad-event-fraction: 0.02 3
prompt-assembly.cost >= 0.99 ratio; p95 priced cost-to-budget ratio <= 0.80 bad-event-fraction: 0.01 3
prompt-assembly.data-boundary-violations = 0 events hard-zero: 0 1
member-domain-tools.availability >= 0.999 ratio bad-event-fraction: 0.001 3
member-domain-tools.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
member-domain-tools.tool-completion >= 0.99 ratio bad-event-fraction: 0.01 2
member-domain-tools.cancel-latency >= 0.99 ratio bad-event-fraction: 0.01 3
member-domain-tools.kill-latency >= 0.99 ratio bad-event-fraction: 0.01 3
member-domain-tools.duplicate-action = 0 events hard-zero: 0 1
member-domain-tools.late-action = 0 events hard-zero: 0 1
member-domain-tools.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
member-domain-tools.cost >= 0.99 ratio; p95 priced cost-to-budget ratio <= 0.80 bad-event-fraction: 0.01 3
member-domain-tools.data-boundary-violations = 0 events hard-zero: 0 1
client-tool-bridge.availability >= 0.999 ratio bad-event-fraction: 0.001 3
client-tool-bridge.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
client-tool-bridge.tool-completion >= 0.99 ratio bad-event-fraction: 0.01 2
client-tool-bridge.cancel-latency >= 0.99 ratio bad-event-fraction: 0.01 3
client-tool-bridge.kill-latency >= 0.99 ratio bad-event-fraction: 0.01 3
client-tool-bridge.duplicate-action = 0 events hard-zero: 0 1
client-tool-bridge.late-action = 0 events hard-zero: 0 1
client-tool-bridge.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
client-tool-bridge.cost >= 0.99 ratio; p95 priced cost-to-budget ratio <= 0.80 bad-event-fraction: 0.01 3
client-tool-bridge.data-boundary-violations = 0 events hard-zero: 0 1
retrieval.availability >= 0.999 ratio bad-event-fraction: 0.001 3
retrieval.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
retrieval.tool-completion >= 0.99 ratio bad-event-fraction: 0.01 2
retrieval.cancel-latency >= 0.99 ratio bad-event-fraction: 0.01 3
retrieval.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
retrieval.grounding >= 0.98 precision; factual-claim coverage >= 0.95 bad-event-fraction: 0.02 3
retrieval.cost >= 0.99 ratio; p95 priced cost-to-budget ratio <= 0.80 bad-event-fraction: 0.01 3
retrieval.data-boundary-violations = 0 events hard-zero: 0 1
operator-tools.availability >= 0.999 ratio bad-event-fraction: 0.001 3
operator-tools.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
operator-tools.tool-completion >= 0.99 ratio bad-event-fraction: 0.01 2
operator-tools.cancel-latency >= 0.99 ratio bad-event-fraction: 0.01 3
operator-tools.kill-latency >= 0.99 ratio bad-event-fraction: 0.01 3
operator-tools.duplicate-action = 0 events hard-zero: 0 1
operator-tools.late-action = 0 events hard-zero: 0 1
operator-tools.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
operator-tools.cost >= 0.99 ratio; p95 priced cost-to-budget ratio <= 0.80 bad-event-fraction: 0.01 3
operator-tools.data-boundary-violations = 0 events hard-zero: 0 1
memory-and-session.availability >= 0.999 ratio bad-event-fraction: 0.001 3
memory-and-session.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
memory-and-session.tool-completion >= 0.99 ratio bad-event-fraction: 0.01 2
memory-and-session.duplicate-action = 0 events hard-zero: 0 1
memory-and-session.late-action = 0 events hard-zero: 0 1
memory-and-session.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
memory-and-session.data-boundary-violations = 0 events hard-zero: 0 1
workbench-intent.availability >= 0.999 ratio bad-event-fraction: 0.001 3
workbench-intent.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
workbench-intent.task-completion >= 0.99 ratio bad-event-fraction: 0.01 2
workbench-intent.cancel-latency >= 0.99 ratio bad-event-fraction: 0.01 3
workbench-intent.kill-latency >= 0.99 ratio bad-event-fraction: 0.01 3
workbench-intent.duplicate-action = 0 events hard-zero: 0 1
workbench-intent.late-action = 0 events hard-zero: 0 1
workbench-intent.queue-age >= 0.99 ratio bad-event-fraction: 0.01 2
workbench-intent.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
workbench-intent.cost >= 0.99 ratio; p95 priced cost-to-budget ratio <= 0.80 bad-event-fraction: 0.01 3
workbench-intent.data-boundary-violations = 0 events hard-zero: 0 1
model-provider.availability >= 0.999 ratio bad-event-fraction: 0.001 3
model-provider.ttft >= 0.99 ratio; p50 <= 1000 ms bad-event-fraction: 0.01 2
model-provider.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
model-provider.task-completion >= 0.99 ratio bad-event-fraction: 0.01 2
model-provider.cancel-latency >= 0.99 ratio bad-event-fraction: 0.01 3
model-provider.kill-latency >= 0.99 ratio bad-event-fraction: 0.01 3
model-provider.duplicate-action = 0 events hard-zero: 0 1
model-provider.late-action = 0 events hard-zero: 0 1
model-provider.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
model-provider.provider-errors >= 0.99 ratio bad-event-fraction: 0.01 2
model-provider.cost >= 0.99 ratio; p95 priced cost-to-budget ratio <= 0.80 bad-event-fraction: 0.01 3
model-provider.data-boundary-violations = 0 events hard-zero: 0 1
voice.availability >= 0.999 ratio bad-event-fraction: 0.001 3
voice.ttft >= 0.99 ratio; p50 <= 1000 ms bad-event-fraction: 0.01 2
voice.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
voice.task-completion >= 0.99 ratio bad-event-fraction: 0.01 2
voice.cancel-latency >= 0.99 ratio bad-event-fraction: 0.01 3
voice.kill-latency >= 0.99 ratio bad-event-fraction: 0.01 3
voice.duplicate-action = 0 events hard-zero: 0 1
voice.late-action = 0 events hard-zero: 0 1
voice.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
voice.provider-errors >= 0.99 ratio bad-event-fraction: 0.01 2
voice.cost >= 0.99 ratio; p95 priced cost-to-budget ratio <= 0.80 bad-event-fraction: 0.01 3
voice.data-boundary-violations = 0 events hard-zero: 0 1
audit-and-evidence.availability >= 0.999 ratio bad-event-fraction: 0.001 3
audit-and-evidence.total-latency >= 0.99 ratio; p50 <= 6000 ms bad-event-fraction: 0.01 2
audit-and-evidence.task-completion >= 0.99 ratio bad-event-fraction: 0.01 2
audit-and-evidence.duplicate-action = 0 events hard-zero: 0 1
audit-and-evidence.late-action = 0 events hard-zero: 0 1
audit-and-evidence.queue-age >= 0.99 ratio bad-event-fraction: 0.01 2
audit-and-evidence.recovery >= 0.99 ratio bad-event-fraction: 0.01 3
audit-and-evidence.data-boundary-violations = 0 events hard-zero: 0 1

Alert contract#

Fractional budgets use paired fast-burn and slow-burn alerts. Hard locks fire on the first event. Each plane availability SLO also alerts on telemetry blindness so an empty dashboard cannot look green. Every alert record carries its owner, severity, runbook, and exact safe degraded mode.

Limitations#

  • This contract defines objectives and response policy; it does not claim that Task 13.2 telemetry propagation, Task 13.4 load/soak measurement, Task 13.5 fault injection, Task 13.6 restore proof, or Task 13.7 alert deployment/game day has occurred.
  • Current SLO attainment is UNKNOWN until production observations reconcile at the named boundaries; missing telemetry never counts as success.
  • The contract does not admit a provider, tool, protocol, channel, DCC, desktop, or long-running capability that another gate keeps blocked.

Machine record: docs/audits/eve-sota-reliability-slo/2026-09-14.json