Disciplines · Decisions (ADRs)

ADR-0086: Eve per-plane reliability objectives

The V1 operations policy defines useful service-level uptime and HTTP latency, but Eve crosses twelve logical planes between session intake, prompt assembly, tools, memory, workbench execution, providers, voice, and evidence.

3sections2 minread

On this page
  • Status: Accepted
  • Date: 2026-09-14
  • Decision owners: SRE Lead and the named plane owner
  • Independent verifier: Charter QA

Context#

The V1 operations policy defines useful service-level uptime and HTTP latency, but Eve crosses twelve logical planes between session intake, prompt assembly, tools, memory, workbench execution, providers, voice, and evidence. A service average can remain green while one of those boundaries loses a task, commits a late effect, drops cost, or crosses a data boundary. Task 13.1 therefore needs a total plane-by-metric contract before trace propagation, fault tests, alert deployment, or game-day evidence can be evaluated consistently.

Decision#

Adopt EVE_SOTA_RELIABILITY_SLO_CONTRACT_2026-09.md and its machine record as the canonical Task 13.1 contract.

The contract makes these choices:

  1. The twelve planes come from the source-verified Task 4.1 threat model. Every plane has exactly one disposition for each of fifteen required reliability dimensions: a concrete SLO or an explicit boundary-based non-applicability reason.
  2. Objectives use a rolling 30-day window. Failed attempts, timeouts, retries, partial outcomes, fallback attempts, and their cost stay in the denominator. Missing or unreconciled telemetry is UNKNOWN, never passing.
  3. Charter targets are floors, not ceilings. The contract retains the scorecard TTFT, total-latency, cancellation, recovery, grounding, and cost guards and sets stricter 99% operational completion/recovery ratios where appropriate.
  4. Duplicate effects, late effects, post-cancel or post-kill effects, fabricated citations, unpriced legs, false recovery success, and data boundary violations are hard locks. A statistical interval, maintenance window, or good aggregate cannot excuse one event.
  5. Fractional objectives receive paired fast- and slow-burn alerts. Each availability objective also receives a telemetry-blindness alert. Every alert names a stable owner, severity, existing runbook, and safe degraded mode.
  6. Threshold, applicability, exclusion, ownership, and degradation changes are prospective versioned decisions approved by SRE and the affected domain owner before affected observations are viewed.

Consequences#

  • Tasks 13.2–13.7 can use stable plane, SLI, RTO, budget, and alert identities.
  • A dashboard with no samples cannot report green, and a retry cannot erase its failed predecessor or spend.
  • Hard-lock incidents fence the affected plane and preserve evidence before recovery verification.
  • This decision does not claim the metrics are instrumented, alerts are deployed, targets are attained, or recovery has been exercised. Those remain the direct evidence obligations of Tasks 13.2–13.7.