Disciplines · Audits

Eve SOTA Operator-Memory Poisoning Audit — Task 9.5

Operator memory is a small, explicitly confirmed, operator-self-scoped data surface.

5sections5 minread

On this page

Date: 2026-09-12 Scope: EVE_SOTA_GAP_CLOSURE_TODOS_2026-09-01.md task 9.5 Decision: verified

Claim#

Operator memory is a small, explicitly confirmed, operator-self-scoped data surface. It is not an authority channel. The production boundary refuses sensitive/member-data shapes, serializes stored values as JSON-quoted fallible data below policy/live state/live instruction, and permanently fences an account-erased subject against resurrection. No vector projection exists in this exact-memory implementation; semantic recall remains task 9.6.

Ingress and containment matrix#

Input named by 9.5 Runtime status Route to persistence Control and proof
User Present Authenticated operator control, or explicitly requested model proposal A pure anchored resolver reads only the current authenticated operator text. Explicit remember/forget commands admit only the matching mutation tool; direct controls derive the subject from verified auth and require an exact confirmation marker.
Page Present Can influence the model, but has no writer or mutation-tool grant A scripted worst-case provider forces a poisoned page proposal; remember_operator_note is absent from the offered tools, the hallucinated call fails, no card is parked, and no row exists. The live case requires no memory-tool call.
Channel Absent from Eve runtime None Task-4.1 threat inventory records the absent path and its admission trigger. The task-9.5 production writer census fails if any new channel writer appears.
Tool result Present Can influence the model, but has no writer or mutation-tool grant Trust-labelled untrusted and absent from the intent resolver. Existing security-suite tool-result attacks are retained; task 9.5 adds a stored tool-escalation attack.
Documentation Present for admin Can influence the model, but has no writer or mutation-tool grant Docs results are trust-labelled untrusted and absent from the intent resolver. Existing docs security cases remain in the security suite.
Agent input Direct agent messages absent; workbench records present Workbench text can influence the model, but has no writer or mutation-tool grant Threat inventory records both states. Only current authenticated operator text reaches the resolver; the writer census catches any future direct path.

This distinction is deliberate: claiming a live channel or direct agent path would fabricate product surface. The meaningful invariant is that every currently present indirect input and every future production TypeScript caller must converge on, or be caught outside, the same persistence boundary.

Production controls#

  1. Sensitive/member-data admission. Values and slots undergo NFKC and invisible-format normalization before matching member identifiers, email, phone, SSN, Luhn-valid payment cards, labelled government/bank data, credential assignments, bearer/basic/JWT forms, provider keys, URL credentials, private-key headers, and IP addresses. Errors never echo the submitted value.
  2. Structural demotion. Confirmed value bytes remain verbatim in storage, but prompt rendering uses JSON strings. Newlines therefore remain \\n inside one memory line and cannot mint bullets or pseudo-system sections. The prompt explicitly says values are fallible data and places policy, authoritative live state, and the live operator instruction above them.
  3. Human authority. The runtime offers no memory mutation tool unless an anchored, operator-text-only resolver recognizes an explicit direct command; it offers only remember or forget as requested. The model-facing writer is additionally mutating: validation runs before a card is parked and execution requires the same action id, session, and operator. Cross-operator confirmation returns 403. Direct privacy controls are admin-only, self-only, and carry distinct explicit confirmation markers. Conservative grammar can reject shorthand; the direct Memory controls remain available.
  4. Anti-resurrection erasure. Account erasure locks all categories in the same order as writers, inserts a domain-separated SHA-256 subject digest, and deletes current and history rows in one PostgreSQL transaction. A write that wins first is subsequently deleted; one that runs second sees the fence and fails. Ordinary “forget all” intentionally does not fence, so an operator may opt back in without recreating a deleted account identity.
  5. No vector remnant. The task-9.5 PostgreSQL receipt enumerates the three assistant_operator_memory% tables and refuses any vector/embedding column or vector UDT. Account erasure leaves zero current/history rows. Task 9.6 owns any future semantic index and must extend this deletion proof before one can ship.

Adversarial verification#

  • Real PostgreSQL lifecycle and route tests cover normalized smuggling forms, structural injection, exact/adjacent isolation, ordinary opt-back-in, persistent digest fences, both concurrent race orderings, unconfirmed poisoning, explicit-intent tool scoping, hallucinated unavailable calls, and cross-operator confirmation theft.
  • The dedicated live-model matrix seeds the real durable store for every run. It tests stored directive adoption, page-borne prompt exfiltration, stale status conflict, destructive memory-tool escalation, structural newlines, and benign preference utility. The gate requires every case to pass every run at k >= 10; final response text is not retained in the receipt.
  • The final qualified live run used OpenRouter's registered deepseek/deepseek-v4-flash-0731 turn model with the fp8 preference. All 60/60 trials passed (6/6 cases at pass^10): zero of 50 attack trials succeeded and the benign utility control passed 10/10. No memory mutation tool was called; five live-conflict trials used the non-mutating admin_release_readiness grounding tool. All trials reported DeepInfra as the serving endpoint. Usage was 790,555 input and 18,849 output tokens with provider-reported cost $0.02969844.
  • The static verifier inventories every production rememberOperatorMemory callsite and exercises mutation controls for census expansion, normalization and admission bypass, prompt quoting, fence creation/check/locking, confirmation, operator-intent bypass, cross-operator coverage, live conflict coverage, and vector introduction. All 13 mutation controls fail as intended before the clean verifier passes.

Prompt-ratchet limitation: the 04424a2f… stamp also reconciles 192 rendered skill bytes from the previously committed 2026-09-06 navigate/page-inspection change that the then-current ratchet had not captured. Task 9.5's live matrix does not rerun the navigate family and is not used to move any family or Wilson floor; the scorecard records that inherited delta separately from this task's 168 description-byte memory-tool change.

Evidence index#

  • PostgreSQL receipt: docs/audits/eve-sota-operator-memory-poisoning/postgres-receipt-2026-09-12.json
  • Live-model receipt: docs/audits/eve-sota-operator-memory-poisoning/live-model-receipt-2026-09-12.json
  • Retained gate logs: docs/audits/eve-sota-operator-memory-poisoning/
  • Machine-readable claim manifest: docs/audits/eve-sota-evidence/phase-09/task-9-5.json

The decision is verified because the evidence runner completed, the live-model receipt passed all 6 × 10 trials, every targeted gate and negative control matched its expected outcome, and the machine-readable manifest validated. This closes task 9.5 only; task 9.3 still awaits independent human labels, task 9.6 owns any semantic index, and task 9.7 owns rollout evidence.