Disciplines · Audits

Eve long-context and context-compaction evaluation — 2026-09

Gap audits and as-built reviews.

0sections1 minread

On this page

Task 15.3 is evaluated at the serving seam, not by counting context-window tokens. The retained matrix drives the production compactor, trust-label derivation, and tool-result digest through seven distinct loss classes and a 40-turn rolling trajectory.

The compactor now retains both edges of an oversized dropped utterance. That preserves late corrections and source references while keeping the same bounded character budget. Every block says that it is untrusted and lossy, numbers dropped turns in chronological order, and requires the model to ask the member or re-run the authoritative tool when an omitted middle, omitted turn, conflict, instruction detail, citation, provenance field, or tool result matters.

The outer tool-result cap no longer emits broken partial JSON. A result without a specific digester is returned as a valid explicit-loss envelope with the original character count, a bounded prefix, and a re-query/no-inference instruction. Search results retain two quotable excerpts and all discovery links; later excerpts require a narrower query.

The seven direct classes are instruction loss, authority/trust-label loss, citation/provenance loss, stale memory, conflicting state, tool-result truncation, and long-trajectory recovery. The last case rolls from 9 through 40 turns, proves the 24-folded/8-recent boundary, counts the eight fully omitted earliest turns, and requires the recovery contract at every roll.

This does not claim unlimited semantic memory. Arbitrary middle content can still be omitted, conflicts are not semantically adjudicated by the compactor, trust labels are not cryptographic provenance, and the trajectory is synthetic and bounded. The regional live receipt retains no prompt or response content. If the reviewed endpoint is unavailable, it is recorded as an admission failure rather than relabelled as a model-quality pass.