Date: 2026-09-12
Decision#
The seven release-scoped eval cases already exist. Task 10.1 verifies that substrate; it does not re-author the cases and does not claim that automatic release-scope activation exists.
The runtime source of truth still declares V1.0, with Veritas and Metis
deferred to V1.2. The complete eval deck contains exactly six family cases and
one golden case with executable releaseBlocked metadata. Every case is
advisory, every current expectation forbids its withheld tool, and every exact
restore expectation calls that tool.
| Case | V1.0 withheld tool | V1.2 restore | Retained raw k=10 |
|---|---|---|---|
veritas-tool-selection |
veritas_top_claims |
call claims; do not substitute Tara recommendations | 10/10 |
family-md-empty-saved-articles |
veritas_saved_articles |
call saved articles; require honest empty result and no fixture headline | 4/10 |
family-md-course-continue |
metis_continue_learning |
call continue learning | 10/10 |
family-md-metis-search |
metis_search_catalog |
call catalog search | 6/10 |
family-md-metis-recommend |
metis_recommended_courses |
call recommendations and ground the fixture course | 10/10 |
family-md-empty-course-search |
metis_search_catalog |
call catalog search; require honest empty result and no fixture course | 8/10 |
family-md-empty-recommend |
metis_recommended_courses |
call recommendations; require honest empty result and no fixture course | 10/10 |
Retained measurement#
The latest retained production-binding measurement is the 2026-08-30 Phase-11 receipt. It contains 70 runs (seven cases at k=10), the raw 58/70 split above, and a current-grader classification of 59/70 after one narrow, logged boundary phrase repair. It records zero provider retries, $0.0295 billed cost, 8.1 s median turn latency, serving through DeepInfra and Baidu, 83.6% cache-read over 63 reporting runs, and complete teardown of both run-unique databases. No withheld tool was admitted.
This is retained sanitized aggregate evidence, not raw transcript evidence. The record preserves per-case counts, providers, usage, cost, latency, retry, and cleanup fields; it does not preserve prompts or response text. Task 10.1 does not rerun the model and moves no full-deck, family, or builder-Wilson floor.
The focused BFF TypeScript project for this audit passes. The complete BFF typecheck is not a green branch-wide signal at this checkpoint: a 6 GiB run reached its heap cap, and a supervised 8 GiB retry completed with existing diagnostics in unrelated Arete/Yemaya libraries and none in the files touched by this task. Those upstream diagnostics are not relabelled as Task 10.1 failures, and the smaller project is retained as the exact changed-source typecheck.
Executable proof#
release-blocked-eval-audit.ts derives the seven cases from the complete deck,
executes the deck's existing admission validator, and independently checks the
exact V1.2 expectations. The deterministic capture binds that projection, the
authoritative release constants, the retained measurement, its verifier, and the
scorecard by SHA-256. The semantic verifier rechecks the current files and
rejects inventory narrowing, an advisory case promoted early, a withheld tool
allowed in V1.0, weakened restore behavior, a fabricated automatic selector, k
below 10, inflated measurement, invented raw transcripts, stale sources, and
widened floor claims. Each control is observed red before the clean regression.
Boundary#
releaseBlocked.restoreExpectation remains metadata. The evaluator still does
not select it from the authoritative runtime release scope; that is Task 10.2.
Task 10.3 owns provider-free dual-scope behavior and inverted/missing mapping
controls. Task 10.4 owns a fresh V1.0 k≥10 rerun and the eventual V1.2 rerun.
Therefore this audit closes Task 10.1 only; Phase 10 and G8 remain open.
Machine-readable audit:
docs/audits/eve-sota-release-scope-eval-audit/2026-09-12.json
Evidence manifest: docs/audits/eve-sota-evidence/phase-10/task-10-1.json