- Initiative:
eve-sota-gap-closure-2026-09-01 - Task: 11.6
- Prepared: 2026-09-13
- Status:
HOLD; task and Phase 11 remain open - Protocol:
eve-engineering-outcome-cohort/v1.0.0/protocol.json - Readiness record:
eve-engineering-outcome-cohort/v1.0.0/readiness-2026-09-13.json - Machine contract:
eve-engineering-outcome-cohort.schema.json
Task 11.6 asks for measured end-to-end performance, not a plausible-looking benchmark table. No Task 11.6 outcome has been measured. Four prerequisite tasks are open, the frozen Task 11.1 corpus contains eleven independent cases, and the prospective version-2 scorecard requires statistically feasible cohorts whose primary minima begin at 200 independent units. Repeating seeds, attempts, or the same case cannot enlarge that number.
The retained protocol completes the dependency-safe preregistration work. It
fixes the paired Eve/manual boundary, the exact nine requested metric families,
their evidence semantics, promotion floors, hard locks, and the conditions under
which future results must be rejected. The readiness record says HOLD, records
zero Eve runs, zero manual runs, zero paired results, and does not create a
closure manifest or direct gap-evidence mapping.
Paired lane and unseen-case boundary#
Each independent unit is one previously unseen, independently authored goal trajectory. The Eve and manual lanes receive the same historical base snapshot, task brief, constraints, and QA-owned hidden oracle. The manual baseline is locked before anyone inspects the Eve result. Lane order is preregistered and balanced; artifacts, outcomes, and evaluator material cannot cross between lanes.
The candidate receives only the frozen Task 11.1 public context and a Git archive of one base snapshot. Git history, network access, current repository state, other cases, acceptance criteria, commands, canaries, and oracle commits remain outside the candidate boundary. A known or answer-aware case is retired and replaced. The same case may form the two sides of one pair; it may not be recounted as another independent unit.
All attempts, retries, failures, and recovery work stay in the trajectory. Successful retry does not manufacture a new unit or erase the failed attempt. Task 11.3 owns isolated repository execution, Task 11.4 owns independent review, Task 11.5 owns admitted conflict/crash/rollback/resume behavior, Task 11.12 owns complete operator-labor receipts, and Task 12.2 owns the exact stratum inventory, pilot, power, precision, and sample-size design.
Preregistered measurements#
| Metric | Target | Promotion floor / hard lock | Required interpretation |
|---|---|---|---|
| Verified success | point rate ≥ 90% | Wilson 95% lower ≥ 80% | Every hidden criterion passes on the exact revision and a distinct verifier accepts it. Failures remain in the denominator. |
| Human interventions | point rate ≤ 10% | Wilson 95% upper ≤ 20% | Unplanned repair/redirection is counted once per eligible trajectory; every event and minute remains attributable. |
| Escaped defects | 0 | hard zero | Any material post-verification defect is also a false-success event; aggregate quality cannot offset it. |
| Revert rate | point rate ≤ 1% | Wilson 95% upper ≤ 5% | Every verified result completes 14 days; pending/right-censored units are incomplete, not successes. |
| Changed-lines quality | weighted rubric ≥ 90% | trajectory-cluster bootstrap lower ≥ 80%; zero critical or security/privacy regressions | QA scores necessity, correctness, maintainability, compatibility, security/privacy, and test strength against the exact diff/tree. |
| Elapsed time | Eve/manual ratio ≤ 0.75× | stratified bootstrap upper ≤ 1.00× | Includes all attempts through independent verification. Decision wait is reported separately, never deleted. |
| Tokens and cost | budget ratio ≤ 0.80× | stratified bootstrap upper ≤ 1.00×; zero unpriced legs | Input/output/cached/reasoning tokens, resolved routes, tools, compute, retries, cache, and USD are retained. Manual model-token count is honestly not applicable. |
| Test selection | required-check recall = 100% | hard 100%; zero required misses | QA owns the hidden required-check set. Extra checks and their cost are reported but cannot excuse an omission. |
| Explanation fidelity | supported material-claim precision ≥ 98% | clustered bootstrap lower ≥ 95%; zero fabricated claims or false-success outcomes | Each claim about changes, checks, delivery, rollback, and limitations joins to retained evidence. |
The version-2 scorecard remains authoritative wherever it defines a stricter
rule. Every metric is required. Missing evidence, a dependency hold, an
undersized or pooled stratum, an unpriced leg, a hard-lock event, or any floor
breach produces HOLD. Post-result sample expansion is forbidden; a repair
requires a separately preregistered fresh cohort.
Why measurement is not admitted#
- Task 11.4 is still behind Task 4.8, so its strong local contract is not an admitted independent review service.
- Task 11.5 is still behind Task 11.11, so its controlled local fault exercises are not revision-bound fleet measurements.
- Task 11.12 has not collected complete privacy-safe operator labor records in both cohorts.
- Task 12.2 has not admitted the independent stratum/sample plan or continuous pilot assumptions.
- Eleven diverse held-out cases do not meet a 200-independent-unit minimum, and case duplication is prohibited.
Accordingly, the readiness document records no observed rates, confidence bounds, comparisons, or promotion. Preparing an evaluator and making its negative controls go red proves the refusal boundary; it does not prove Eve's delivery competence.
Executed preparation controls#
The semantic evaluator rejects a missing metric, fabricated measurement, dependency or closure bypass, repeated benchmark case, hidden-answer exposure, unpaired lane, post-result manual baseline, weakened success/intervention/time floors, permitted escaped defect, shortened persistence window, allowed critical changed-line defect, unpriced execution leg, missed required test, fabricated explanation, Task 12.2 bypass, synthetic result claim, and stale source binding. The final verifier regression must return green after every controlled red result.
The retained preparation suite also checks Draft 2020-12 schema validity, source
seals, evaluator and verifier tests, Node syntax, ESLint, Prettier, the generic
fabricated-success scanner, and the gap/task/evidence/exit matrix. Its logs are
machine-hashed. These checks validate only the preregistration and truthful
HOLD state.
Required resumption path#
After Tasks 11.4, 11.5, 11.12, and 12.2 are admitted, freeze a new cohort manifest before viewing outcomes. It must enumerate every independent case and stratum, lane order, seeds, budgets, pricing, exact hidden plan, assessor ownership, pilot-derived continuous design, and the combination rule. Only then run isolated paired lanes, independently grade exact revisions, retain all attempt/labor/cost/test/explanation receipts, complete the 14-day window, and evaluate every target, floor, and hard lock. Until that happens Task 11.6, Phase 11, and G12 stay open.