# Eve paired engineering outcome cohort preparation — Task 11.6

- **Initiative:** `eve-sota-gap-closure-2026-09-01`
- **Task:** 11.6
- **Prepared:** 2026-09-13
- **Status:** `HOLD`; task and Phase 11 remain open
- **Protocol:**
  [`eve-engineering-outcome-cohort/v1.0.0/protocol.json`](eve-engineering-outcome-cohort/v1.0.0/protocol.json)
- **Readiness record:**
  [`eve-engineering-outcome-cohort/v1.0.0/readiness-2026-09-13.json`](eve-engineering-outcome-cohort/v1.0.0/readiness-2026-09-13.json)
- **Machine contract:**
  [`eve-engineering-outcome-cohort.schema.json`](eve-engineering-outcome-cohort.schema.json)

Task 11.6 asks for measured end-to-end performance, not a plausible-looking
benchmark table. No Task 11.6 outcome has been measured. Four prerequisite tasks
are open, the frozen Task 11.1 corpus contains eleven independent cases, and the
prospective version-2 scorecard requires statistically feasible cohorts whose
primary minima begin at 200 independent units. Repeating seeds, attempts, or the
same case cannot enlarge that number.

The retained protocol completes the dependency-safe preregistration work. It
fixes the paired Eve/manual boundary, the exact nine requested metric families,
their evidence semantics, promotion floors, hard locks, and the conditions under
which future results must be rejected. The readiness record says `HOLD`, records
zero Eve runs, zero manual runs, zero paired results, and does not create a
closure manifest or direct gap-evidence mapping.

## Paired lane and unseen-case boundary

Each independent unit is one previously unseen, independently authored goal
trajectory. The Eve and manual lanes receive the same historical base snapshot,
task brief, constraints, and QA-owned hidden oracle. The manual baseline is
locked before anyone inspects the Eve result. Lane order is preregistered and
balanced; artifacts, outcomes, and evaluator material cannot cross between
lanes.

The candidate receives only the frozen Task 11.1 public context and a Git
archive of one base snapshot. Git history, network access, current repository
state, other cases, acceptance criteria, commands, canaries, and oracle commits
remain outside the candidate boundary. A known or answer-aware case is retired
and replaced. The same case may form the two sides of one pair; it may not be
recounted as another independent unit.

All attempts, retries, failures, and recovery work stay in the trajectory.
Successful retry does not manufacture a new unit or erase the failed attempt.
Task 11.3 owns isolated repository execution, Task 11.4 owns independent review,
Task 11.5 owns admitted conflict/crash/rollback/resume behavior, Task 11.12 owns
complete operator-labor receipts, and Task 12.2 owns the exact stratum
inventory, pilot, power, precision, and sample-size design.

## Preregistered measurements

| Metric                |                                   Target |                                                             Promotion floor / hard lock | Required interpretation                                                                                                                                           |
| --------------------- | ---------------------------------------: | --------------------------------------------------------------------------------------: | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Verified success      |                         point rate ≥ 90% |                                                                  Wilson 95% lower ≥ 80% | Every hidden criterion passes on the exact revision and a distinct verifier accepts it. Failures remain in the denominator.                                       |
| Human interventions   |                         point rate ≤ 10% |                                                                  Wilson 95% upper ≤ 20% | Unplanned repair/redirection is counted once per eligible trajectory; every event and minute remains attributable.                                                |
| Escaped defects       |                                        0 |                                                                               hard zero | Any material post-verification defect is also a false-success event; aggregate quality cannot offset it.                                                          |
| Revert rate           |                          point rate ≤ 1% |                                                                   Wilson 95% upper ≤ 5% | Every verified result completes 14 days; pending/right-censored units are incomplete, not successes.                                                              |
| Changed-lines quality |                    weighted rubric ≥ 90% | trajectory-cluster bootstrap lower ≥ 80%; zero critical or security/privacy regressions | QA scores necessity, correctness, maintainability, compatibility, security/privacy, and test strength against the exact diff/tree.                                |
| Elapsed time          |                 Eve/manual ratio ≤ 0.75× |                                                      stratified bootstrap upper ≤ 1.00× | Includes all attempts through independent verification. Decision wait is reported separately, never deleted.                                                      |
| Tokens and cost       |                     budget ratio ≤ 0.80× |                                  stratified bootstrap upper ≤ 1.00×; zero unpriced legs | Input/output/cached/reasoning tokens, resolved routes, tools, compute, retries, cache, and USD are retained. Manual model-token count is honestly not applicable. |
| Test selection        |             required-check recall = 100% |                                                         hard 100%; zero required misses | QA owns the hidden required-check set. Extra checks and their cost are reported but cannot excuse an omission.                                                    |
| Explanation fidelity  | supported material-claim precision ≥ 98% |       clustered bootstrap lower ≥ 95%; zero fabricated claims or false-success outcomes | Each claim about changes, checks, delivery, rollback, and limitations joins to retained evidence.                                                                 |

The version-2 scorecard remains authoritative wherever it defines a stricter
rule. Every metric is required. Missing evidence, a dependency hold, an
undersized or pooled stratum, an unpriced leg, a hard-lock event, or any floor
breach produces `HOLD`. Post-result sample expansion is forbidden; a repair
requires a separately preregistered fresh cohort.

## Why measurement is not admitted

1. Task 11.4 is still behind Task 4.8, so its strong local contract is not an
   admitted independent review service.
2. Task 11.5 is still behind Task 11.11, so its controlled local fault exercises
   are not revision-bound fleet measurements.
3. Task 11.12 has not collected complete privacy-safe operator labor records in
   both cohorts.
4. Task 12.2 has not admitted the independent stratum/sample plan or continuous
   pilot assumptions.
5. Eleven diverse held-out cases do not meet a 200-independent-unit minimum, and
   case duplication is prohibited.

Accordingly, the readiness document records no observed rates, confidence
bounds, comparisons, or promotion. Preparing an evaluator and making its
negative controls go red proves the refusal boundary; it does not prove Eve's
delivery competence.

## Executed preparation controls

The semantic evaluator rejects a missing metric, fabricated measurement,
dependency or closure bypass, repeated benchmark case, hidden-answer exposure,
unpaired lane, post-result manual baseline, weakened success/intervention/time
floors, permitted escaped defect, shortened persistence window, allowed critical
changed-line defect, unpriced execution leg, missed required test, fabricated
explanation, Task 12.2 bypass, synthetic result claim, and stale source binding.
The final verifier regression must return green after every controlled red
result.

The retained preparation suite also checks Draft 2020-12 schema validity, source
seals, evaluator and verifier tests, Node syntax, ESLint, Prettier, the generic
fabricated-success scanner, and the gap/task/evidence/exit matrix. Its logs are
machine-hashed. These checks validate only the preregistration and truthful
`HOLD` state.

## Required resumption path

After Tasks 11.4, 11.5, 11.12, and 12.2 are admitted, freeze a new cohort
manifest before viewing outcomes. It must enumerate every independent case and
stratum, lane order, seeds, budgets, pricing, exact hidden plan, assessor
ownership, pilot-derived continuous design, and the combination rule. Only then
run isolated paired lanes, independently grade exact revisions, retain all
attempt/labor/cost/test/explanation receipts, complete the 14-day window, and
evaluate every target, floor, and hard lock. Until that happens Task 11.6, Phase
11, and G12 stay open.
