Disciplines · Audits

Eve creative-workflow benchmark preparation — 2026-09

Gap audits and as-built reviews.

0sections3 minread

On this page

Task 12.9 now has a deterministic, content-bound benchmark candidate for four separate creative families: reference-to-editable motion, original brief-to-editable motion, measured-house reconstruction, and interactive architectural walkthrough. It remains blocked, rather than preregistered, because Task 17.7 has not supplied and ratified the authoritative creative delivery contracts and coverage map. The absent contract is represented as a null digest, the limitation is explicit, and the verifier refuses a promoted status while that dependency is absent.

Each family fixes its independent case identity before variants or output inspection, prohibits derivatives and practiced capstones crossing splits, and requires at least 200 independent cases with at least 40 per case class. Three fixed model seeds are nested under a case rather than counted as extra evidence; deterministic runtime paths require two replays. Decisions must remain visible by case class, complexity, native/runtime format, provider route or no-model route, and fresh versus corrected input.

The benchmark measures quality at the boundary that can establish it. Reference motion uses decoded frames and native-timeline readback for layout, timing, and Bezier fidelity. Original motion adds blinded brief, brand, composition, and originality review. Reconstruction compares native geometry and topology with independent dimensional ground truth and hard-locks unsupported load-bearing claims. Walkthroughs are exercised in the actually served runtime, including input-to-motion latency, frame rate, navigation containment, reconnect behavior, accessible alternatives, and dimensional readback after engine import. Native editability is checked after save and reopen; an unopenable project or unserved build is a zero-tolerance failure.

Every family also carries source-to-verified-delivery measures for success, failure or abandonment, paired human labor, revisions, wall-clock latency, and provider/render budget. Route identity comes from runtime receipts joined to the current approved model registry. Asset metadata, UI/marketing labels, and a headline clip duration are expressly inadmissible. Model, rendering, storage, streaming, retry, queue, and revision legs remain in their respective cost and latency boundaries, and any unpriced provider or render leg hard-locks a family. Ratios use the same held-out case completed by an independently assigned qualified professional in conventional native tools without Eve or model help. Lane order is randomized across at least three professionals, acceptance inputs and limits are identical and blinded, and a preregistered stopping rule retains all attempts instead of selecting the best output after inspection.

Visual-judge calibration is itself held out by source identity and requires at least 100 independent cases per criterion and case class, three blinded human raters per case, Krippendorff alpha of at least 0.67, Spearman correlation of at least 0.80, false accepts no greater than 0.05, and false rejects no greater than 0.10. The model judge is advisory. Humans remain required for missing or failed calibration, drift, unresolved appeals, dimensional decisions, high-risk cases, and hard-lock-relevant decisions. Human ratings use fixed anchors for unusable, substantively revision-requiring, and delivery-ready work.

The decision rule is conjunctive: every applicable target and floor must pass in every preregistered stratum, all hard locks must remain zero, and all four families must pass before release. Missing cases, prices, receipts, native readbacks, labels, repeats, failed attempts, or denominators are INCOMPLETE and block. Amendments are prospective, independently approved, content-hashed, and made before affected outputs are visible. Criterion/stratum denominators are locked before execution; Wilson 95% intervals govern binary measures, while continuous measures use a separately held non-promoting pilot and at least 10,000 case-clustered BCa resamples.

The schema and verifier enforce the exact family membership, case classes, sample and repeat rules, criteria, numeric thresholds, measurement boundaries, judge calibration, routing and costing rules, source hashes, dependency status, and record digest. The focused suite includes a deterministic regeneration check and real-CLI negative controls for a missing family, pooled repeats, undersized samples, missing dimensions, weakened thresholds, nonzero hard locks, unvalidated judges, metadata attribution, unpriced cost, hidden dependency status, stale source hashes, fabricated ratification, and output-visible amendments, an assisted comparator, and unanchored human scores.

Verify the retained blocked candidate with:

bash
node tools/eve-everywhere/verify-creative-workflow-benchmark.mjs

Regenerate it deterministically with:

bash
node tools/eve-everywhere/generate-creative-workflow-benchmark.mjs

This artifact contains no held-out creative case, human label, model run, native project, runtime receipt, quality score, cost, labor, latency, or operator decision. It is preregistration machinery awaiting Task 17.7, not evidence that any creative family, Task 12.9, Phase 12, G7, or G14 passes.