- Status: approved (2026-08-14)
- Priority: EXT — non-blocking (no shipped gate is scored against a gold set that has no expert-decided cases in it)
- Decision owner: @GreyChimp
- Drafted: 2026-08-11 by Claude Code (Opus 5)
- Approval: approved as recommended by @GreyChimp on 2026-08-14; outcome recorded in decision-log.json; review by 2027-08-14
Question#
The evaluation programme needs film educators, directors, actors, animators, editors, cinematographers, sound designers, costume specialists, game designers, accessibility experts, cultural reviewers, learners, and statistical owners — people, engaged on terms, paid or not, named or not. Who is recruited, on what terms, and who signs for the money and the confidentiality?
This is not a scheduling question. Until somebody with standing over the material has decided a case, every rate the programme reports is a rate over answers the system itself supplied (YSD-18011), and the apparatus that can hold expert judgement holds none.
Recommendation#
Recruit against the procedures already contracted under YSD-18003 — conflict of interest, compensation, attribution, confidentiality, review, withdrawal — and start with the two roles that unblock the most: a cultural reviewer and a domain expert for the first-release lens set. Do not begin annotation before YSD-19005 licenses the slices, or the panel annotates material it may not keep.
Options considered#
- Recruit broadly up front — rejected: thirteen roles engaged before a single licensed slice exists means paying people to wait.
- Use internal staff as stand-in experts — rejected: the evaluation's whole claim is that the judgement came from someone with standing, and an internal annotator is the checked party supplying both sides.
- Two roles first, against licensed slices — recommended.
Consequences#
- Gold sets stay apparatus-only until the panel exists; the §18.1 case-content items (YSD-18010, YSD-18011, YSD-18012) cannot be marked before then.
- Compensation and confidentiality become live commitments with a named signer.
- The evaluation reports carry a panel roster, which is itself disclosable.
Machine-enforced outcome#
Expert-decided cases enter the gold sets carrying their annotator's declared experience, and the scorers separate expert from beginner divergence rather than averaging them into one number.
Control in force#
The apparatus still refuses to pretend the panel exists, and that does not
change with approval — only recruitment changes it. ANNOTATOR_EXPERIENCE
distinguishes beginner from expert rather than merging them, so an
unrecruited expert cannot be silently counted as one, and assertGoldCase
refuses a case whose truth is not properly stated. No release gate is scored
against a populated expert corpus, because there is not one. What approval
settles is the TERMS recruitment runs under, not that it has happened.