Disciplines · Proposals

Study annotation guide — study-annotation-guide@1.0

RFCs and design proposals (including this one).

1section12 minread

On this page

Published 2026-08-10. The first version — no labels predate it.

Every label in an evaluation dataset cites the version of this guide it was written under. The citation is checked: a version nobody published, a task the cited version does not define, or an adjudication the task's answer model does not permit is a finding against the dataset, not a note on it.

Two rules apply to every task below.

Abstention is an answer. Where the material does not carry enough to decide, decline. A guess is indistinguishable from a reading once it is written down, and a set that rewards guessing measures the annotator rather than the material.

Record what is there, not what follows from it. The material shows a boundary, a word, a voice, a shape, a joint. What any of it means, who anybody is, and how anybody felt are separate questions with their own origin rules, and several of them this workspace refuses to answer at all.

Tasks#

shot-boundary#

Question. At which frame does one shot end and the next begin?

Unit. one boundary between two shots

Answer model. single-answer — one label is right, and agreement statistics apply where the answers can be compared as categories.

Readings recorded. None — a label under this task is a measurement, not a reading of what a work means.

Label. A position on frame number in the edition being annotated. Correct within ±1 frame for a hard cut, ±2 frames for a boundary derived from a gradual transition.

Comparing two annotators. Two answers agree when they fall within ±1 frame for a hard cut, ±2 for a boundary derived from a gradual transition. A category statistic would count a difference inside the tolerance as a disagreement, so none is published for this task.

Where one label stops. A hard cut is marked on the FIRST frame of the incoming shot. A gradual transition — dissolve, wipe, fade through black — has no such frame, so it is marked at the MIDPOINT of the transition and its length in frames is recorded alongside. The midpoint is the convention because it is the only frame of a dissolve that both annotators can find without agreeing first on where the transition started.

When to decline. Decline when the incoming shot cannot be located to within the tolerance — a whip pan into another whip pan, a flash frame, a transition that continues past a cut away. An abstention here is a fact about the material and is scored as correct on an insufficient-context item.

Never record.

  • a camera move, rack focus, or lighting change inside one continuous take is not a shot boundary, however strongly it reads as one
  • a cut in the sound is not a boundary in the picture; the two are separate tracks and separate tasks

Common errors.

  • marking every frame of a dissolve rather than its midpoint
  • marking the last frame of the outgoing shot instead of the first of the incoming one — a consistent one-frame offset that a tolerance of ±1 will hide until it is compounded
  • following the edit as it feels rather than as it is: a match cut on movement is one boundary, not none

transcript-segment#

Question. What was said, and between which times?

Unit. one continuous stretch of speech from one speaker

Answer model. single-answer — one label is right, and agreement statistics apply where the answers can be compared as categories.

Readings recorded. None — a label under this task is a measurement, not a reading of what a work means.

Label. Free text — the words as spoken, in the language spoken, with no correction of grammar, dialect, register, or repetition.

Comparing two annotators. Open text: two transcriptions of one utterance differ in punctuation, elision and hyphenation without differing about what was said. No chance-corrected agreement statistic is defined for this task.

Where one label stops. A segment ends on the release of its final word, not at the start of the silence after it, and not at the point the next speaker begins. A pause inside one utterance stays inside one segment unless it exceeds one second, at which point it is two.

When to decline. Mark inaudible speech as inaudible with its time range. Never transcribe what was probably said: a plausible reconstruction is indistinguishable from a heard word once it is written down, and the whole set is then a measurement of the annotator.

Never record.

  • correcting a speaker: "gonna" is not "going to", and a false start is part of the record
  • transcribing intent — what a speaker meant, implied, or was about to say is not speech
  • attaching a name to a voice; who spoke is a separate task with its own origin rules

Common errors.

  • ending a segment at the start of the following silence, which inflates every duration
  • merging an overlap into one segment because it is easier to type
  • normalising numbers and dates to a written form the speaker did not use

speaker-turn#

Question. Which stretches of this recording are the same voice?

Unit. one turn, belonging to one voice cluster local to this source

Answer model. single-answer — one label is right, and agreement statistics apply where the answers can be compared as categories.

Readings recorded. None — a label under this task is a measurement, not a reading of what a work means.

Label. Free text — a cluster label local to this recording — "voice A", "voice B". The label is an index, and it means nothing outside this source.

Comparing two annotators. The labels are local: the cluster names are indices one annotator assigned to one recording, so two annotators who split the audio identically may share not one label. Agreement is measured over the partition and never over the label names.

Where one label stops. Overlapping speech belongs to BOTH turns and is recorded as two overlapping turns. It is never split down the middle: the midpoint of an overlap is a place neither speaker started or stopped, so a split there is an invented boundary in a task about real ones.

When to decline. Where two voices cannot be told apart — the same speaker after a cold, a deliberate impression, a processed voice — leave the stretch unassigned rather than assigning it to the nearer cluster.

Never record.

  • naming a cluster after a person. A cluster is a continuity claim about a voice and never an identity; an identity may be entered by a human or read from source metadata, and never derived from how a voice sounds
  • inferring age, gender, origin, or health from a voice — none of it is what the task asks and all of it is a demographic inference this workspace refuses

Common errors.

  • creating a new cluster when a speaker changes tone, and reusing one when two speakers happen to sound alike
  • letting the picture decide the cluster — the task is over the audio, and an off-screen line belongs to whoever said it

object-instance#

Question. Which objects are visible in this frame, and where is each one?

Unit. one instance of one object in one frame

Answer model. single-answer — one label is right, and agreement statistics apply where the answers can be compared as categories.

Readings recorded. None — a label under this task is a measurement, not a reading of what a work means.

Label. Free text — a term from the project's own object namespace; a term outside it is a namespace request, not a label.

Comparing two annotators. Both annotators draw from the project's own object namespace, so a chance-corrected agreement statistic over this task is defined.

Where one label stops. The box encloses the visible extent of the object, not its inferred extent: a figure cut by the frame is boxed to the frame edge, and a chair behind a table is boxed around the part that can be seen. A localisation is correct at an intersection-over-union of 0.5 or better against the reference box.

When to decline. Where the object cannot be told from the material — a shape in shadow, a reflection, a blur — leave it unboxed. An uncertain box and a confident one look identical downstream.

Never record.

  • boxing an object the annotator knows is there from an earlier frame but cannot see in this one
  • labelling a person with any attribute; a person is an instance of a person

Common errors.

  • one box around a group rather than one box per instance
  • boxing the inferred whole of a partially occluded object
  • letting box tightness drift over a long session, which shows up as a systematic intersection-over-union decline nobody attributes to the annotator

pose-track#

Question. Where are the body keypoints of each tracked person in this frame?

Unit. one keypoint set for one person in one frame

Answer model. single-answer — one label is right, and agreement statistics apply where the answers can be compared as categories.

Readings recorded. None — a label under this task is a measurement, not a reading of what a work means.

Label. A position on image coordinates, normalised to the frame. Correct within within 5% of the subject height on the axis being measured, per keypoint.

Comparing two annotators. Two answers agree when they fall within 5% of the subject height on the axis being measured, per keypoint. A category statistic would count a difference inside the tolerance as a disagreement, so none is published for this task.

Where one label stops. A keypoint is placed at the joint centre as it appears, not where anatomy says it should be. An occluded keypoint is marked OCCLUDED and left unplaced — never interpolated from the frames around it, because an interpolated point is a smooth track of a measurement nobody made.

When to decline. A person too small, too blurred, or too occluded for the joints to be located is not tracked in that frame. A partial track is honest; a completed one is not.

Never record.

  • reading an emotion, intent, attitude, or state of mind from a pose. A pose is body geometry, and emotion-state inference is a prohibited analysis in this workspace whatever the posture appears to show
  • identifying a person from their geometry or gait, or carrying a track across sources

Common errors.

  • placing occluded joints where they "must" be rather than marking them occluded
  • swapping left and right on a subject facing away
  • continuing one track through a shot boundary onto a different person

cross-reference-relation#

Question. What relation, if any, holds between these two records — and is it one a reader should be shown?

Unit. one ordered pair of records

Answer model. interpretive — where more than one reading is defensible, all of them are recorded and none is adjudicated away.

Readings recorded. None — a label under this task is a measurement, not a reading of what a work means.

Label. One of: same-function, productive-contrast, bridge-reference, keep-distinct.

Comparing two annotators. Both annotators draw from RELATION_JUDGEMENTS, so a chance-corrected agreement statistic over this task is defined.

Where one label stops. Where more than one judgement is defensible, record ALL of them as a bounded set rather than choosing. A bounded set needs at least two readings, each with its grounds in the material, and a statement of what puts a reading OUTSIDE the set — without that statement the set is not bounded and any label could be added later to make a failing system pass.

When to decline. Where the pair carries nothing either way, "keep-distinct" is the answer and not an abstention: it is a positive judgement that the relation must not be drawn. Abstain only when the records cannot be read at all.

Never record.

  • settling an arguable pair on the annotator's own authority. Adjudicating an interpretive item to one label makes the score a measure of conformity to whichever reading won, and it rises as the system gets narrower
  • recording resemblance as identity. Resemblance is what proposes an identity edge and is never what settles one

Common errors.

  • reading "same-function" off surface similarity — an analogy has to reach across something, and two records sharing every axis are merely similar
  • reading "productive-contrast" off two records with nothing in common; a contrast needs a shared frame to be a contrast within
  • calling an adjacent connection a bridge reference, when the distance is the whole content of the claim

performance-choice-function#

Question. What is this performance choice doing in the scene, and what carries it?

Unit. one choice by one performer, inside one continuous stretch of action

Answer model. interpretive — where more than one reading is defensible, all of them are recorded and none is adjudicated away.

Readings recorded. performance-turn, motivation, subtext, emotion, mental-state — kinds the product never lets a machine assert without a human, so what is written here is what a machine suggestion is weighed against.

Label. Free text — a reading stated as what the depiction DOES, with the frames or lines it rests on. "Withholding the look until the second question" is a reading; "sad" is a verdict with nowhere to point.

Comparing two annotators. Open text: a reading is stated in the annotator's own words, and two people describing the same choice share no string to be counted against each other. No chance-corrected agreement statistic is defined for this task.

Where one label stops. Where more than one reading is defensible, record ALL of them as a bounded set and adjudicate none. Two readings are two readings when they rest on DIFFERENT material or point at different moments; the same reading in other words is ONE reading, and splitting it widens the set until anything a system says falls inside. The set states what puts a reading outside it — normally the extent of the material the annotator was given.

When to decline. Two different declines, recorded differently. INSUFFICIENT CONTEXT when this material cannot carry the reading: the face is out of frame, the line is inaudible, the choice is only legible against a scene that was not supplied. INTENTIONAL UNKNOWN when the material is sufficient and the answer is still not established: nobody wrote it down, and the people who were there remember it differently. The first is a fact about the material and its right answer is "cannot tell"; the second is a fact about the record and has no right answer to hold anyone to.

Never record.

  • attributing a state to the PERFORMER. The task is over the depiction — "the performance plays the fear as impatience" is a reading; "the actor was frightened" is an inference about a real person's interior, which this workspace refuses outright
  • reading the choice off the script rather than off the performance. What the line says is available to every reading and settles none of them
  • settling an arguable moment on the annotator's own authority. Adjudicating an interpretive item to one label makes the score a measure of conformity to whichever reading won, and it rises as a system gets narrower

Common errors.

  • writing the outcome instead of the choice: what the scene later reveals is not what the performance is doing here, and a set annotated with hindsight cannot show whether a system reads the moment or the plot
  • recording a technical adjustment as a choice about the character — a re-angle for the light, a pause held for an overlapping line, a mark hit late
  • one bounded set per annotator rather than per item: the set is what the material supports, not a list of who said what