GuidesAnnotations & review

Annotations & review

Human labels are the ground truth for Neens: a reviewer looks at a session and records a pass or fail verdict with a critique. The Review page queues the sessions most worth a human look, the Annotations tab is the browsable history of every label, and the Alignment view measures how well your judges agree with those humans — closing the loop that keeps automated scoring honest.

At a glance

WhereReview page — Queue and Annotations tabs; Alignment tab on the Judges page
Label storeOne shared ground-truth store: a label written in the queue is the same record the Annotations tab and alignment metrics read
Verdictspass / fail, plus a free-text critique and an optional failure-mode or cluster tag
RolesReviewer (any annotator) and principal (expert — their labels and gold labels define truth)
AlignmentPrecision, recall, specificity, and Cohen’s κ per judge, recomputed nightly (3:00 UTC) or on demand

The Review queue

The Queue tab is a prioritized worklist, not a raw session list. Each candidate session is ranked by five signals — Impact, Uncertainty, Novelty, Disagreement, and Redundancy — so reviewer time goes where a human verdict changes the most: high-impact sessions, scores sitting near their pass/fail threshold, traffic unlike anything labeled before, and places where judges contradict each other, while clusters that already have several labels are de-prioritized.

Configure the queue

The queue configuration is saved per agent (open the config drawer on the Queue tab):

SettingDefaultMeaning
Statuseserror, failed, failureWhich session statuses are eligible.
Score filtersnonee.g. only sessions with a metric lt some value.
Failure modes / clustersallRestrict the queue to specific failure modes or clusters.
Only unlabeledoffDrop sessions that already have a label.
Novelty budget5Max novelty-ranked items per queue.
Queue size20 (max 200)How many items the queue returns.

Label a session

Pick an item

Select a queue item to open the full session — every span, tool call, and message — so the verdict is grounded in what actually happened (see Traces & sessions).

Record the verdict

Choose Pass or Fail (keyboard: P / F), write a critique explaining why, and optionally tag a taxonomy failure mode or a discovered cluster.

Submit

The label lands in the shared ground-truth store with your identity attached. Relabeling the same session supersedes your earlier label rather than deleting it — the audit trail is kept.

Gold labels and principals. Labels from principal reviewers (experts) and labels explicitly marked gold define the expert truth used for alignment and for grading other annotators. A principal can also adjudicate a disputed session, superseding all conflicting labels with one authoritative verdict.

The Annotations tab

The Annotations tab on the Review page browses the same label store: every label with its target, verdict, critique, failure-mode/cluster tag, gold flag, timestamp, and who wrote it — each label permanently carries its reviewer’s name, and superseded labels remain inspectable.

It also hosts the Annotator leaderboard: per reviewer, label volume, pass/fail split, gold labels authored, and — where a reviewer has labeled targets that also have expert truth — their agreement rate and Cohen’s κ against it. Volume gets recognition; agreement tells you whose labels to trust.

One store, two views. The Review queue and the Annotations tab read and write the same ground-truth labels — there is no separate “annotations” dataset to keep in sync. Those same labels are what the alignment metrics below are computed against.

Judge ↔ human alignment

Once you have expert labels, the Alignment tab on the Judges page measures each judge against them. A judge’s most recent verdict per session is paired with the expert truth for that session, treating failure as the positive class:

MetricQuestion it answers
PrecisionWhen the judge flags a failure, how often is it really one?
Recall (TPR)Of the real failures, how many does the judge catch?
Specificity (TNR)How well does it avoid false alarms on good sessions?
Cohen’s κChance-corrected overall agreement (0 ≈ random, 1 = perfect).

Convergence chart

Each recompute persists a measurement, so the chart plots every judge as a line over time — toggle between κ, Precision, and Recall, and narrow the window with the time-range picker (Today / 24h / 7d / 30d / All / Custom). A rising line means your judge edits are converging on human judgment; a flat low line means the rubric still doesn’t match what your experts consider a failure.

Alignment is recomputed automatically every night at 3:00 UTC, and on demand with the refresh action. With no gold/expert labels yet, it degrades gracefully — you’ll see an empty trend with the reason, never fabricated numbers.

Disagreement drill-down

Select a judge to see exactly where it diverged from the experts, split into:

  • Missed failures — the judge passed a session an expert failed (the costly kind).
  • False alarms — the judge failed a session an expert passed.

Each row shows the judge’s score and reasoning next to the expert’s critique. If the human was wrong (or the case is genuinely ambiguous), click Relabel to write a corrected gold label in place — then hit Recompute alignment to refresh the metrics immediately.

Closing the loop: iterate the judge

Alignment turns judge editing from guesswork into a measured cycle:

  1. Measure — check κ/precision/recall for the judge on the Alignment tab.
  2. Diagnose — read the disagreements: is the judge missing a failure type your experts catch, or nitpicking things they don’t care about?
  3. Edit — update the judge’s instructions or criteria; saving publishes a new immutable version (score history from older versions is preserved). See Judges.
  4. Re-run — run the new version over labeled sessions and recompute alignment.
  5. Repeat until the trend converges — then trust the judge at scale, including as a gate in pre-prod evaluations.
API reference
RoutePurpose
GET /review-queueThe ranked review queue.
GET / PUT /review/queue-configRead / save the per-agent queue configuration.
POST /reviewSubmit a label (target_id, verdict pass|fail, critique?, failure_mode_id?, cluster_id?, is_gold?).
GET /review/labelsThe unified label history (the single ground-truth store).
GET /review/goldActive gold-standard labels.
POST /review/adjudicatePrincipal override — supersede conflicting labels with one authoritative verdict.
GET /review/leaderboardPer-annotator volume + gold-agreement + κ.
GET /review/alignmentPersisted alignment trend (?range=, ?judge_id=).
POST /review/alignment/refreshRecompute + persist alignment for every judge with scores.
GET /review/alignment/disagreements?judge_id=…A judge’s disagreements, split missed-failures / false-alarms.

Troubleshooting

SymptomCause → fix
Alignment tab is emptyNo gold/principal labels yet — label sessions in the Review queue (mark decisive ones gold), then refresh alignment.
Queue is emptyNo sessions match the configured statuses/filters, or Only unlabeled is hiding everything already labeled — loosen the queue config.
A judge’s κ is high but it still feels wrongκ is computed only on sessions with expert truth; label more (and more varied) sessions so the sample represents real traffic.
Leaderboard shows no agreement for an annotatorThey haven’t labeled any targets that also have expert truth — agreement needs overlap with gold/principal labels.

See also the FAQ.