Review
The Review page is where people label agent runs pass or fail. Neens builds you a queue of the items where a human verdict is worth the most, shows you what happened, and records your verdict as ground truth — the reference your judges are measured against, the source of your regression sets, and the confirmation that a failure mode is real.
At a glance
| Where | Review in the sidebar. Two tabs: Queue (your worklist) and Labelled (every label so far, with the labelling trend) |
| What you label | Traces (one agent run) or Sessions (a whole conversation) — switch at the top of the page |
| Verdicts | Pass, Fail, Skip, or Flag for discussion, plus an optional failure mode, critique and gold mark |
| Queue size | 1–200 items, default 25 |
| Queue sources | Smart mix (default), Recent failures, A failure mode or issue, Low judge scores, Judge disagreements, Random sample |
| Scope | Each queue belongs to one person, one agent and one grain (traces or sessions) |
| API | POST /api/review/queues, GET /api/review/queues/current, POST /api/review — see Review via the API |
What your verdicts feed
Every verdict you record lands in the same shared ground-truth store the rest of Neens reads. Nothing needs to be exported or synced.
| Your verdict… | …is used by |
|---|---|
| Pass / Fail on a trace | Judge alignment: precision, recall and Cohen’s κ for each judge are computed against human verdicts, so you know whether to trust a judge before it scores or gates anything |
| Pass / Fail that contradicts a judge | The agent’s regression set — a trace where a judge got it wrong is captured automatically, so the next judge version is tested on exactly that case |
| Failure mode tag on a fail | Confirmation that a discovered failure mode is real, with real examples behind it |
| Gold mark | Expert truth: gold labels are what judge accuracy and other reviewers’ agreement are graded against |
| Critique | The “why” a judge author reads when a judge disagrees with you |
| Pass / Fail on a session | Session-level ground truth: did the user get what they needed across the whole conversation |
Traces or sessions
Use the Traces / Sessions switch at the top of the page to choose what you are judging. Each has its own queue, so switching never loses your place in the other.
| Traces | Sessions | |
|---|---|---|
| One item is | One agent run | A whole conversation (every trace that shares a conversation id) |
| Best for | A single tool call, answer or step: “was this refund issued correctly?” | The outcome: “did the user get what they needed?” |
| Viewer tabs | Summary, Spans, Graph, Agent map, Raw | Transcript, Traces, Agent map, Raw |
| Feeds judge alignment and regression sets | Yes | No — stored as session-level ground truth |
A session verdict is not copied onto its traces. A conversation can fail while every one of its traces looks fine on its own (the agent was polite and wrong four times in a row), and a conversation can succeed despite one failed step. So Neens stores a session verdict as a label on the conversation, separate from trace labels. Your trace-level judge alignment and regression sets are never changed by session reviews. If you want both views, review the same conversation in both grains.
Sessions only appear if your agent sends a conversation id with its traces — see Traces & sessions → Sessions.
The workspace
The Queue tab has three columns.
- Queue list (left). Every item in your queue, in order, with its age, agent and a reason chip — Errored, Judge on the fence, Judges disagree, New pattern, Low score or Random check. A dot shows each item’s state (to do, pass, fail, skipped). Filter between To do and Done. A progress bar above shows how many you have reviewed, the pass / fail / skipped split, and an estimate of the time left.
- What happened (centre). For a trace: what the user asked, what the agent answered, and the steps it took (LLM calls, tool calls) with durations; errors are shown in place. Open the Spans tab to inspect a step’s inputs and outputs. For a session: the transcript of the whole conversation, and the list of traces it is made of. When Neens knows what to check, a Look for hint sits on top (for example “the refund for the duplicate charge was never issued”), taken from the judge that flagged the item or from the error.
- Verdict panel (right).
- Pass or Fail.
- Skip when you can’t tell; Not sure: flag for discussion to mark it for a teammate.
- Failure mode — if it failed, pick the failure mode. Neens pre-selects a Suggested one when the item belongs to a known failure mode.
- What went wrong? — an optional critique. Keep it to one sentence a judge author can act on.
- Mark as a gold example — tick this when you are confident and the case is a good yardstick; gold labels are used to check judge accuracy.
- Why you’re seeing this — the item’s rank in the queue and the signals that put it there (see below).
Recording a verdict moves you to the next to-do item. Changing your mind later is safe: relabelling the same item replaces your earlier label but keeps the old one in the history.
Build a queue
Open the queue builder from the controls at the top of the Queue tab (the source, time window and size chips, or More filters).
What do you want to judge?
Traces or Sessions (see above).
Where should items come from?
Pick one of six sources:
| Source | What it picks | Use it when… |
|---|---|---|
| Smart mix (default) | Failures, borderline judge scores, conflicts and new patterns, ranked by how much a label helps | You have a few minutes and want them spent well. The right choice most days |
| Recent failures | Errored or failed runs, newest first | Something just broke and you want to see it with your own eyes |
| A failure mode or issue | Everything in one failure mode or cluster | You want to confirm a failure mode is real before anyone fixes it |
| Low judge scores | Runs a chosen judge scored below a threshold (default 0.5) | You suspect one judge is too harsh or too lenient on its low end |
| Judge disagreements | Runs where judges conflict with each other, or with the run’s own status | You want the fastest route to finding wrong judges |
| Random sample | An unbiased sample, not ranked | You want to measure how accurate your judges really are. Ranked queues over-sample hard cases, which skews accuracy numbers |
Narrow it down
- Time window — Last 24 hours, Last 7 days (default), Last 30 days or All time.
- Agent version — limit to one version of your agent.
- Failure mode — limit to one failure mode (required for A failure mode or issue).
- Skip items I or a teammate already labelled — on by default. Turn it off to re-review.
- Add a few unusual items — reserves a share of the queue for items unlike anything labelled
before, so new failure patterns get seen. Turned on by default at 20% of the queue (the API
accepts 0–50% via
exploreShare). Those slots are counted inside the queue size: a 25-item queue with 20% unusual items is 20 ranked items plus 5 unusual ones, still 25 in total.
How many?
Choose 10, 25 (default), 50 or 100, or type any number from 1 to 200. The builder estimates how long the queue will take.
Preview and build
The footer previews the result before anything is built: “142 traces match. You’ll review the top 25”, with the mix (for example 14 failures · 6 judge borderline · 3 conflicts · 2 unusual). Tick Save as my default to make these settings the default for everyone on this agent who has not built a queue yet, then Build queue.
How queues behave
- Your queue is yours. Each person has their own queue per agent and per grain (one for traces, one for sessions). Teammates’ queues can overlap, so keep Skip items I or a teammate already labelled on to avoid labelling the same item twice.
- Order is stable. A queue stays in the same order across reloads and devices until you regenerate it, so you can stop and pick up where you left off.
- Regenerate keeps your labels. Regenerate rebuilds the queue with the current settings. Items you already labelled stay, in order; only unlabelled items are replaced with fresh candidates.
- Labels made elsewhere show up. If someone labels an item from the Traces page, it shows as done in your queue too.
- No queue yet? The first time you open Review, Neens builds one from the agent’s saved default (or Smart mix, last 7 days, 25 items).
Why you’re seeing this
Every item in a Smart mix queue is ranked by five signals. The verdict panel shows the item’s rank (“Ranked #3 of 25”) and each signal rated High, Medium, Low or None relative to the other items in the queue, with the evidence behind it. How items are ranked opens the full breakdown.
| Signal | Previously called | Question it answers | Example evidence |
|---|---|---|---|
| How much it hurt | Impact | Did it fail, and how costly was it? | The run errored at tool.issue_refund and used 18.4k tokens, more than 90% of runs |
| Judges unsure | Uncertainty | Is a judge score sitting near its pass line? | ”Resolution quality” scored 0.68 against a 0.70 pass line |
| Signals conflict | Disagreement | Do judges disagree with each other or with the run’s status? | The run errored, but “Resolution quality” judged it a pass. One of them is wrong |
| How unusual | Novelty | Is it unlike anything labelled before? | Close to the centre of the “Wrong resolution offered” failure mode |
| Already covered | Redundancy | Have people already labelled plenty like it? | Nobody has labelled this failure mode yet, so nothing counts against it |
The priority score combines them:
priority = 0.35 × impact + 0.25 × uncertainty + 0.25 × novelty + 0.15 × disagreement − 0.20 × redundancyEach signal is between 0 and 1, so a run that errored badly (impact 0.95), with a judge near its
line (uncertainty 0.55), that judges disagree about (disagreement 1), in a familiar failure mode
(novelty 0.05), that nobody has labelled yet (redundancy 0) scores
0.33 + 0.14 + 0.01 + 0.15 − 0 ≈ 0.63. Items below about 0.30 usually teach judges little.
Already covered is the only signal that lowers priority: once a failure mode has several labels, more of the same teaches little, so the queue moves on to something new.
Smart mix ranks by this score. Recent failures sorts by time and Random sample doesn’t rank at all, so you get an unbiased spot check; their items have no priority score.
Keyboard shortcuts
Keyboard shortcuts work on the Queue tab whenever you are not typing in a text box.
| Key | Action |
|---|---|
J / K | Next / previous item |
P | Pass |
F | Fail |
S | Skip |
C | Jump to the critique box |
Esc | Leave the critique box |
? | Show all shortcuts |
When the queue is complete
When the last item is done, Neens shows a summary of what your labels changed:
- Reviewed — the count, the pass / fail / skipped split and how long it took.
- Judge disagreements found — verdicts that contradicted a judge. Each was added to the regression set automatically.
- Judge agreement — each affected judge’s agreement with people, before and after your labels, and how many labels it is now based on (for example “Resolution quality” agreement 71% → 76%, now based on 41 labels (was 16)).
- Failure modes confirmed — the failure modes you tagged, with counts.
Then pick what’s next: Build another queue, Check judge alignment, Review the skipped items, or review the next batch when more candidates still match your filters (“142 traces still match this queue’s filters, and 117 of them are unlabelled”).
How-tos
Confirm a failure mode before fixing it
A failure mode is a hypothesis until a person has looked at it. Confirm it before you build a regression set or start a fix.
Build a focused queue
Open the queue builder, choose Traces and A failure mode or issue, and pick the failure mode. Set the size to 10–25: enough to see the pattern, short enough to finish.
Label and tag
For each trace, decide whether it really shows this failure. On Fail, keep the Suggested failure mode if it fits, or pick the right one if the trace is a different problem. Pass the ones that were clustered in by mistake.
Read the summary
Failure modes confirmed shows how many traces you confirmed. If most passed, the failure mode is not real or not one problem — dismiss or split it on Issues & failure modes. If most failed, freeze them into a regression set on Datasets.
Calibrate a judge with a random sample
Ranked queues are deliberately full of hard cases, which makes them great for teaching judges and bad for measuring them. To measure a judge’s real accuracy, label a random sample.
Build a random queue
Choose Traces, Random sample, Last 30 days, and a size of 50. Untick Add a few unusual items so the sample stays unbiased.
Label without looking at the judge first
Decide pass or fail from what happened, then move on. Mark the clear-cut cases gold.
Check alignment
From the completion summary, choose Check judge alignment. The judge’s precision, recall and κ against your labels are now based on a representative sample. See Judge ↔ human alignment.
Review whole conversations for a support agent
For a support agent, the question that matters is often whether the customer’s problem got solved, not whether each step looked right.
Switch to Sessions
Choose Sessions at the top of the Review page, then build a Smart mix or Recent failures queue.
Read the transcript
The Transcript tab shows the whole conversation. Ask one question: did the user get what they needed? Use the Traces tab when you need to see which run went wrong.
Record the outcome
Pass if the problem was resolved, Fail if not, with a one-line critique (“refund promised in turn 2, never issued”). The verdict is stored for the conversation as a whole.
Go deeper where needed
If a failed conversation points at one broken step, switch to Traces and label that run too, so it can feed judge alignment and the regression set.
Review via the API
Everything the workspace does is available over the API, relative to your Neens host. Calls use
an API key from Settings → API keys, sent as a bearer token. Every route is also available
without the /api prefix.
Build (or regenerate) a queue. The body is the same as the builder’s settings; only the fields you want to change are needed.
curl -X POST https://your-neens/api/review/queues \
-H "Authorization: Bearer $NEENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"grain": "trace",
"source": "failure_mode",
"window": "7d",
"size": 25,
"filters": { "failureModeId": "fm_wrong_resolution" },
"skipLabelled": true,
"exploreShare": 0.2
}'The response is the queue plus its items, in order. Each item carries the targetType and
targetId to label, its status, its rank and priority, and the plain-language reasons:
{
"queue": { "id": "rq_8f2c…", "grain": "trace", "source": "failure_mode", "size": 25,
"candidateCount": 142, "mix": { "errored": 14, "borderline": 6, "conflict": 3, "novel": 2 } },
"items": [
{ "position": 0, "targetType": "session", "targetId": "29c14f8f…", "status": "todo",
"rank": 1, "priority": 0.75,
"primaryReason": { "kind": "errored", "label": "Errored" },
"reasons": [ { "signal": "impact", "level": "high", "title": "It failed",
"detail": "tool.issue_refund returned an error; 18.4k tokens" } ] }
],
"progress": { "total": 25, "todo": 25, "pass": 0, "fail": 0, "skipped": 0, "flagged": 0 }
}In the API, a trace is labelled with targetType: "session" and a session (conversation) with
targetType: "conversation". Always send back the targetType the queue item gives you.
Record a verdict. Pass the queue_id so the item is marked done in the queue:
curl -X POST https://your-neens/api/review \
-H "Authorization: Bearer $NEENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"queue_id": "rq_8f2c…",
"target_type": "session",
"target_id": "29c14f8f…",
"verdict": "fail",
"critique": "Offered a credit instead of refunding the duplicate charge.",
"failure_mode_id": "fm_wrong_resolution",
"is_gold": false
}'All Review queue routes
| Route | Purpose |
|---|---|
GET /review/queues/current?grain=trace|session | Your current queue for this agent and grain; builds one from the saved default if you have none |
POST /review/queues | Build or regenerate a queue (keeps labelled items). Add "saveAsDefault": true to save the settings as the agent’s default |
POST /review/queues/preview | Same body; returns candidateCount, willReview and mix without building anything |
PATCH /review/queues/{queue_id}/items/{position} | Set an item’s status to skipped, flagged or back to todo. Pass and fail are set only by POST /review |
GET /review/queues/{queue_id}/summary | The completion summary: counts, duration, disagreements captured, failure modes confirmed, judge agreement before/after, remaining candidates |
POST /review | Record a verdict (target_type, target_id, verdict, optional queue_id, critique, failure_mode_id, cluster_id, is_gold) |
GET /review/labels/trend | Labels over time, as shown on the Labelled tab |
Queue body fields: grain (trace | session), source (smart, recent_failures,
failure_mode, low_score, disagreement, random), window (24h, 7d, 30d, all),
size (1–200, default 25), filters (failureModeId or clusterId for failure_mode;
metric and scoreBelow for low_score; agentVersion), skipLabelled (default true),
exploreShare (0–0.5, default 0.2), fresh (default false; start a new queue without carrying
labelled items over, used by Review the next N). Full schemas are in the API reference.
Troubleshooting
| Symptom | Cause → fix |
|---|---|
| The queue is empty or shorter than the size you chose | Fewer items match than you asked for. Widen the Time window, choose another source, or turn off Skip items I or a teammate already labelled |
| Sessions shows nothing | Your traces don’t carry a conversation id, so there are no sessions to review. See Traces & sessions |
| My session verdicts don’t change judge alignment | By design: session verdicts are stored for the conversation, not copied to its traces. Review the traces too if you want them to count |
| Low judge scores returns nothing | Pick the judge (metric) to filter on; the default threshold is 0.5 |
| Regenerate replaced items I was about to label | Regenerate keeps only items that already have a pass or fail verdict. Label an item first if you want to keep it |
See also Annotations & alignment and the FAQ.