Scorer lifecycle
A scorer lifecycle is a lightweight way to curate your judges as they mature — from a prompt you’re still tuning, to a trusted metric everyone relies on, to one you’ve retired. It’s a curation signal only: staging a judge never changes what it scores. What it does change is which scores show up by default, so an agent’s Scores page stays focused on the metrics that matter instead of every ad-hoc experiment.
At a glance
| Where | Stage a judge on the Judges page; the Finalized / Experimental / All filter and per-run views live on the Scores page and each judge’s run history |
| Key API routes | PATCH /judges/{id}/lifecycle, GET /scores?lifecycle=, GET /scores?runId=, GET /judges/{id}/compare, GET /eval-runs/export |
| Scope | Agent-scoped — you can only stage a judge visible to your agent; comparisons and exports are validated against your agent |
| Changes scores? | No. Every capability here is curation, filtering, comparison, or export — it never re-grades a trace or edits a stored score |
The stages
A judge moves through three stages:
| Stage | What it means | Default view |
|---|---|---|
| Experimental | A work-in-progress scorer you’re still tuning. New judges you create start here. | Hidden from the Scores page by default |
| Finalized | Reviewed and trusted for everyday use. The built-in library judges (Faithfulness, Primary Score, and the rest) ship Finalized. | Shown by default |
| Archived | Retired — its score history is kept, but it’s out of your working set. | Hidden from the Scores page by default |
Change a judge’s stage
Open the Judges page
Find the judge you want to promote or retire on the Judges page.
Set its stage
Change the judge’s lifecycle stage — typically promoting an experimental scorer to finalized once you trust its verdicts, or archiving one you no longer use. The move is one-directional in spirit (experimental → finalized → archived), but you can set any stage at any time, including re-opening an archived judge back to experimental.
The change is recorded in the activity/audit trail and takes the same permission as configuring a judge, so a viewer can’t silently re-stage someone else’s scorer.
Via the API, PATCH /judges/{id}/lifecycle with a body of
{"stage": "experimental" | "finalized" | "archived"} (any other value is a 422). A
agent-scoped caller can only re-stage a judge visible to its agent; anything else is a 404.
Keep the Scores page focused
Because new judges start experimental, the Scores page defaults to showing Finalized scorers only — so an in-progress judge doesn’t clutter everyone’s metrics until you promote it. A Finalized / Experimental / All filter switches the whole page:
- Finalized (default) — trusted, everyday metrics.
- Experimental — only the work-in-progress scorers you’re tuning.
- All — every judge, regardless of stage.
A score whose judge can’t be resolved (legacy rows, built-in ragas metrics) is treated as
Finalized, so relaxing the filter only ever adds experimental rows — it never hides a
legitimate metric. See Scores → Lifecycle filter.
Compare, scope, and export runs
For ad-hoc batch eval runs — the manual Run now runs, each tied to one judge version — a judge’s run history lets you audit and compare what individual runs produced. Each of these reads the run’s own captured snapshot, so a later run overwriting the shared score row never corrupts an older run’s numbers.
Scope to one run
View exactly the scores a single ad-hoc run produced instead of the metric’s current live scores.
API: GET /scores?runId=<evalRunId>.
Compare two versions
Iterating on a judge’s prompt? Run the old and new versions,
then compare their two runs to see, per target, which sessions improved, regressed, or
stayed unchanged — with each side’s score and reason and the average delta. API:
GET /judges/{id}/compare?runA=<runId>&runB=<runId>. A run from another judge or agent is a 404.
Export a run with input and output
Download a run’s (or several runs’) scores as CSV or JSON, with each target’s input and
the agent’s output alongside the score, label, threshold, and reason — for spreadsheets, offline
review, or sharing a regression. API: GET /eval-runs/export?runs=<id>,<id>&format=csv|json.
Ad-hoc runs, not continuous scorers. Per-run scoping, version comparison, and export are built for bounded batch runs. A continuous scorer re-grades traces in place and overwrites each trace’s score row, so a single “run” isn’t a meaningful boundary for it — compare and export are aimed at the manual runs you launch to evaluate a change.
How it works
Scorer lifecycle is entirely additive on top of judges and scores:
- Stage lives on the judge as
lifecycle_stage. Pre-existing judges (and any judge whose name can’t be resolved to a stage) are treated as Finalized, so nothing disappears when the feature first appears. - The Scores filter resolves which judge names in your agent are experimental or archived and filters score rows by that set — a purely read-side filter that works the same on SQLite and ClickHouse deployments.
- Per-run scoping, comparison, and export read the run’s per-target snapshot (not the live, overwrite-in-place score table), so the numbers you compare or export always reflect what that run actually scored.
Related
- Judges — create, version, enable, and run scorers.
- Scores — the score catalogue, the lifecycle filter, and per-run views.
- Continuous evaluation — the always-on scoring these ad-hoc tools contrast with.
- Pre-prod evaluations — run your trusted (finalized) judges against a candidate release before you ship.