Cost–quality frontier
The proof tabs on Cost & Quality rest on one picture: a cost–quality frontier. Every candidate model is a dot — the cheaper it is, the further left; the better it scores, the higher up. The line through the best dots is the efficient frontier: the most quality you can buy at each price. It answers the question underneath both proof tabs — are we overpaying for quality, and what is the single move that saves money without moving the quality needle?
The same maths drives two tabs, pointed at two different models:
- The Judge model tab ranks your LLM judges’ scorer models — quality is how well each scorer agrees with your human ground truth. See Cost & Quality and Cost Optimization.
- The Agent model tab ranks candidate agent models — quality is a stable pass rate on a frozen golden suite. See Model sweeps and Sweep decisions.
Both read the same frontier, so “clears the bar”, “recommended”, and “no measurable quality loss” mean exactly one thing across the two tabs.
At a glance
| Where | The Judge model and Agent model tabs on the Cost & Quality page |
| Key API routes | GET /cost-optimization/frontier (judge scorer models), GET /model-sweeps/{id}/frontier (agent models) |
| Both need a quality axis | The Cost frontier needs human gold labels (see below); the Sweeps frontier needs a finished sweep with judges |
| Reads only | Nothing is stored and nothing is spent — every number is computed on read. Neens never switches a model for you |
| Scope | The agent you’re viewing; anyone with read access can see it |
A frontier is a recommendation, not an action. Neens names the cheapest model that clears your bar without a measurable quality loss; it never switches a scorer, promotes an agent model, or changes a deployment. You make the move — and the honest way to make it is behind a pre-prod eval so it can’t regress in production.
How to read the frontier
The axes are always the same shape — cost on the X-axis (right = more expensive), quality on the Y-axis (up = better) — and every dot falls into one role.
| On the chart | What it means |
|---|---|
| Your model now | The one you’re running today — the incumbent every recommendation is measured against. |
| On frontier | Nothing cheaper is at least as good. You’re getting the best quality available at that price — there’s no free saving here. |
| Recommended | The single best move: the cheapest model that is strictly cheaper than your current one, clears the acceptance bar, and is not measurably worse (see the noise band below). |
| Overpaying | A model off the frontier: something cheaper is at least as good, so the extra spend buys you nothing. |
| Quality risk / below the bar | Below the acceptance bar — cheap, but not good enough to consider. |
| Quality unknown | Its cost is known but there’s no quality signal for it yet, so it can’t be placed on the frontier. It’s shown, never guessed at. |
The acceptance bar
The acceptance bar is the minimum quality a model must clear to be a candidate at all — a horizontal line on the chart. A model below it is a quality risk no matter how cheap.
- Cost tab (judges): the bar is agreement with your experts, default 80%. Override it
per view with the bar control (or
?bar=on the API, a0–1rate). - Sweeps tab (agent models): the bar is the sweep’s pass-rate gate — the one you committed to when you launched it, defaulting to the 90% deployment success threshold. See Set the quality bar.
The noise band — why “cheaper” needs a confidence interval
Quality is measured on a sample, so it is never a single exact number — it’s a point estimate with a confidence interval around it (a 95% Wilson interval, the same statistic the sweep comparison uses). That interval is the noise band.
This is what makes an honest “no quality loss” claim possible:
- If the cheaper model’s interval overlaps your current model’s, the two are not statistically distinguishable at this sample size — switching is a no measurable quality loss move. That is the honest version of “cheaper, without compromising quality.”
- If the intervals don’t overlap, Neens shows the real, signed quality delta instead of hiding it — the cheaper model either measurably improves or measurably regresses.
- A model that is distinguishably worse is never recommended, however much it saves.
Unpriced is never $0, and never wins “cheapest.” If a model has no rate in your
price table, Neens doesn’t know its cost — it’s shown as unpriced
(—), excluded from the frontier line, and can never be named the cheapest move. A self-hosted
model you run yourself is a legitimately priced $0, which is a different thing entirely. Add a rate
under Settings → Model pricing to bring an unpriced model onto the chart.
Cost tab: the judge frontier
On the Cost tab, each dot is a scorer model one of your judges can run on, and quality is that scorer’s agreement with your experts. The chart answers: can this judge do its grading on a cheaper model without disagreeing with my humans any more than it already does?
There is one independent frontier per judge — you can only swap a judge’s scorer model within that judge, so a recommendation never proposes moving one judge’s grading onto a different judge’s model.
You need gold labels first
Without human gold labels there is no quality axis, so there is no frontier. Judge quality here is judge↔expert agreement, and that only exists once experts have labelled sessions. A judge with no gold-backed labels shows quality unknown for every scorer model and proposes no switch — an honest empty state, never a fabricated agreement number.
To give the Cost frontier a quality axis, produce expert truth:
Label sessions in the Review queue
Open the Review queue and record pass/fail verdicts on the sessions it prioritizes. See Annotations & review.
Make the decisive ones gold
Mark authoritative labels gold, or have a principal (expert) reviewer label them — gold and principal labels are the expert truth the frontier measures against.
Let alignment compute
Judge↔expert alignment recomputes nightly (and on demand). Once a judge has gold-backed alignment on two or more scorer models, its frontier can propose a switch between them.
Find and apply a “switch & save”
Open the Judge model tab
Cost & Quality → Judge model. Candidate scorer models are ranked against the bar; the Recommended moves list on the Overview tab ranks the available moves across every judge by monthly saving, each showing the quality delta against that judge’s noise band.
Read the recommended move
The recommendation names the cheapest scorer model that is strictly cheaper than the judge’s current one, clears the agreement bar, and sits inside the current model’s noise band. It comes with the projected monthly saving (the window’s spend scaled to 30 days) and the quality delta — with a no measurable quality loss claim when the intervals overlap, or the real signed delta when they don’t.
Corroborate before you trust it
Read the point’s failure-class metrics — precision and recall — next to the headline agreement (see the honest limitation below). A high agreement on a stream with few real failures deserves a second look.
Make the switch, then verify it
Neens doesn’t flip the scorer for you. Reassign the judge to the cheaper scorer model yourself (Cost Optimization → Point a judge at the new scorer), and gate the change behind a pre-prod evaluation so it can’t quietly regress your scoring in production. The saving shows up on your LLM-judge cost the next time the page loads.
An honest limitation you should know
The two scorer models are compared over the sessions each happened to score, not a shared holdout. Neens keeps one score per session, so a judge’s scorer models are each measured over their own gold-labelled sessions — the comparison is confounded by which sessions each one scored. This is surfaced, not hidden: each point carries its own gold count (a thin point is visible), and a recommendation is only ever “cheaper and within the noise band.” Treat a recommended switch as a strong hint, corroborated by precision/recall, not a proof — and confirm it with a pre-prod eval before you rely on it. The current scorer model shown is a heuristic: the busiest one in the window.
Sweeps tab: the agent-model frontier
On the Sweeps tab, each dot is a model sweep arm — a candidate model for your agent — and quality is the arm’s stable-pass rate (the composite judge score over the frozen eval set). The chart answers: which cheaper model still clears the bar for this agent?
Because a sweep already measures quality with a Wilson interval, cost per case, and a pass-rate gate, the frontier reuses those numbers directly — it recomputes nothing. Read it exactly like the judge frontier:
- The recommended arm is the cheapest one that beats your baseline (the declared incumbent — the arm you listed first, overridable), clears the bar, and is inside the baseline’s noise band.
- Because a sweep is an experiment, not production traffic, the recommendation reports a per-case saving and, honestly, no monthly figure — inventing production volume for an experiment would be a lie.
- An arm with no data (its endpoint was unreachable) is shown as such, never as 0%.
For the full leaderboard, per-agent verdicts, regression drill-down, and the shareable decision, use the comparison view — see Sweep decisions.
A void sweep draws no frontier. If a sweep’s arms didn’t run under identical conditions, it is void: the arms’ own numbers are still shown as facts, but the frontier is empty, no arm is classified, and no switch is proposed. A comparison of arms that measured different things is the most convincing wrong answer this feature could produce, so Neens declines it rather than draw a chart that would be screenshotted into a decision doc. Fix what drifted and run a fresh sweep.
How it works
- Cost frontier (
GET /cost-optimization/frontier) joins the per-scorer cost from Cost Optimization with judge↔expert agreement from alignment, one frontier per judge, over a window (?window=7dby default, or?start=&end=). Agreement carries a Wilson confidence interval — the noise band. - Sweeps frontier (
GET /model-sweeps/{id}/frontier) is derived entirely from the sweep’s own comparison — same reconcile, same bar, same baseline — and recomputes no quality. - The recommended move is, in both cases, the cheapest priced model that is strictly cheaper than your current one, clears the acceptance bar, and whose quality interval overlaps the current one’s (or is better). Overlapping intervals ⇒ no measurable loss; a measurable gap ⇒ the real delta; a distinguishably worse model ⇒ no recommendation.
- Everything is computed on read. Changing the window or the bar recomputes the frontier; there’s nothing to refresh or rebuild, and nothing is spent.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| The Cost frontier is empty / every dot says quality unknown | No gold/expert labels for these judges yet | Label sessions and mark them gold in the Review queue, then let alignment compute |
| A judge shows a single dot and no switch | It has only run on one scorer model, or only one has gold-backed quality | Score the same judge on a second scorer, over sessions that have expert labels |
| A model shows — in cost and never wins | Its model has no price in your catalogue | Add a rate under Settings → Model pricing; it is unknown, not free |
| The recommendation says “no measurable quality loss” but the delta isn’t zero | The quality intervals overlap, so the sample can’t distinguish the two | That is the honest claim. For more certainty, gather more gold labels (Cost) or golden items (Sweeps) |
| The Sweeps frontier is blank with a void banner | The sweep is void | Fix what drifted mid-sweep and run a fresh one; a void sweep is intentionally not comparable |
| You applied a recommended switch and quality dropped | A recommendation is a strong hint, not a proof — especially on the Cost tab’s disjoint-sample join | Always gate a switch behind a pre-prod eval; revert if it regresses |
Related
- Cost Optimization — the Cost tab: what your agent and judges spend, and how to add a cheaper scorer
- Model sweeps / Sweep decisions — the Sweeps tab: running the comparison and reading the verdict
- Annotations & review — how to produce the gold labels the Cost frontier needs
- Cost & model pricing — where a price comes from, and how unpriced models are handled
- Pre-prod evaluations — how to ship a model switch behind a gate that blocks a real regression