Wasted spend
Wasted spend is what your agent spent after a session had already gone wrong. It adds up the cost of every step from the first failing step onward, in sessions that failed. Use it to put a dollar figure on a failure mode and to see which fix saves the most.
At a glance
| Where you see it | The Wasted spend tile on Agent Home; each cluster’s panel in Triage and the Failure atlas; the side panel of a failure mode’s Where it breaks tab |
| Key API route | GET /stats/wasted-spend?range=7d |
| What it needs | Token counts and model names on your LLM spans, and prices for those models (Settings → Model pricing, see Cost & model pricing) |
| Scope | Per agent, over the selected time window |
The Agent Home tile
The tile shows the total for the window, the share of all spend it represents, the change from the previous window, and a see clusters link to the Triage view. The i next to the title explains what is counted, names the biggest source (the cluster with the most wasted spend), and lists what was left out:
- failed sessions with no failing step to measure from, and
- models with no price, with a Set prices link.
The tile never shows $0 just because nothing could be measured. It shows — instead, with the reason:
| You see | Why |
|---|---|
| — and Set prices to see wasted spend | Failed sessions have steps after their first failure, but none of those steps ran on a priced model |
| — and No failing step to measure from | The window has failed sessions, but none of them has a failing step, for example silent failures where every step succeeded |
| ≥ $12.40 | Some steps after a failure ran on a model with no price. The figure covers only the priced steps, so the real total is at least this much. No change against the previous period is shown, because comparing two partial figures would be misleading |
| $0.00 | There were no failed sessions in the window, or the steps after each failure cost nothing (for example, tool calls that use no tokens) |
When the window has a very large number of failed sessions, the figure covers the most recent ones and the i says so. Neens does not scale a sampled figure up to guess the total.
A worked example
Take the seven-step session from Where it breaks. The
agent calls check_inventory, gets a 504, tries again, and answers anyway. A judge marks the answer
as failed, so the session is a silent failure. Its first failing step is step 3.
Suppose the LLM steps run on a model priced at $3.00 per million input tokens and $15.00 per million output tokens (example rates; yours come from your price table). Tool and guardrail steps here carry no tokens, so they cost $0.
| # | Step | Tokens in / out | Cost | Counted? |
|---|---|---|---|---|
| 1 | plan (LLM) | 1,800 / 150 | $0.00765 | No, before the failure |
| 2 | lookup_order (tool) | none | $0 | No, before the failure |
| 3 | check_inventory (tool) | none | $0 | Yes, the failing step |
| 4 | handle_tool_error (LLM) | 2,600 / 120 | $0.00960 | Yes |
| 5 | check_inventory (tool) | none | $0 | Yes |
| 6 | synthesize_answer (LLM) | 3,400 / 400 | $0.01620 | Yes |
| 7 | output_guardrail (guardrail) | none | $0 | Yes |
The arithmetic for the two LLM steps that count:
step 4: (2,600 / 1,000,000) × $3.00 = $0.00780
( 120 / 1,000,000) × $15.00 = $0.00180
--------
$0.00960
step 6: (3,400 / 1,000,000) × $3.00 = $0.01020
( 400 / 1,000,000) × $15.00 = $0.00600
--------
$0.01620
wasted spend = $0 (step 3) + $0.00960 + $0 + $0.01620 + $0 = $0.02580The whole session cost $0.03345, so 77% of it was spent after the session had already failed. Across the 286 sessions in the cluster that week, sessions shaped like this one add up to about 286 × $0.0258 ≈ $7.38.
If the failing step itself is an LLM call (for example, a model call that times out after consuming its input), its cost is included, because the failing step always counts.
How is it defined
Wasted spend for one session is the cost of every step that started at or after the first failing step, including the failing step itself. It applies only to sessions whose outcome is errored or silent failure (see outcome definitions).
- Steps are LLM calls, tool calls, retrievals and guardrail checks. Wrapper spans (agent, chain) are not steps.
- The first failing step is the first step, in start order, whose status is
error. - Cost is tokens × your model prices, the same way every other cost in Neens is computed. Steps with no tokens, such as most tool calls, cost $0.
- Recovered sessions are not counted. The agent handled the error, so the steps after it did useful work.
- Not confirmed sessions are not counted. Nothing yet says they failed.
- Failed sessions with no failing step are left out and counted separately. A silent failure where every step succeeded has no point to measure from. The tile and the API report how many there were.
A cluster’s wasted spend is the sum over its sessions. The Agent Home figure is the sum over all failed sessions in the window. For very large windows, Neens measures a sample of the most recent failed sessions and says that the figure is sampled.
Partial pricing
Neens never guesses a price (see Unpriced is a real state). When some steps ran on a model with no price:
- the figure covers only the priced steps and is marked partial (on the Agent Home tile, as ≥ before the amount),
- the unpriced models are named, with a Set prices link.
In the example above, if synthesize_answer ran on an unpriced model, wasted spend would show
$0.0096, partial, naming that model. If no contributing model is priced, there is no figure
at all: the tile shows —, and the API returns usd: null, not 0.
Prices apply to history. Once you set a price, wasted spend for past sessions uses it too (subject to the price’s effective date).
Reduce wasted spend
Find the biggest source
The tile’s i names the cluster that wastes the most. see clusters opens Triage, where each cluster’s panel shows its own wasted spend.
Look at the steps after the failure
Open the cluster’s Where it breaks tab. The hatched cells are the wasted steps. A long tail of LLM steps after the failure is where the money goes.
Make the agent stop or fall back sooner
Common fixes:
- Fail fast. When a required tool fails, end the run with a clear message instead of re-planning and calling the model again.
- Retry with backoff, then stop. A retry that fails the same way (step 5 in the example) adds cost and nothing else. Cap it.
- Use a fallback. Answer from cached or partial data, and say so, rather than generating a full answer on missing inputs.
- Check earlier. Move a guardrail or validation before the expensive steps, not after.
Fix the root cause
Stopping sooner reduces waste. Fixing the choke point removes it. Generate a remediation for the cluster and prove it with an eval-verified PR.
Watch it drop
After the fix merges, the cluster moves to Fix merged · watching in Triage, and the tile’s change from the previous window shows the effect. See Fix outcomes.
Wasted spend covers the cost of your agent’s own steps. It doesn’t include what Neens spends running judges over those sessions. For that, see Cost & Quality.