Triage your failures
The Triage view answers the question you have after every deploy: which failures should we
work on first? It places each active failure cluster by how many sessions it reaches and how fast
it is growing. It marks which clusters your agent can fix, and it ranks the fixable ones in a
Fix queue. It also counts silent failures, meaning sessions that returned ok but gave
the user a broken answer.
At a glance
| Where it lives | Diagnose → Failure Modes → Triage |
| Window | Follows the page’s time picker (Today, 24h, 7d, 30d, All time, custom). The default is 7d, compared with the 7 days before it |
| Key API routes | GET /clusters/triage?range=7d, GET /sessions?outcome=silent |
| What it needs | At least one completed clustering run (see Failure clustering). Fixability comes from each cluster’s remediation |
| Scope | Per agent |
The Failure Modes page has four views: Failure Modes (the cluster cards), Triage (this page), Atlas (see Failure atlas) and Issues & Taxonomy (see Issues & failure modes). Old links to the retired Scatter view open Triage.
A worked example
Take an e-commerce support agent. In the last 7 days it handled 4,812 sessions. One cluster,
check_inventory 5xx/504 cascade, collects sessions where the inventory service returned 504s
and the agent carried on anyway. It had 104 sessions in the previous 7 days and 286 in
this window. 121 of those 286 sessions ended with status ok, but a judge marked the answer as
failed, because the agent told customers an item was in stock without checking.
Here is how that cluster appears:
| Measure | Value | How it’s computed |
|---|---|---|
| Share | 5.9% | 286 ÷ 4,812 sessions in the window |
| Growth | +175% | (286 − 104) ÷ max(104, 5) × 100 |
| Silent share | 42% | 121 ÷ 286 |
| Fixability | Agent-fixable | Its remediation is a tool_resilience change in the agent: retry with backoff, then answer from cached stock |
On the matrix it sits high and to the right, in the FIX NOW quadrant. It gets a solid ring because the agent can fix it, and a centre dot because more than 25% of its sessions are silent failures.
Two other clusters in the same window show why it tops the Fix queue:
| Cluster | Sessions (prev) | Share | Growth | Silent | Priority |
|---|---|---|---|---|---|
| check_inventory 5xx/504 cascade | 286 (104) | 5.94% | +175% | 121 | 5.94 × 2.75 × 1.42 = 23.3 |
| Final answer contradicts tool result | 64 (0, new) | 1.33% | NEW | 52 | 1.33 × 3 × 1.81 = 7.2 |
| refund_api amount sent in dollars, not cents | 188 (171) | 3.91% | +10% | 54 | 3.91 × 1.10 × 1.29 = 5.5 |
The new cluster is smaller but gets the maximum growth weight. The refund cluster is chronic: it is large but barely growing. The inventory cluster is large, growing fast and mostly silent, so it ranks first. The formula is described in How the fix queue is ranked.
Read the triage view
The summary cards
| Card | What it shows |
|---|---|
| Failing sessions | Sessions in the window that errored or are silent failures, as a count and a percentage of all the agent’s sessions, with the change in percentage points from the previous window |
| Silent failures | Failing sessions that returned ok but carry a failure signal, and their share of all failing sessions. The bar under the number splits the clustered sessions by outcome. Click the card to open Where failures end up |
| New clusters | Clusters first seen in this window that had no sessions in the previous one |
| Agent-fixable | The share of clustered sessions that sit in clusters whose remediation is a change to the agent, plus how many clusters still need triage and how many fix runs are active |
The triage matrix
Each bubble is one active cluster with sessions in the window. Its size is the session count.
- Horizontal position is growth: the change in sessions compared with the previous window. The axis runs from −60% to +320%. Clusters beyond either end are drawn at the edge.
- Vertical position is share: the cluster’s sessions as a percentage of all the agent’s sessions in the window.
- The NEW lane on the right holds clusters first seen in this window with no sessions in the previous one. They have no growth figure, so they get a lane of their own and are not placed on the growth axis. With All time there is no previous window, so no cluster has a growth figure or counts as new: every bubble sits on the 0% line and the NEW lane stays empty.
Two dashed lines split the chart into four quadrants. The vertical line is 0% change. The horizontal line is the median share of the clusters on the chart, so “large” and “small” are relative to your other clusters in the same window. The FIX NOW quadrant is shaded.
| Quadrant | Position | What it usually means |
|---|---|---|
| FIX NOW | Growing, large share | A problem that is getting worse and already reaches many users. Start here |
| CHRONIC | Flat or shrinking, large share | A long-standing problem. Worth fixing, but it is not a new regression |
| EMERGING | Growing, small share | Small today and rising. Check it before it moves up |
| FADING | Shrinking, small share | Going away on its own or after a fix |
The ring around a bubble shows fixability:
| Ring | Meaning |
|---|---|
| Solid | Agent-fixable: the fix is a change to your agent (prompt, tool handling, schema) |
| Dashed | Downstream service: the failure lives in a service the agent calls. Neens tracks it, but the fix happens elsewhere |
| Dotted | Guardrail working as intended: a guardrail blocked something it should block |
| Dash-dot | Not fixable in this agent: the remediation says it can’t be fixed in the agent for another reason |
Grey outline, faint fill and a ? | Needs triage: no remediation yet, so Neens can’t say whether it is fixable |
A centre dot means at least 25% of the cluster’s sessions are silent failures. A green trail from the cluster’s previous position means a fix was merged and Neens is watching whether it holds (see Fix outcomes).
Hover a bubble for its numbers. Click it, or focus it with Tab and press Enter, to open its panel. The bucket chips above the matrix (Tool failures, Timeouts and so on) filter the matrix, the queue and the sessions table to one bucket.
The fix queue
The panel next to the matrix has three parts:
- Fix queue: agent-fixable and needs-triage clusters, ranked by priority. Each row shows its inputs (growth, sessions, silent count), so you can see why it ranks where it does.
- Fix merged · watching: clusters with a merged fix whose effect Neens is still measuring. They leave the queue while they are watched.
- Not fixable in this agent: downstream-service, guardrail and other not-fixable clusters.
The cluster panel
Clicking a cluster shows its session count (and the previous window’s), its share, its wasted spend, a sessions-over-time sparkline, and a What the user got bar that splits its sessions by outcome. It also shows the current remediation and what state it is in: proposed, in a fix run, watching after merge, verified, or regressed.
From the panel you can generate a remediation, start or view a fix run, build a regression set, or open Where it breaks → (see Where it breaks).
Silent failures and where failures end up
Click the Silent failures card to open Where failures end up. It is a flow chart from each failure bucket on the left to what the user got on the right: Errored, Silent failure, Recovered or Not confirmed. Band widths are session counts. Click a band to filter the sessions table to that bucket and outcome. Click an outcome on the right to filter by outcome only.
The sessions table
The sessions table under the view follows the window and filters above it. On this page it
includes an Outcome column by default. When you filter it to Silent failure, a banner
offers Add to dataset so you can collect those sessions for a regression set. To get the same
filter from the API, pass outcome=errored|silent|recovered|unconfirmed to GET /sessions. Each
returned row also carries its outcome.
How-tos
Triage a new cluster after a deploy
Set the window to cover the deploy
Pick 24h (or a custom range starting at the deploy) on the time picker. The comparison window is always the equal-length window immediately before it.
Look at the NEW lane first
Anything that appeared since the deploy lands there. A cluster first seen minutes after a release is the strongest lead you have. Its panel shows First seen so you can line it up with the deploy time.
Check how it fails
In the panel, read What the user got. A new cluster made mostly of silent failures means the agent answered wrongly without erroring. Those failures won’t show up in error-rate alerts.
Decide whether it’s fixable
A new cluster usually shows Needs triage. Click Generate remediation to draft one. Neens then classifies it as agent-fixable, downstream or guardrail, and the ring and queue update. See Remediations.
Find the step that broke
Open Where it breaks → to see which step fails first. See Find the choke point before fixing.
Find silent failures and add them to a dataset
Open the flow
Click the Silent failures card. Where failures end up opens under the cards.
Pick a band
Click the band from the bucket you care about to Silent failure, for example Agent output issues → Silent failure. The sessions table filters to that bucket and outcome.
Check a few sessions
Open two or three sessions from the table and confirm the answers really are wrong. A silent failure comes from a judge score, a human annotation or an issue label, so it is only as good as that signal.
Add them to a dataset
Click Add to dataset in the banner above the table and pick a dataset, or create one. See Datasets. To freeze them as a regression set for a fix, use Build regression set from the cluster’s panel instead.
How is it defined
How outcomes are defined
Every failing session gets exactly one outcome:
| Outcome | Definition |
|---|---|
| Errored | The session’s status is error. The user saw an error |
| Silent failure | Status ok, but something marks the answer as failed: a judge score of fail or below the agent’s Score threshold (see Clustering), a human annotation of fail, or an issue label other than none |
| Recovered | Status ok, no failure signal, and at least one span errored (a step, or a wrapper span around steps). The agent hit an error and handled it |
| Not confirmed | Everything else: status ok with no error and no failure signal yet, or no status. It is clustered with failures but nobody has confirmed it failed |
A silent failure can have no errored step at all, for example when the tool calls succeeded and the final answer was wrong. Not confirmed sessions often become silent failures once a judge runs over them or someone reviews them. See Judges and Review & annotations.
How growth is defined
Growth compares the cluster’s sessions in this window (n) with the equal-length window just
before it (prev):
growth = (n − prev) ÷ max(prev, 5) × 100The floor of 5 previous sessions keeps tiny clusters from swinging wildly. A cluster that goes from 1 session to 4 shows +60% ((4 − 1) ÷ 5), not +300%.
A cluster is new when it was first seen inside this window and had no sessions in the previous one. New clusters have no growth figure and go in the NEW lane. A cluster that existed before but had no sessions in the previous window is not new; its growth uses the floor, so 8 sessions from 0 shows +160%.
How fixability is defined
Fixability comes from the cluster’s top remediation:
| Remediation | Fixability |
|---|---|
| Actionable in the agent | Agent-fixable |
| Advisory (the cause is a downstream service) | Downstream service |
| Not actionable, because a guardrail blocked what it should | Guardrail working as intended |
| Not actionable for another reason | Not fixable in this agent |
| No remediation yet | Needs triage |
Regenerating or editing the remediation changes the cluster’s fixability. See Remediations.
How the fix queue is ranked
priority = share × growth weight × (1 + silent share)
share = cluster sessions ÷ all sessions in the window, as a percentage
growth weight = 3 for a new cluster
= 1 + min(2, max(0, growth) ÷ 100) otherwise
silent share = silent failures ÷ cluster sessionsSo growth can at most triple a cluster’s weight, and shrinking clusters are never pushed below their share. A cluster where every session is a silent failure counts double. The same score orders the queue in the UI and in the API response, so both always agree.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| The matrix is empty | No clustering run has finished, or no active cluster has sessions in the window | Widen the window, or run clustering (see When clustering runs) |
Every bubble shows ? | No cluster has a remediation yet | Click Generate remediation on the clusters at the top of the queue |
| Silent failures is 0 but users complain | Nothing scores the answers yet | Deploy a judge or review sessions, then check again |
| Many sessions are Not confirmed | Clustered sessions have no error and no signal either way | Same fix: a judge or human review turns them into silent failures or clears them |