GuidesWhere it breaks

Where it breaks

The Where it breaks tab on a failure mode shows the cluster’s recent sessions as a trace barcode: one row per session, one cell per step. Failures that look random in a trace list usually line up here, and the column where they line up is the step to fix.

At a glance

Where it livesOpen a failure mode (from the Failure Modes cards, or Where it breaks → in the Triage panel), then the Where it breaks tab
Key API routeGET /clusters/{cluster_id}/trajectories?limit=50 (up to 200)
What it showsThe cluster’s most recent sessions, 50 by default
ScopePer agent. A cluster from another agent returns 404

Read the barcode

Rows and cells

Each row is a session, with a coloured marker at the left showing its outcome. Rows are grouped by outcome (errored, then silent failure, recovered and not confirmed) and are newest first within each group. Each cell is a step: an LLM call, tool call, retrieval or guardrail check, in the order they started. Agent, chain and other wrapper spans are left out, because they contain the steps rather than being steps.

CellMeaning
GreyThe step succeeded
Solid bucket colourThe first failing step: the first step whose status is error
Light bucket colourA later step that also errored
HatchedA step that ran after the failure in a session that did not recover. This is where wasted spend comes from
Green outlineA step that ran after the failure in a session that recovered

A row with a dashed outline has no failing step. See Rows with no failing step.

Column headers show the most common step name in that column. Sessions with more than 40 steps have their middle collapsed into one marker showing how many steps were hidden. Scroll sideways inside the card to see long sessions. Click a row to open that session.

Alignment

The Align toggle changes how rows line up:

  • Align at start puts every session’s first step in the first column. Use it to compare what sessions did overall.
  • Align at first failure shifts each row so the first failing steps share one column. Steps before the failure are to its left and steps after it are to its right. Use it to see whether sessions break at the same step and what they do next. A dashed line marks the aligned column.

Rows with no failing step can’t be aligned on a failure, so they stay aligned at their start.

The side panel

ItemWhat it tells you
Choke pointThe step name that is most often the first failing step, and the share of the sessions with a failing step that fail there first
After the failureThe average number of steps that ran after the first failure, in sessions that did not recover
No failing stepHow many of the shown sessions have no errored step
Wasted spendWhat the whole cluster spent from the first failing step onward over the last 7 days, and whether the figure is partial. It shows — with Set prices to see wasted spend when none of the steps could be priced, and — with No failing step to measure from when none of the failed sessions has a failing step. For a very busy cluster it is based on the most recent sessions, and says so

A histogram above the rows shows where first failures happen: the number of sessions whose first failure is at each step position.

A worked example

The check_inventory 5xx/504 cascade cluster from the e-commerce support agent in Triage your failures has a typical session with seven steps:

#StepKindStatus
1planLLMok
2lookup_ordertoolok
3check_inventorytoolerror (504)
4handle_tool_errorLLMok
5check_inventorytoolerror (504)
6synthesize_answerLLMok
7output_guardrailguardrailok

In the barcode, cell 3 is solid blue (Tool failures), cell 5 is light blue, and cells 4 to 7 are hatched, because the session ended as a silent failure: it told the customer the item was in stock. Its steps after the failure count is 4.

With Align at first failure on, most rows line up on check_inventory, and the side panel reads:

  • Choke point: check_inventory, where 81% of the sessions with a failing step fail first, usually on the first call to it (step 3).
  • After the failure: 4.1 steps. Unrecovered sessions keep going, including a retry that fails the same way.
  • No failing step: 9 of 50. These are silent failures where every step succeeded and the issue classifier flagged the answer.

This tells you two things. The fix belongs at step 3, in how the agent handles a check_inventory error, not in the final answer prompt. And the retry at step 5 makes things worse, because it adds steps (and cost) without ever succeeding.

Find the choke point before fixing

Open the failure mode

From Triage, click the cluster and choose Where it breaks →. Or open the failure mode from its card and pick the Where it breaks tab.

Align at first failure

Switch the toggle to Align at first failure. If most rows line up on the same step name, that step is the choke point, and the side panel names it.

Check the share

A choke point at 80% or more is a single fix target. If the share is low and first failures are spread across several columns, the cluster may mix different problems. Compare it with its neighbours in the Failure atlas before you write a fix.

Look at what happens after

Hatched cells after the failure show work the agent did after the session was already lost. A long hatched tail suggests the fix should also stop the run early, for example by failing fast or answering from a fallback when a required tool errors.

Check the rows with no failing step

If many rows have a dashed outline, part of the cluster fails without any errored step. Fixing the choke point won’t help those sessions. Open a few to see what the answer got wrong.

Fix it

Put the step name in your remediation or fix run so the fix targets that step. See Remediations and Eval-verified PR.

Rows with no failing step

A session is in a failure cluster because it failed, but not every failure has an errored step. The common case is a silent failure: every tool call succeeded and the final answer was wrong. A judge, a reviewer or the issue classifier caught it.

These rows:

  • have a dashed outline and no solid cell,
  • stay aligned at their start when you align at first failure,
  • are left out of the choke point and the steps-after-failure average,
  • are not counted in wasted spend, because there is no step to measure from.

If most of a cluster’s rows look like this, the problem is in what the agent says rather than a step that breaks. Look at the answers, and consider a judge for this failure.