GuidesIssues & failure modes

Issues & failure modes

A failure mode is a named, defined way your agent goes wrong — “Tool-call argument hallucination”, “Refund policy misquoted”. The set of them is your agent’s taxonomy. An Issue is a failure mode in operation: the Neens classifier assigns each failing trace to a mode, and the volume rolls up into Issue cards with a lifecycle, a trend, and a path to a fix. Both live on the Failure Modes page, under the Issues & Taxonomy tab.

At a glance

Where it livesDiagnose → Failure Modes → Issues & Taxonomy
Key API routesGET /taxonomy/failure-modes, POST /taxonomy/failure-modes/import, POST /taxonomy/failure-modes/import-file, GET /taxonomy/issues, PATCH /taxonomy/issues/{id}/lifecycle, POST /flywheel/failure-modes/{id}/generate-eval, POST /taxonomy/failure-modes/{id}/links
What it needsA taxonomy (import the starter library or confirm clusters); an agent LLM connection (Settings → LLM connections) for the classifier
ScopePer agent

The Failure Modes page has two tabs. The Failure Modes tab (covered in Failure clustering) shows what the machine found by clustering raw traces. This page covers the Issues & Taxonomy tab: the curated, named catalogue your team owns — plus the operational Issues rollup of how often each named failure is actually happening.

Neens organizes failures at two levels. A fixed set of about ten broad buckets (Timeouts, Rate limits, Auth failures, Guardrails tripped, Upstream service errors, Tool failures, Data/parsing errors, Agent output issues, Other) groups the top-level rollups and chart legends. Underneath, the open-ended failure modes are the specific patterns — the level you name, classify against, and manage.

The taxonomy

Every failure mode carries a name, a definition (“what distinguishes this failure mode?”), a curated severity (High / Medium / Low), an optional owner, compliance tags, a curation status, and a provenance badge:

Provenance badgeMeaning
Platform defaultShipped with the Neens starter library (or bulk-imported from your own label list)
CustomAdded or confirmed by your team
Auto-discoveredSurfaced by failure clustering
Curation statusMeaning
CandidateSuggested, not yet reviewed
ConfirmedReviewed and adopted by your team
MonitoringKept under watch
ArchivedRetired — excluded from classification and the Issues rollup

Build your taxonomy

The Issues & Taxonomy tab header carries two buttons for managing the catalogue itself:

  • + New issue type — the primary button. Opens a dialog to add a single failure mode by hand: a name, definition, severity, optional owner, and compliance tags (POST /taxonomy/failure-modes).
  • Manage taxonomy — opens the full Taxonomy table: every mode with its provenance, curation status, and bulk actions (import, edit, merge, split, delete — see Curate it over time below). Use this when you’re tidying the catalogue rather than adding one issue.

Ways to populate the taxonomy:

  • Import starter library — the fastest start. On the empty state (or via POST /taxonomy/failure-modes/import), seed a set of common agent failure modes, or post your own label list in the same call. Imports are idempotent: re-importing refreshes definitions without duplicating, and never overwrites a status your team has since changed.
  • Confirm candidate / Promote to issue — turn a discovered cluster on the Failure Modes tab into (or link it to) a named issue. You must review at least 3 exemplar traces from the cluster first, so every mode is grounded in real evidence. See Promote a cluster to an issue for the cluster-side how-to.
  • Create manually — the + New issue type button above, or POST /taxonomy/failure-modes directly.

Import a taxonomy file (CSV / JSON)

Most teams already keep their failure taxonomy in a spreadsheet. Click Import on the Issues & Taxonomy tab to upload it as a CSV or JSON file — a thin path over the same idempotent bulk-import, so re-uploading updates modes in place instead of duplicating.

The CSV columns are:

ColumnRequiredNotes
nameYesThe failure mode’s name
definitionWhat distinguishes this failure mode
severityOne of high, medium, low (blank is allowed)
compliance_tagsOne or more tags separated by ; or ,
keyA stable id used for idempotent re-import (defaults to a slug of the name)

Common spreadsheet header names are accepted as aliases — for example Failure Mode → name, Description → definition, and Tags → compliance_tags. Grab the exact header row from the Download template link in the Import modal (backed by GET /taxonomy/failure-modes/import-template).

name,definition,severity,compliance_tags,key
Hallucinated Price,"Agent states a product price not present in the catalog tool result",high,accuracy;grounding,hallucinated_price
Wrong Order Lookup,"Agent looks up the wrong order id for a returns request",medium,reliability,wrong_order_lookup

Malformed rows never sink the upload. A row missing name or carrying an invalid severity is reported back individually — you get a created / updated / skipped summary where each skipped row lists its row number and the reason — while every valid row still imports. Nothing is silently dropped.

Imported modes land as Candidate (provenance: seeded) and are idempotent by key/name: re-uploading the same file refreshes definitions and bumps each mode’s version rather than creating duplicates, and it never overwrites a status your team has since changed.

Import from the command line or CI

For automation, post the file body directly to the raw-body endpoint (the request body is the file — this is not a multipart form upload). Requires the manage_taxonomy permission — any member or admin can import, edit, and delete taxonomy modes; viewers are read-only:

curl -X POST "$NEENS_BASE_URL/taxonomy/failure-modes/import-file?format=csv" \
  -H "Authorization: Bearer $NEENS_API_KEY" \
  -H "Content-Type: text/csv" \
  --data-binary @failure-modes.csv

format is optional (csv or json) — Neens infers it from the Content-Type and content when omitted — and status defaults to candidate. The response is {created, updated, skipped: [{row, name, reason}], failureModes, format}.

The pre-existing strict-JSON endpoint POST /taxonomy/failure-modes/import (body {"modes": [...]}, or {"include_starter_library": true}) remains the programmatic path when you’re generating the list rather than uploading a file.

Curate it over time

Click any row in the Taxonomy table to open its detail drawer:

  • Edit name, definition, severity, owner, and status (every change bumps the mode’s version).
  • Merge two modes — the source’s cluster evidence and exemplars fold into the target and the source is archived.
  • Split a mode into a new one, moving selected cluster links across.
  • Delete a mistaken or obsolete mode outright (DELETE /taxonomy/failure-modes/{id}) — distinct from merge (which archives the source) and archive (a status change); its cluster-evidence links are removed with it.
  • Cluster evidence — the clusters backing a mode. Links marked auto were proposed by Neens (after each clustering run, a cluster is linked to a mode when a clear majority of its classified members carry that mode); links you create yourself are never overwritten by the machine.

Every taxonomy change — create, import, edit, delete — is recorded in the tenant audit log (GET /audit, admin-only) with who did it and when.

How traces get classified

Classification is done by the built-in Issue Classification judge, provisioned enabled for every agent:

  • Every 30 minutes, Neens takes the agent’s recent failure-set sessions (the same selection criteria clustering uses — errors, below-threshold scores, human “fail” labels) that haven’t been classified yet. Healthy traffic is never classified, and each session is classified once.
  • For each session, the classifier shows an LLM the session’s evidence (input, output, messages, metadata) and the list of your non-archived failure modes with their definitions (up to 40), and asks for the single best-matching mode — or none if no listed mode clearly applies. Answers are snapped to exact mode names; the model can never invent a mode.
  • Each classification is stored as a score (metric_key = issue_class) carrying the mode name as its label, the classifier’s confidence, and a one-sentence reason.
  • Spend is bounded: at most 500 sessions per day per agent by default, drawn from the classifier’s own daily budget.

The classifier uses the agent’s LLM connection, like every LLM feature. With no connection configured, no classification happens — the taxonomy still exists and is fully editable, but Issues show zero classified traces. With no (non-archived) taxonomy, the classifier doesn’t run at all, so it never burns budget with nothing to classify against.

Issues

The Issues rollup shows each non-archived failure mode as a card with its lifecycle state, severity, classified trace count, mean classifier confidence, linked clusters, a trend sparkline, and its fix / eval-gate status. A mode with zero classified traces still appears — an Issue is a mode, whether or not it’s currently occurring. Scope the rollup with the time-range picker (Today, 24h, 7d, 30d, All time, or a custom range; default 7d), and filter by search, severity, and lifecycle state.

The analytics summary

A summary strip sits above the cards: total issues, active vs watching counts, affected traces (sessions caught by an active issue in the window), and a severity / lifecycle breakdown across the active ones. These numbers are computed on the server over your entire taxonomy in the selected window, before the search/severity/state filters are applied — so narrowing the list to severity=high never changes what the summary reports. Widening or narrowing the time range does change it, since it re-windows which sessions count as “active”.

Active vs Watching

Every issue falls into exactly one of two buckets for the selected window:

  • Active — classified traces > 0 in the window. Active issues render as full cards, sorted by classified trace count descending — the noisiest problem is always first, so triage order tracks actual impact rather than alphabetical order or creation date.
  • Watching — zero classified traces in the window. These are modes in your taxonomy that simply haven’t happened (yet, or in this window). They render as a collapsed, compact one-line list — a severity dot, the name, its one-line definition, and “0 traces” — so a taxonomy of a hundred named modes doesn’t bury the handful that actually need attention. Expand the Watching section to see them all, or use Expand all / Collapse all at the top of the tab to open or close every group (active cards and the watching list) at once.

Example. Your taxonomy has 30 failure modes. This week, 6 of them classified at least one trace — one dominant mode with hundreds of sessions, the rest much smaller — and the other 24 classified none. The summary reads 6 active · 24 watching; the 6 active cards are sorted loudest-first, and the 24 watching modes sit collapsed underneath until you expand them or widen the window.

A Watching row’s action is deliberately lighter than an active card’s: a muted Definition link (so you can remind yourself what it means without opening the drawer) and an accent Create eval → link — pre-emptively locking the failure in as a regression test before it ever fires in production. View traces and Generate fix don’t apply with zero traces, so they’re not shown until the mode goes active.

Lifecycle

StateMeaning
OpenNewly surfaced, not yet triaged
AcknowledgedSomeone has seen it and taken ownership
InvestigatingBeing actively looked into
MitigatingA fix is in progress
ResolvedAddressed — no longer expected to recur
MutedDeliberately silenced

Move an Issue between states with the Lifecycle actions menu on its card (or PATCH /taxonomy/issues/{id}/lifecycle). Any state can transition to any other; the one exception is a mode whose curation status is Archived — it’s out of the operational loop entirely. Resolved and Muted Issues are hidden from the default list and suppressed from dashboards and insight highlights; select those states in the filter to bring them back.

Neens reconciles the lifecycle for you. Every 30 minutes: an Issue whose linked remediation reaches verified (its proof eval passed) is auto-transitioned to Resolved (unless you muted it), and a Resolved Issue that receives new classified traces is reopened to Open — a fixed failure that recurs never stays silently “resolved”.

From an Issue to action

An active Issue card carries three actions:

  • View traces — opens the traces the classifier assigned to the Issue in the selected window, each with its classification confidence and reason.
  • Generate fix → View fix — a stateful action: Generate fix drafts a typed remediation for the Issue when it has none; once one exists, the same spot becomes View fix, linking to it with its current status. See Remediations.
  • Create eval — locks the failure in as a regression test (below). Once the Issue already has a gate, the card shows the gate’s status here instead.

A watching Issue’s row keeps only Definition and Create eval — see Active vs Watching above.

Create an eval from an Issue

Click Create eval on the Issue card

Neens needs real evidence first: if the mode has no classified traces, linked clusters, or exemplars yet, the call is rejected with a clear message rather than creating an eval that would trivially pass forever.

Neens builds the pieces

One guided step (POST /flywheel/failure-modes/{id}/generate-eval) materializes a dataset captured from the failing traces, drafts an LLM judge that scores whether the failure is absent (high score = good), and wires them together into an eval gate.

Enable the gate

The gate starts as a draft. Enable it on the Eval Gates page to start catching regressions — if the failure creeps back, the gate fails. An Issue that already has a gate shows it on the card instead of the Create eval action.

Troubleshooting

SymptomCauseFix
”No issue taxonomy yet”The agent has no failure modesClick Import starter library, Import a CSV/JSON file of your own taxonomy, add one manually, or confirm a cluster
Issues all show zero tracesNo LLM connection (classifier can’t run), or no failing traffic in the selected windowConfigure Settings → LLM connections; widen the time range
”No classified issues in this window”Taxonomy exists but nothing was classified into it recentlyNormal when the agent is healthy; check View traces on the Failure Modes cards for unclassified failures
Everything is in Watching, nothing ActiveNo traces were classified into any mode in the current windowWiden the time range, or check the LLM connection / classifier as above
Summary tiles don’t match the cards you can seeThe summary counts the whole taxonomy in the window, before your search/severity/state filtersExpected — clear the filters to see the same population the summary describes
Create eval fails with “no evidence sessions”The mode has no classified traces, linked clusters, or exemplarsWait for classification, link a cluster, or add exemplars first
A resolved Issue reopened by itselfNew traces were classified into it after it was resolvedThat’s the recurrence guard working — investigate the new traces