Issues & failure modes
A failure mode is a named, defined way your agent goes wrong — “Tool-call argument hallucination”, “Refund policy misquoted”. The set of them is your agent’s taxonomy. An Issue is a failure mode in operation: the Neens classifier assigns each failing trace to a mode, and the volume rolls up into Issue cards with a lifecycle, a trend, and a path to a fix. Both live on the Failure Modes page, under the Issues & Taxonomy tab.
At a glance
| Where it lives | Diagnose → Failure Modes → Issues & Taxonomy |
| Key API routes | GET /taxonomy/failure-modes, POST /taxonomy/failure-modes/import, POST /taxonomy/failure-modes/import-file, GET /taxonomy/issues, PATCH /taxonomy/issues/{id}/lifecycle, POST /flywheel/failure-modes/{id}/generate-eval, POST /taxonomy/failure-modes/{id}/links |
| What it needs | A taxonomy (import the starter library or confirm clusters); an agent LLM connection (Settings → LLM connections) for the classifier |
| Scope | Per agent |
The Failure Modes page has two tabs. The Failure Modes tab (covered in Failure clustering) shows what the machine found by clustering raw traces. This page covers the Issues & Taxonomy tab: the curated, named catalogue your team owns — plus the operational Issues rollup of how often each named failure is actually happening.
Neens organizes failures at two levels. A fixed set of about ten broad buckets (Timeouts, Rate limits, Auth failures, Guardrails tripped, Upstream service errors, Tool failures, Data/parsing errors, Agent output issues, Other) groups the top-level rollups and chart legends. Underneath, the open-ended failure modes are the specific patterns — the level you name, classify against, and manage.
The taxonomy
Every failure mode carries a name, a definition (“what distinguishes this failure mode?”), a curated severity (High / Medium / Low), an optional owner, compliance tags, a curation status, and a provenance badge:
| Provenance badge | Meaning |
|---|---|
| Platform default | Shipped with the Neens starter library (or bulk-imported from your own label list) |
| Custom | Added or confirmed by your team |
| Auto-discovered | Surfaced by failure clustering |
| Curation status | Meaning |
|---|---|
| Candidate | Suggested, not yet reviewed |
| Confirmed | Reviewed and adopted by your team |
| Monitoring | Kept under watch |
| Archived | Retired — excluded from classification and the Issues rollup |
Build your taxonomy
The Issues & Taxonomy tab header carries two buttons for managing the catalogue itself:
- + New issue type — the primary button. Opens a dialog to add a single failure mode by
hand: a name, definition, severity, optional owner, and compliance tags
(
POST /taxonomy/failure-modes). - Manage taxonomy — opens the full Taxonomy table: every mode with its provenance, curation status, and bulk actions (import, edit, merge, split, delete — see Curate it over time below). Use this when you’re tidying the catalogue rather than adding one issue.
Ways to populate the taxonomy:
- Import starter library — the fastest start. On the empty state (or via
POST /taxonomy/failure-modes/import), seed a set of common agent failure modes, or post your own label list in the same call. Imports are idempotent: re-importing refreshes definitions without duplicating, and never overwrites a status your team has since changed. - Confirm candidate / Promote to issue — turn a discovered cluster on the Failure Modes tab into (or link it to) a named issue. You must review at least 3 exemplar traces from the cluster first, so every mode is grounded in real evidence. See Promote a cluster to an issue for the cluster-side how-to.
- Create manually — the + New issue type button above, or
POST /taxonomy/failure-modesdirectly.
Import a taxonomy file (CSV / JSON)
Most teams already keep their failure taxonomy in a spreadsheet. Click Import on the Issues & Taxonomy tab to upload it as a CSV or JSON file — a thin path over the same idempotent bulk-import, so re-uploading updates modes in place instead of duplicating.
The CSV columns are:
| Column | Required | Notes |
|---|---|---|
name | Yes | The failure mode’s name |
definition | What distinguishes this failure mode | |
severity | One of high, medium, low (blank is allowed) | |
compliance_tags | One or more tags separated by ; or , | |
key | A stable id used for idempotent re-import (defaults to a slug of the name) |
Common spreadsheet header names are accepted as aliases — for example Failure Mode → name,
Description → definition, and Tags → compliance_tags. Grab the exact header row from
the Download template link in the Import modal (backed by
GET /taxonomy/failure-modes/import-template).
name,definition,severity,compliance_tags,key
Hallucinated Price,"Agent states a product price not present in the catalog tool result",high,accuracy;grounding,hallucinated_price
Wrong Order Lookup,"Agent looks up the wrong order id for a returns request",medium,reliability,wrong_order_lookupMalformed rows never sink the upload. A row missing name or carrying an invalid
severity is reported back individually — you get a created / updated / skipped summary
where each skipped row lists its row number and the reason — while every valid row still
imports. Nothing is silently dropped.
Imported modes land as Candidate (provenance: seeded) and are idempotent by key/name:
re-uploading the same file refreshes definitions and bumps each mode’s version rather than
creating duplicates, and it never overwrites a status your team has since changed.
Import from the command line or CI
For automation, post the file body directly to the raw-body endpoint (the request body is
the file — this is not a multipart form upload). Requires the manage_taxonomy permission —
any member or admin can import, edit, and delete taxonomy modes; viewers are read-only:
curl -X POST "$NEENS_BASE_URL/taxonomy/failure-modes/import-file?format=csv" \
-H "Authorization: Bearer $NEENS_API_KEY" \
-H "Content-Type: text/csv" \
--data-binary @failure-modes.csvformat is optional (csv or json) — Neens infers it from the Content-Type and content
when omitted — and status defaults to candidate. The response is
{created, updated, skipped: [{row, name, reason}], failureModes, format}.
The pre-existing strict-JSON endpoint POST /taxonomy/failure-modes/import (body
{"modes": [...]}, or {"include_starter_library": true}) remains the programmatic path when
you’re generating the list rather than uploading a file.
Curate it over time
Click any row in the Taxonomy table to open its detail drawer:
- Edit name, definition, severity, owner, and status (every change bumps the mode’s version).
- Merge two modes — the source’s cluster evidence and exemplars fold into the target and the source is archived.
- Split a mode into a new one, moving selected cluster links across.
- Delete a mistaken or obsolete mode outright (
DELETE /taxonomy/failure-modes/{id}) — distinct from merge (which archives the source) and archive (a status change); its cluster-evidence links are removed with it. - Cluster evidence — the clusters backing a mode. Links marked auto were proposed by Neens (after each clustering run, a cluster is linked to a mode when a clear majority of its classified members carry that mode); links you create yourself are never overwritten by the machine.
Every taxonomy change — create, import, edit, delete — is recorded in the tenant audit log
(GET /audit, admin-only) with who did it and when.
How traces get classified
Classification is done by the built-in Issue Classification judge, provisioned enabled for every agent:
- Every 30 minutes, Neens takes the agent’s recent failure-set sessions (the same selection criteria clustering uses — errors, below-threshold scores, human “fail” labels) that haven’t been classified yet. Healthy traffic is never classified, and each session is classified once.
- For each session, the classifier shows an LLM the session’s evidence (input, output,
messages, metadata) and the list of your non-archived failure modes with their definitions
(up to 40), and asks for the single best-matching mode — or
noneif no listed mode clearly applies. Answers are snapped to exact mode names; the model can never invent a mode. - Each classification is stored as a score (
metric_key=issue_class) carrying the mode name as its label, the classifier’s confidence, and a one-sentence reason. - Spend is bounded: at most 500 sessions per day per agent by default, drawn from the classifier’s own daily budget.
The classifier uses the agent’s LLM connection, like every LLM feature. With no connection configured, no classification happens — the taxonomy still exists and is fully editable, but Issues show zero classified traces. With no (non-archived) taxonomy, the classifier doesn’t run at all, so it never burns budget with nothing to classify against.
Issues
The Issues rollup shows each non-archived failure mode as a card with its lifecycle state, severity, classified trace count, mean classifier confidence, linked clusters, a trend sparkline, and its fix / eval-gate status. A mode with zero classified traces still appears — an Issue is a mode, whether or not it’s currently occurring. Scope the rollup with the time-range picker (Today, 24h, 7d, 30d, All time, or a custom range; default 7d), and filter by search, severity, and lifecycle state.
The analytics summary
A summary strip sits above the cards: total issues, active vs watching counts,
affected traces (sessions caught by an active issue in the window), and a severity /
lifecycle breakdown across the active ones. These numbers are computed on the server over your
entire taxonomy in the selected window, before the search/severity/state filters are applied
— so narrowing the list to severity=high never changes what the summary reports. Widening or
narrowing the time range does change it, since it re-windows which sessions count as “active”.
Active vs Watching
Every issue falls into exactly one of two buckets for the selected window:
- Active — classified traces > 0 in the window. Active issues render as full cards, sorted by classified trace count descending — the noisiest problem is always first, so triage order tracks actual impact rather than alphabetical order or creation date.
- Watching — zero classified traces in the window. These are modes in your taxonomy that simply haven’t happened (yet, or in this window). They render as a collapsed, compact one-line list — a severity dot, the name, its one-line definition, and “0 traces” — so a taxonomy of a hundred named modes doesn’t bury the handful that actually need attention. Expand the Watching section to see them all, or use Expand all / Collapse all at the top of the tab to open or close every group (active cards and the watching list) at once.
Example. Your taxonomy has 30 failure modes. This week, 6 of them classified at least one trace — one dominant mode with hundreds of sessions, the rest much smaller — and the other 24 classified none. The summary reads 6 active · 24 watching; the 6 active cards are sorted loudest-first, and the 24 watching modes sit collapsed underneath until you expand them or widen the window.
A Watching row’s action is deliberately lighter than an active card’s: a muted Definition link (so you can remind yourself what it means without opening the drawer) and an accent Create eval → link — pre-emptively locking the failure in as a regression test before it ever fires in production. View traces and Generate fix don’t apply with zero traces, so they’re not shown until the mode goes active.
Lifecycle
| State | Meaning |
|---|---|
| Open | Newly surfaced, not yet triaged |
| Acknowledged | Someone has seen it and taken ownership |
| Investigating | Being actively looked into |
| Mitigating | A fix is in progress |
| Resolved | Addressed — no longer expected to recur |
| Muted | Deliberately silenced |
Move an Issue between states with the Lifecycle actions menu on its card (or
PATCH /taxonomy/issues/{id}/lifecycle). Any state can transition to any other; the one
exception is a mode whose curation status is Archived — it’s out of the operational loop
entirely. Resolved and Muted Issues are hidden from the default list and suppressed
from dashboards and insight highlights; select those states in the filter to bring them back.
Neens reconciles the lifecycle for you. Every 30 minutes: an Issue whose linked remediation reaches verified (its proof eval passed) is auto-transitioned to Resolved (unless you muted it), and a Resolved Issue that receives new classified traces is reopened to Open — a fixed failure that recurs never stays silently “resolved”.
From an Issue to action
An active Issue card carries three actions:
- View traces — opens the traces the classifier assigned to the Issue in the selected window, each with its classification confidence and reason.
- Generate fix → View fix — a stateful action: Generate fix drafts a typed remediation for the Issue when it has none; once one exists, the same spot becomes View fix, linking to it with its current status. See Remediations.
- Create eval — locks the failure in as a regression test (below). Once the Issue already has a gate, the card shows the gate’s status here instead.
A watching Issue’s row keeps only Definition and Create eval — see Active vs Watching above.
Create an eval from an Issue
Click Create eval on the Issue card
Neens needs real evidence first: if the mode has no classified traces, linked clusters, or exemplars yet, the call is rejected with a clear message rather than creating an eval that would trivially pass forever.
Neens builds the pieces
One guided step (POST /flywheel/failure-modes/{id}/generate-eval) materializes a
dataset captured from the failing traces, drafts an LLM
judge that scores whether the failure is absent (high score = good), and
wires them together into an eval gate.
Enable the gate
The gate starts as a draft. Enable it on the Eval Gates page to start catching regressions — if the failure creeps back, the gate fails. An Issue that already has a gate shows it on the card instead of the Create eval action.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| ”No issue taxonomy yet” | The agent has no failure modes | Click Import starter library, Import a CSV/JSON file of your own taxonomy, add one manually, or confirm a cluster |
| Issues all show zero traces | No LLM connection (classifier can’t run), or no failing traffic in the selected window | Configure Settings → LLM connections; widen the time range |
| ”No classified issues in this window” | Taxonomy exists but nothing was classified into it recently | Normal when the agent is healthy; check View traces on the Failure Modes cards for unclassified failures |
| Everything is in Watching, nothing Active | No traces were classified into any mode in the current window | Widen the time range, or check the LLM connection / classifier as above |
| Summary tiles don’t match the cards you can see | The summary counts the whole taxonomy in the window, before your search/severity/state filters | Expected — clear the filters to see the same population the summary describes |
| Create eval fails with “no evidence sessions” | The mode has no classified traces, linked clusters, or exemplars | Wait for classification, link a cluster, or add exemplars first |
| A resolved Issue reopened by itself | New traces were classified into it after it was resolved | That’s the recurrence guard working — investigate the new traces |