Traces & Sessions
Traces and Sessions are where you see what your agents actually did. A trace is one agent run — a request that came in, the steps the agent took, and the response it produced. A session is a conversation: one or more traces grouped together. Start here to inspect a single run, follow a multi-turn interaction, or hunt down the runs behind a failure.
If you haven’t sent any data yet, head to Send traces first — the Traces page will walk you through it.
At a glance
| Where | Sidebar → Observe → Traces / Sessions |
| Key API routes | GET /sessions (list traces), GET /sessions/{id} (trace detail), GET /conversations (session rollup), GET /conversations/{id}, GET /agents |
| Scope | Always scoped to the current agent; you can only open traces in agents you belong to |
| Needs | Ingested trace data — nothing else. No LLM connection required to browse |
The trend header
Above the filters on both the Traces and Sessions pages sits a trend header — a rich, at-a-glance read on how your volume, token usage, and cost are moving over time, before you touch a single row. Use it to spot a spike or a creeping trend first, then dig into the table to find the runs behind it.
The header has three parts, plus one control that ties everything below it together.
The KPI strip
Four tiles run across the top. Each shows the value for the selected time window, a period-over-period delta, and a sparkline of the trend:
| Tile | What it shows |
|---|---|
| Traces/day (Sessions page: Sessions/day) | The average number of runs ingested per day in the window. |
| Total tokens | Input + output tokens across every run in the window. |
| p50 latency | The median end-to-end run duration — half your runs finished faster, half slower. |
| Est. spend | Estimated cost across the window, derived from tokens × each run’s model price (see Token usage & cost). |
Reading the delta. Each tile compares the selected window to the immediately-preceding window of equal length — so on 7d, “this week vs. last week”; on 24h, “today vs. yesterday”. The delta is the change between the two. A rising Traces/day or Total tokens means more activity; a rising p50 latency or Est. spend is usually worth a look. The direction is what matters — read it as “up since last period” or “down since last period”, not as an absolute score.
A blank metric means “no data”, never zero. If a window has nothing to compute a tile from —
for example Est. spend when every run used an unpriced model, or a delta with no preceding
period to compare against — the tile shows —. Neens never shows a fake 0 in place of a
number it doesn’t have.
The token trend chart
Below the tiles is a stacked-area chart of token throughput over the selected window, split into two bands:
- Input tokens — tokens sent to the models (prompts, context, tool results).
- Output tokens — tokens the models generated back.
The two bands stack, so the top edge of the chart is your total token volume over time — the same figure the Total tokens tile reports for the window — and the split shows the input/output mix behind it. That mix is worth watching because output tokens usually cost more than input tokens, so a run whose output share is climbing gets more expensive faster than its raw volume suggests. Hover any point to read the exact Input and Output counts at that moment.
Segment the trend by model, agent, status or source
Use the Group by control in the header to break the volume into stacked series instead of the input/output token split. Pick a dimension:
- Model — one band per model, so you can see a migration land or a model’s share grow.
- Agent — one band per agent, to compare how much traffic each is handling.
- Status —
okvs.errorvs.unknownvolume side by side. - Source — one band per ingestion source.
Each band is a count over time, and the bands stack to the same total volume as the ungrouped view — so the shape of the whole is unchanged, only its composition is revealed. When a dimension has many values, the busiest few are shown individually and the rest fold into a single Other band so the chart stays readable; the legend names every band. Set Group by back to None to return to the token chart. Your choice is remembered per page.
Overlay the eval pass-rate (semantic health)
When judges or scorers have graded runs in the window, the header offers a Pass rate view — the share of scored runs that pass their threshold, which is the meaningful “is quality holding up over time?” signal (and not the same as operational status: a run can complete cleanly yet still be a wrong answer).
- A Pass rate KPI tile appears in the strip, with its period-over-period delta.
- Click Pass-rate trend to reveal a small pass-rate % line beneath the main chart. It sits on its own chart with its own 0–100% axis — a percentage and a token count can’t share a scale, so it is never overlaid as a second axis. A period with no graded scores shows a gap in the line rather than a misleading 0%.
If nothing has been scored in the window, the tile, the toggle and the line are simply not shown — there’s no empty health claim to read into.
The time-window control — one window, one truth
A single time-window control in the header scopes the chart and the list table below it together. Change it in one place and everything on the page — the sparklines, the KPI deltas, the token chart, and the rows in the table — moves to the same window. There is no way for the chart and the table to disagree about which slice of time you’re looking at.
The presets are Today, 24h, 7d, 30d, and All, plus Custom for an explicit from/to range. (See Time ranges for exactly what each preset covers.)
Drag across the chart to zoom
You don’t have to reach for the picker to narrow the window. Drag horizontally across the token chart to select a span of buckets, and on release the page zooms to exactly that range — the chart, the KPI tiles, and the list table below all snap to it together, because the selection sets the page’s one time window (as a Custom range). The selection includes the full width of the last bucket you drag over, so dragging across two days gives you both days end-to-end.
A plain click (without dragging) does nothing — it won’t collapse the window to a single instant. To zoom back out, pick any preset (7d, 30d, All, …); selecting a preset replaces the custom range the drag created.
Collapse it when you need the room
The whole header is collapsible. Click the collapse control to fold it down to a single-line trend — the trace count over the window (on Sessions, the session count) — reclaiming the vertical space for the table when you’re deep in the rows. Expand it again whenever you want the full picture back. Neens remembers your choice, so the header stays the way you left it the next time you open the page.
The same header on other list pages
The trend header isn’t only on Traces and Sessions. Several other list pages carry a count variant of it — the same time-window control, sparklines, collapse toggle and ”— means no data” rule, but built around how many of something happened over time rather than tokens and cost. Wherever you land on one of these pages, you get the “what’s the trend?” read before you scan the rows:
| Page | What the header trends |
|---|---|
| Activity | Event volume — activity runs per day (or per hour on a short window). |
| Audit log | Auth and admin events per day, over the events you’re allowed to see. |
| Remediations | Fix-task throughput — how many fixes were created per day. |
| Pre-prod evals | Run cadence plus a pass-rate tile, so you can see both how often you’re running evals and whether they’re passing. |
| Review queue | Annotation throughput — how many human labels the review loop is producing per day. |
Each header has a /day tile and a total tile for the window; Pre-prod evals adds the
pass-rate tile. As on Traces and Sessions, the chart is a plain volume-over-time area — it is
not a success/error split, because a run’s pipeline status isn’t the same thing as a semantic
failure (that lives in Scores, Failure Modes and the clusters). A window with nothing to show reads
as — or an empty chart, never a fabricated 0.
On Pre-prod evals, the pass-rate tile reads — for a window in which no items were scored yet — a run that’s still awaiting traces contributes to the cadence count but not to the pass-rate until its items are judged.
Worked example: chasing rising spend
Say the token chart shows the Output tokens band climbing steadily across the week while Input tokens stays flat. Here’s how you’d run it down without leaving the page:
- Set the time-window control to 7d so the whole climb is in view, and cross-check the Est. spend tile — a rising output share usually shows up there as an upward delta, since output tokens cost more.
- Hover the tail of the chart to read the exact Input / Output counts and confirm the mix really is shifting toward output.
- Because the window scopes the table too, the rows below already cover the same period. Sort by Tokens out (add the column from Columns if it isn’t shown) to surface the runs generating the most output.
- Open one of the heaviest runs to inspect its spans and see where the output is coming from — a verbose model call, a retry loop, or a tool returning large payloads back into the prompt.
The header turned “spend is creeping up” into a ranked list of the exact runs driving it.
Traces
The Traces page lists one row per trace, newest first, 50 rows per page.
Finding traces
- Search the loaded page by trace ID with the Search loaded traces by ID… box (matches within the rows already loaded).
- Filter with the Filters button. It opens a panel of grouped filters; edits apply when you click Apply. Active filters show as removable chips next to the button, with a count badge and a Clear all link.
- Sort by clicking a sortable column header (sorts the loaded page).
- Columns are configurable via the Columns button, and your choice is remembered. Beyond the fixed fields, Neens automatically offers a column for every top-level key in your traces’ metadata (under Metadata) and for every enrichment output field (under Enrichment) — see Enrichments.
Full-text search
The search box on the Traces and Sessions pages doesn’t just match trace IDs — type words
and Neens searches the content of your traces: the message text and tool-call content inside
the spans. Use it to find the runs where an agent said (or a tool returned) a particular phrase —
"refund policy", an error string, a customer name — without knowing the trace ID.
Content search combines with every other filter: search timeout with the Status filter set
to error and the time window on Last 24h to see only recent errored runs that mention a timeout.
How it works. Your query is split into terms that are ANDed together — a trace matches only if all its terms appear (up to 12 terms; extra words are dropped). On the ClickHouse backend the search is accelerated by a token bloom-filter index over the span and tool-call bodies; it also works on the SQLite/Postgres backends. Each term resolves up to 10,000 matching traces, so a search on a very broad word may return a truncated set.
On the API, pass the query as the q parameter to GET /sessions or GET /conversations,
alongside any of the other filter parameters.
Every available filter
| Filter | What it matches |
|---|---|
| Status | Trace status: ok, error, or unknown |
| Time | When the trace started: Today, Last 24h, 7 days, 30 days, All time, or Custom (see Time ranges) |
| Agent | The agent name on the trace |
| Agent Version | The version label the run was tagged with (from the neens.version_label attribute) |
| Source | Ingest format: otlp, openinference, or raw |
| Conversation ID | Traces belonging to a conversation |
| Session ID | One specific trace by its ID |
| Failure cluster | Members of an active failure cluster (see Failure clustering) |
| Score | Traces carrying a score for a metric, optionally Pass/Fail against its threshold (see Scores) |
| Issues & Taxonomy | Traces classified into an issue (see Taxonomy & Issues) — distinct from Failure cluster above |
| Turns, Tokens in, Tokens out, Total tokens, Duration, Spans | Min–max numeric ranges |
| Model, Tool, Span kind, Span status | Span-level multi-selects — a trace matches if any of its spans matches (OR within a filter, AND across filters) |
| Tags | Labels you’ve applied to traces |
The same vocabulary is available on the API as GET /sessions query parameters
(status, agent_name, source, model, tool_name, span_kind, span_status, tags,
cluster_id, conversation_id, session_id, score_metric, score_status, score_label,
issue_mode, version, started_after/started_before (ISO-8601), min_*/max_* ranges,
plus page and page_size, default 50, max 200). You can also filter by an enrichment value
with enrichment=<enrichment_id>|<field>|<value> (repeatable; values AND together).
The values inside each dropdown (agents, versions, models, tools, tags, score metrics, issues, …) are listed alphabetically, so a value stays in the same place as more traces arrive.
What the columns show
The default columns give an at-a-glance read on each run:
| Column | What it tells you |
|---|---|
| ID | The trace identifier (always shown; click the row to inspect). |
| Status | Whether the run ended ok, error, or unknown. |
| Model | The primary model used. |
| Agent Version | The version label the run was tagged with, if any (blue pill). |
| Issue | The issue this trace was classified into, if any (amber pill). |
| Failure Cluster | The active failure cluster this trace belongs to, if any (purple pill) — the ML-grouped cluster, distinct from the classifier’s Issue. |
| Turns | Number of conversational turns. |
| Tokens in / Tokens out | Input and output token counts. |
| Duration | End-to-end latency. |
| Timestamp | When the run started. |
Optional columns include Tags, Agent, Source, Total tokens, Ended, Cost (USD, derived from tokens × the model’s price — see Token usage & cost), Conversation, and any discovered metadata or enrichment columns.
Issue vs. Failure Cluster. The Issue column (and the Issues & Taxonomy filter) surface the classifier’s closed-set label, while the Failure Cluster column groups similar failures discovered by ML clustering. They’re separate signals — a run can carry one, both, or neither. Both lead into Failure clustering.
Inspecting a trace
Click any row to open the inline inspector beside the list (click again to close). The header shows the trace ID, agent, duration, and total tokens; applied tags render as pills underneath. Traces can also open on their own page, where the header adds the start time, span count, and — when the trace has been classified — a Classified as band naming the failure mode, its confidence, its lifecycle state, and the classifier’s reasoning.
From either view you can act on the trace with the header buttons:
- Add to dataset — put this trace in a dataset for evaluation.
- Annotate — attach a human judgment (see Annotations & review).
- Assign label — tag the trace for later filtering (the Tags filter/column).
System prompts in effect
When a trace carries system prompts, a collapsible System prompts in effect panel appears above the span explorer. It shows the actual instructions the agent ran under for this trace, recovered from its LLM spans and attributed per model call. Identical prompts shared across spans (e.g. a supervisor and its workers) collapse into one row; genuinely distinct prompts each appear. The panel is hidden when the trace has none.
The span explorer
The heart of the inspector is a five-tab view of everything that happened in the run. A fullscreen toggle in the tab bar gives you more room (Esc exits).
- Spans — a waterfall timeline plus a clickable span list on the left; the right pane shows the full trace by default and narrows to a single span when you click one (← Back to full trace restores it). Each span’s detail includes its Input, Output, error message (if any), per-span token counts, and its Tool Calls — each tool invocation’s name, arguments, result, error, and duration. The split between panes is draggable and remembered.
- Conversation — the human/assistant transcript, cleaned up so you read the dialogue rather than internal plumbing. It is a reconstruction from the spans, and the same derivation is what a dataset capture and a pre-prod baseline record — see Conversation transcript for exactly which spans and payload shapes contribute a turn.
- Graph — a turn-grouped diagram of the run (below).
- Agent Map — the run’s agent/tool topology at a glance, with the cost/latency bottleneck highlighted (below).
- Raw — the underlying trace JSON, unmodified.
The Graph tab
Long agent runs produce hundreds of spans; laid out flat they’re unreadable. The Graph tab groups spans into turns — a top-level LLM exchange plus everything it triggered — and renders one compact, collapsible card per turn showing its number, status, a message snippet, span count, tokens, and duration. Expanding a turn draws its internal span graph to the right; turns containing errors start expanded. Nothing is lost: every span is reachable, and clicking a span node opens its detail (including tool calls).
The toolbar offers a Search spans / turns… box, an Errors only toggle, and Expand all / Collapse all; a minimap and zoom controls help navigate large runs. In a multi-trace conversation, turns are grouped under a header per member trace.
The Agent Map tab
Where the Graph tab shows every span, the Agent Map answers a different question: who calls whom, and which node dominates cost and latency? It rolls the run up into its distinct actors — each agent, each tool, and each LLM model becomes a single node — and draws the weighted hand-offs between them, laid out left→right (a supervisor on the left flowing out to the workers, tools, and models it drives).
- Edge thickness encodes how heavy each hand-off is. Toggle whether that’s driven by Calls, Latency, or Tokens from the toolbar.
- The bottleneck is highlighted. The single node responsible for the most time gets a red ring and a ⚡ badge, and it’s named in the toolbar. Neens attributes self time — a node’s own time minus the time spent inside the children it called — so the honest culprit (usually a model or a slow tool) is flagged, not the outer agent that merely wraps everything.
- Nodes carry their rolled-up invocation count, self time and share, tokens, cost, and error count; click any node for the full breakdown. (A node whose model has no price shows — for cost — see Cost & model pricing.)
This is the fastest way to see the shape of a multi-agent run and spot the node worth optimizing first.
Agent Map across a cohort
A single run is an anecdote. To see whether a bottleneck is systemic — which node dominates cost and latency across the last 7 days, a failure cluster, or an agent version — open the Agent Map page (under Diagnose in the sidebar). It rolls a whole cohort of runs into one topology, using the same filters as the Traces and Sessions lists (time window, cluster, agent version, failure set, model, tool, and the rest). From the failure Clusters view you can also click “View agent map” to jump straight to the aggregate topology for that cluster.
- Per-run averages by default. Node and edge numbers read as a typical run (total ÷ number of runs), so cohorts of different sizes stay comparable. Toggle to Totals to see where the whole cohort spent its time and money.
- Prevalence. Click a node to see how many of the cohort’s runs it appears in (e.g. “appears in 48 of 500 runs”) — a tool that shows up in every run reads very differently from one that fires rarely.
- Error rate, not raw counts. Each node and edge carries the share of its calls that errored, so a handful of errors over thousands of calls doesn’t masquerade as a hot spot.
- Versions split by model. Because LLM nodes are keyed by model, a model change across the cohort shows up as two distinct nodes — the honest way to compare, say, opus vs. sonnet.
Very large cohorts are capped (the newest runs are rolled up first); when that happens Neens tells you how many runs were left out so you can narrow the window.
Token usage & cost
Every trace records input and output token counts (per span and rolled up per trace). The Cost column is derived from those counts at read time — tokens × the price of the model that trace actually ran on, from a dated price table you can override with your own negotiated or self-hosted rates. There is no stored cost column, so correcting a price corrects the number everywhere.
If a trace’s model has no price, the Cost cell shows —, never $0.00: Neens reports an
unpriced model honestly rather than applying a generic rate. Set a price under
Settings → Model pricing. See Cost & model pricing for the
full model, and Dashboards for spend breakdowns by model.
Sessions
The Sessions page rolls traces up into conversations — one row per conversation. Every
trace that shares a conversation_id is grouped together; a trace without one is its own
single-trace session. This is the right view for multi-turn agents, where one user interaction
spans several back-and-forth runs.
Neens captures the conversation ID at ingest from standard attributes
(gen_ai.conversation.id, session.id, conversation.id, or thread.id — see
Send traces).
Traces vs. Sessions. A trace is one run. A session is a conversation — a group of related traces. If your agent answers in a single run, one trace is one session; if it takes several turns, those traces roll up into one session row.
Columns and filtering
Sessions share the Traces page’s Filters and Columns controls, and a filter matches a conversation when any of its member traces matches — but the row’s totals always reflect the full conversation. The one exception is the numeric-range filters (Turns, Tokens in/out, Total tokens, Duration): on Sessions these test the conversation’s aggregate total, so they line up with the summed value the row shows rather than any single member trace. Columns specific to the rollup:
- Conversation ID — the grouping key (always shown).
- Session IDs — the member trace IDs: the full ID for a single-trace session, else the first
ID plus a
+Ncount with the complete list on hover. - Traces — how many runs make up the session.
Status is the worst member status (error wins), Model, Agent, and Agent
Version are the most common values across members, Turns and token counts are summed,
Duration spans the first member’s start to the last member’s end, and Failure mode
shows the most recent member classification.
Inspecting a session
Click a row to open the conversation drawer. It uses the same five-tab span explorer as a single trace, but stitched across every member trace:
- Spans lays out one continuous waterfall on the conversation’s wall-clock timeline, with member traces as labelled groups in the span list.
- Conversation, Graph, and Agent Map present the whole multi-turn interaction as a single flow.
The drawer header lists each member trace as a chip — click one to drill into that trace on its own. It also carries the same actions as a single trace, applied to the whole conversation:
- Add to dataset — add every member trace to a dataset for evaluation.
- Annotate — attach a human judgment to the conversation.
- Assign tag — tag the conversation (the Tags filter/column).
Because a session is a rollup, each of these fans out across all member traces; the dialog confirms that it “applies to all N traces in this conversation” before you commit.
Time ranges
The Time filter on Traces and Sessions offers the standard presets used across Neens:
| Preset | Window |
|---|---|
| Today | Since the start of the current calendar day |
| Last 24h | Trailing 24 hours |
| 7 days / 30 days | Trailing 7 / 30 days |
| All time | No time bound |
| Custom | Explicit from/to date-times (either side may be left open) |
The trend header’s time-window control draws on this same set of ranges, and it is the page’s window: changing it moves the header’s chart and the list table together, so you never have to set the time in two places.
Today is a calendar-day window — subtly different from the trailing Last 24h. In the
Traces/Sessions filter it starts at midnight in your browser’s timezone; on server-windowed
pages that accept a range parameter (dashboards, overview, activity metrics), Today starts
at 00:00 UTC. Pick whichever matches how you’re reasoning about the data.
Reference
Key API endpoints
| Endpoint | What it returns |
|---|---|
GET /sessions | Paged trace list (page, page_size default 50 / max 200) with per-row labels, failure mode, and enrichment values |
GET /sessions/filter-options | The distinct values behind the filter dropdowns (agents, versions, models, tools, span kinds, tags, clusters, score metrics) |
GET /sessions/{id} | One trace: header, spans, tool calls, system prompts, and its issue classification |
GET /sessions/{id}/similar | Up to limit (default 10) semantically similar traces, by embedding distance |
GET /conversations | Paged conversation rollup, same filter vocabulary as GET /sessions |
GET /conversations/{id} | One conversation: rollup, member traces, and stitched spans/tool calls |
POST /conversations/{id}/labels | Tag a conversation — fans the label out to every member trace |
POST /conversations/{id}/annotations | Annotate a conversation — fans the judgment out to every member trace |
GET /agents | The distinct agent names seen in the agent |
All timestamps are returned as UTC ISO-8601.
Troubleshooting
- The page shows a setup guide instead of a table — no traces have arrived yet. The zero state includes copy-paste snippets pre-filled with your agent’s ingest key and auto-refreshes every few seconds; see Send traces.
- A session shows one trace per row instead of grouping — the member traces aren’t sharing a conversation ID. Ensure your instrumentation sets one of the conversation attributes listed above on every run in the interaction.
- Model column is empty on old traces — the per-trace model rollup is stamped at ingest; traces ingested before it existed derive a model in the detail view but may show — in the list. Re-ingesting fills it in.
- The Conversation tab is empty, or shows a reply the Raw tab contradicts — the transcript is derived from the spans with a documented set of rules (which span kinds contribute, which payload shapes are understood). See Conversation transcript → Troubleshooting for the cause and the instrumentation change that fixes it.
- A content search looks like it’s missing results — a search on a very common word may show a truncated set (each term caps at 10,000 matches). Narrow it with more terms, or add filters (status, time, agent) to bring the result set under the cap.