NeensDocs
⬡ Observe · Diagnose · Evaluate · Fix

Make every agent failure its last.

Neens turns your agent’s traces into answers: see every run, find why it failed, score it with automated judges, and close the loop with fixes that can’t regress.
Get started →Send your first trace
Observe → Diagnose → Evaluate → Fix
👁️
Observe
Inspect every trace and session — spans, tool calls, inputs, outputs, latency, and cost.
🩺
Diagnose
Recurring failures are clustered into failure modes and tracked as Issues you can manage.
⚖️
Evaluate
Score traces with automated judges, curate golden datasets, and align judges to human labels.
🔧
Fix
Turn failures into typed remediations, simulate a fix before shipping, and gate releases on your own evals.
Start here
🚀
Getting started
Sign in, create an agent, and send your first trace in a few minutes.
📖
Core concepts
Traces, sessions, judges, scores, failure modes — the vocabulary everything builds on.
🔌
Send traces
The ingestion API, with copy-paste examples for OpenTelemetry, Python, and curl.
Go deeper
🧪
Pre-prod evaluations
Replay a golden dataset against a candidate agent version and catch regressions before they ship.
📊
Dashboards & insights
Slice quality, cost, and volume with the metrics catalogue, and let detectors surface anomalies.
🏢
Administration
Orgs, agents, members and roles, API keys, LLM connections, and the audit log.

Neens — make every agent failure its lastDocumentation build v0.16.6