The Neens loop
Neens turns your agent’s failures into regression tests, so the ones you fix can’t come back. This page shows how that happens: six steps that start and end in production, and that every feature in these docs belongs to. Each turn of the loop takes one real failure, proves a fix for it, and leaves a regression test behind that every later release has to pass. That is the flywheel: the more it turns, the more of your agent’s past mistakes are locked out for good.
choose-model (which model) and reliability-review (how are we doing).Every step has a guide, and a Neens skill your coding agent can run for you. You can drive the whole loop from the Neens app, from your coding agent or harness over MCP, or mix the two.
People stay in charge of three things. You confirm which failures are real, you are the ground truth a judge is measured against, and a person merges every fix. Neens and your coding agent diagnose, implement and prove. Neither merges or deploys.
1. Production traces
| What Neens does | Receives standard OpenTelemetry traces, rolls them into sessions, and scores a sample automatically with your judges |
| What you do | Point your agent’s OTLP exporter at Neens, once |
| Guide | Send traces, Framework quickstarts, Continuous evaluation |
| Skill | instrument-agent — wires OpenTelemetry and proves a trace arrived with what Neens needs |
Everything downstream is only as good as what arrives here. A trace that carries the input, output, model, tool calls and a conversation id is one Neens can cluster, judge and replay.
2. Failure modes
| What Neens does | Clusters failing traces into recurring failure modes, labels them, and says where each one lives: your agent, a downstream service, or a guardrail doing its job |
| What you do | Triage: open the representative traces, confirm the modes that are real, dismiss the ones that aren’t, and record your verdicts |
| Guide | Failure modes & clustering, Issues & failure modes, Review & annotations |
| Skill | triage-failures — reads the failure modes and real traces, and records your confirmed verdicts as ground truth |
A cluster is a hypothesis. Your confirmation is what turns it into a failure mode worth a test.
3. Frozen golden set
| What Neens does | Snapshots the failing traces, plus passing controls, into an immutable golden version |
| What you do | Decide what belongs in the set, and fill in expected outputs where you have them |
| Guide | Datasets → Golden datasets and versions |
| Skill | build-regression-set — freezes a confirmed failure mode into a golden regression set |
Version items never change, so every later run is measured against exactly the same examples. The passing controls matter as much as the failures: they are what catches a fix that breaks something that used to work.
4. A judge aligned to people
| What Neens does | Runs a pass/fail judge for one failure mode and measures its agreement with your reviewers’ verdicts |
| What you do | Label a handful of traces, read where the judge disagrees with you, and revise it until it agrees |
| Guide | Judges → Judge ↔ human alignment, Review & annotations |
| Skill | write-judge — writes a binary judge for one confirmed failure mode and checks it against your own verdicts |
A judge scores automatically only after it agrees with people. That is also what lets it verify a fix later: a judge nobody has checked cannot anchor a proof.
5. A proven fix
| What Neens does | Drafts a typed remediation with a root cause; hands it to your coding agent as a fix bundle, or has the fix engine open an eval-verified PR; proves the candidate with pass^k against the failure’s golden set and your accumulated regression set |
| What you do | Review the proof and merge. At every autonomy level, a person merges |
| Guide | Remediations, Fix bundles, Eval-verified PR, Autonomy levels |
| Skill | fix-failure — implements a remediation in your repo, proves it with a real Neens evaluation, and reports the PR and proof back |
pass^k means the candidate must pass k independent runs, all green, with zero regressions. Two out of three is a fail. Unit tests passing is not proof.
6. Release gate
| What Neens does | Replays every golden and regression set against the candidate version in a pre-prod evaluation, compares it with the current version, and returns GO or NO-GO with the regressing trajectories as evidence |
| What you do | Run the gate in CI, or ask for it before you ship. Read the evidence behind a NO-GO |
| Guide | Pre-prod evaluations, Eval gates |
| Skill | gate-release — returns GO, NO-GO or NO DATA with the evidence |
There are two ways to run the gate. Your own eval harness
can run the candidate over the golden prompts with neens eval run (by hand or in CI), and Neens
scores the traces it emits. Or Neens calls your agent’s endpoint itself (push mode, neens eval push), with no harness on your side.
Back to production
A GO ships. After the merge, fix outcomes check that the failure’s volume in production actually dropped, and which business KPIs moved. New traces keep arriving, the next failure mode surfaces, and the loop turns again.
The difference is what the last turn left behind. The failure you just fixed is now a regression test in the set every future release replays, so it cannot silently come back.
Side loops
Two skills sit beside the loop rather than on it:
| Skill | When | Guide |
|---|---|---|
choose-model | Which model should this agent run on? A model sweep replays one golden set against each candidate model and ranks them on pass^k, cost and latency | Model sweeps, Cost & Quality |
reliability-review | How is the agent doing? A one-page report: KPIs against target, quality trends, top failure modes, the fix backlog, and which fixes held | Business KPIs, Fix outcomes |
Start turning it
- New to Neens? Getting started sends your first trace.
- Want your coding agent to run the loop? Connect it and add the
skills. If you’re not sure where to begin, ask it to run
neens-start. - Want the vocabulary first? Read Core concepts.