The Neens loop

The Neens loop

Neens turns your agent’s failures into regression tests, so the ones you fix can’t come back. This page shows how that happens: six steps that start and end in production, and that every feature in these docs belongs to. Each turn of the loop takes one real failure, proves a fix for it, and leaves a regression test behind that every later release has to pass. That is the flywheel: the more it turns, the more of your agent’s past mistakes are locked out for good.

The Neens loopA clockwise loop of six steps: 1 production traces, 2 failure modes, 3 a frozen golden set, 4 a judge aligned to people, 5 a proven fix that a person merges, 6 a release gate that returns GO or NO-GO, and back to production. Each step names the Neens skill that drives it. In the centre: each turn adds a regression test. Two side loops, choose-model and reliability-review, sit outside the main loop.+1 regression testevery turn of the loop1Production tracesinstrument-agent2Failure modestriage-failures3Frozen golden setbuild-regression-set4Aligned judgewrite-judge5Proven fix · you mergefix-failure6Release gate: GO/NO-GOgate-releasechoose-modelreliability-review
The loop, clockwise from the top. Each box names its step and the skill that drives it. Side loops: choose-model (which model) and reliability-review (how are we doing).

Every step has a guide, and a Neens skill your coding agent can run for you. You can drive the whole loop from the Neens app, from your coding agent or harness over MCP, or mix the two.

People stay in charge of three things. You confirm which failures are real, you are the ground truth a judge is measured against, and a person merges every fix. Neens and your coding agent diagnose, implement and prove. Neither merges or deploys.

1. Production traces

What Neens doesReceives standard OpenTelemetry traces, rolls them into sessions, and scores a sample automatically with your judges
What you doPoint your agent’s OTLP exporter at Neens, once
GuideSend traces, Framework quickstarts, Continuous evaluation
Skillinstrument-agent — wires OpenTelemetry and proves a trace arrived with what Neens needs

Everything downstream is only as good as what arrives here. A trace that carries the input, output, model, tool calls and a conversation id is one Neens can cluster, judge and replay.

2. Failure modes

What Neens doesClusters failing traces into recurring failure modes, labels them, and says where each one lives: your agent, a downstream service, or a guardrail doing its job
What you doTriage: open the representative traces, confirm the modes that are real, dismiss the ones that aren’t, and record your verdicts
GuideFailure modes & clustering, Issues & failure modes, Review & annotations
Skilltriage-failures — reads the failure modes and real traces, and records your confirmed verdicts as ground truth

A cluster is a hypothesis. Your confirmation is what turns it into a failure mode worth a test.

3. Frozen golden set

What Neens doesSnapshots the failing traces, plus passing controls, into an immutable golden version
What you doDecide what belongs in the set, and fill in expected outputs where you have them
GuideDatasets → Golden datasets and versions
Skillbuild-regression-set — freezes a confirmed failure mode into a golden regression set

Version items never change, so every later run is measured against exactly the same examples. The passing controls matter as much as the failures: they are what catches a fix that breaks something that used to work.

4. A judge aligned to people

What Neens doesRuns a pass/fail judge for one failure mode and measures its agreement with your reviewers’ verdicts
What you doLabel a handful of traces, read where the judge disagrees with you, and revise it until it agrees
GuideJudges → Judge ↔ human alignment, Review & annotations
Skillwrite-judge — writes a binary judge for one confirmed failure mode and checks it against your own verdicts

A judge scores automatically only after it agrees with people. That is also what lets it verify a fix later: a judge nobody has checked cannot anchor a proof.

5. A proven fix

What Neens doesDrafts a typed remediation with a root cause; hands it to your coding agent as a fix bundle, or has the fix engine open an eval-verified PR; proves the candidate with pass^k against the failure’s golden set and your accumulated regression set
What you doReview the proof and merge. At every autonomy level, a person merges
GuideRemediations, Fix bundles, Eval-verified PR, Autonomy levels
Skillfix-failure — implements a remediation in your repo, proves it with a real Neens evaluation, and reports the PR and proof back

pass^k means the candidate must pass k independent runs, all green, with zero regressions. Two out of three is a fail. Unit tests passing is not proof.

6. Release gate

What Neens doesReplays every golden and regression set against the candidate version in a pre-prod evaluation, compares it with the current version, and returns GO or NO-GO with the regressing trajectories as evidence
What you doRun the gate in CI, or ask for it before you ship. Read the evidence behind a NO-GO
GuidePre-prod evaluations, Eval gates
Skillgate-release — returns GO, NO-GO or NO DATA with the evidence

There are two ways to run the gate. Your own eval harness can run the candidate over the golden prompts with neens eval run (by hand or in CI), and Neens scores the traces it emits. Or Neens calls your agent’s endpoint itself (push mode, neens eval push), with no harness on your side.

Back to production

A GO ships. After the merge, fix outcomes check that the failure’s volume in production actually dropped, and which business KPIs moved. New traces keep arriving, the next failure mode surfaces, and the loop turns again.

The difference is what the last turn left behind. The failure you just fixed is now a regression test in the set every future release replays, so it cannot silently come back.

Side loops

Two skills sit beside the loop rather than on it:

SkillWhenGuide
choose-modelWhich model should this agent run on? A model sweep replays one golden set against each candidate model and ranks them on pass^k, cost and latencyModel sweeps, Cost & Quality
reliability-reviewHow is the agent doing? A one-page report: KPIs against target, quality trends, top failure modes, the fix backlog, and which fixes heldBusiness KPIs, Fix outcomes

Start turning it

  • New to Neens? Getting started sends your first trace.
  • Want your coding agent to run the loop? Connect it and add the skills. If you’re not sure where to begin, ask it to run neens-start.
  • Want the vocabulary first? Read Core concepts.