Coding agents
Run the failure→fix loop from the coding agent you already work in — Claude Code, Codex, Cursor, VS Code or any other harness that speaks MCP. Your agent reads the traces, freezes the failures into regression sets, writes the judges, implements the fix in your repo and proves it against a real evaluation. You review the evidence and merge.
Three layers
| Layer | What it gives your agent | Where |
|---|---|---|
| The MCP server — the tools | 60 tools, 38 read and 22 write, over the same permission and tenancy checks as the app. Four families (the fix loop, orchestrated fix runs, prompt optimization and model sweeps) appear only when those features are enabled for your workspace. No tool deletes anything | Connect · MCP tools |
| The Neens skills — the judgment | Which call comes first, what counts as evidence, what a number means, and when to stop and ask a person | Skills |
| llms.txt — the docs as context | These docs in the llms.txt format: a map at /docs/llms.txt and the full text at /docs/llms-full.txt, for an agent that needs to look something up | llms.txt · llms-full.txt |
The tools alone let an agent do anything the loop allows. The skills are what make it do the right thing in the right order: read evidence before naming a failure, check a judge against your verdicts before trusting it, and never call a fix done without a passing evaluation.
The loop, step by step
| Step | What happens | Skill | Guide |
|---|---|---|---|
| 1 · Instrument | Your agent sends OpenTelemetry traces to Neens | instrument-agent | Send traces |
| 2 · Triage | Recurring failures are clustered into failure modes; you confirm which are real | triage-failures | Issues & failure modes |
| 3 · Freeze | Failing traces plus passing controls become an immutable golden set | build-regression-set | Datasets |
| 4 · Judge | A pass/fail judge per failure mode, checked against your reviewers | write-judge | Judges |
| 5 · Fix | A remediation is implemented in your repo and proven with pass^k before merge | fix-failure | Remediations |
| 6 · Gate | Every future release replays every regression set: GO or NO-GO | gate-release | Pre-prod evaluations |
Two more skills sit beside the loop: choose-model runs a model sweep
when you want to change models, and reliability-review writes the weekly status report.
neens-start is the entry point when you don’t know where to begin.
What it looks like
Asking a coding agent with the Neens skills installed whether a branch is safe to ship:
› Can we ship branch fix/refund-policy?
gate-release replaying 3 frozen golden sets against the candidate
candidate fix/refund-policy@4e1c9a baseline prod:7d k=3
refund-policy-v4 14/24 → 22/24 +8 0 regressions
tool-contradiction 24/24 → 24/24 0 regressions
escalation-controls 10/10 → 9/10 −1 1 regression
└ trace 7f3a… agent escalated a refund under $50
VERDICT NO-GO policy: ≥ 90% pass^3, 0 regressions
The fix works on its own failure but broke one control.
Want me to look at trace 7f3a before you decide?The numbers are illustrative. The shape is real: one row per regression set, the regression named with its trace, a verdict against your gate policy, and the call left to you.
The ground rule
Your agent diagnoses, implements and proves. A person merges. Neens never merges or deploys, and neither do the skills. A fix arrives as a pull request with its evaluation attached; what ships is your decision. Autonomy levels control how much of the rest runs unattended, and the merge stays a human step at every level.
Everything your agent does happens as you. It signs in through your browser, is limited to the one agent you pick and to your role, and you can revoke it at any time under Settings → MCP access.