Skills
The MCP tools are the raw capability. The Neens skills are the
judgment about how to use them: which call comes first, what counts as evidence, what a number
means, and when to stop and ask a person. They are plain SKILL.md files, open source at
github.com/neens-ai/skills, and work in any coding agent that
reads skills.
At a glance
| Source | github.com/neens-ai/skills |
| Skills | Nine — one per step of the loop, plus an entry point and a status report |
| Install | As a Claude Code plugin, or with npx skills add for any agent that reads SKILL.md |
| Needs | A Neens MCP connection, and an LLM connection in Settings → LLM providers |
The eight rules
Every skill follows the same eight rules. They are also a fair summary of how Neens expects the loop to be run.
- Evidence before opinion. Read real traces before naming a failure, writing a judge or proposing a fix. Every claim cites a trace id or a run id.
- A person owns the verdict. The agent proposes. Ground-truth labels, “this is a real failure” and “ship it” come from a human.
- One failure, one test. A judge checks one failure mode, returns pass or fail, and is checked against a person before it gates anything.
- Freeze before you measure. Runs replay immutable golden versions, so results stay comparable over time.
- Every fix ships with its proof. A fix is done when a real evaluation of the candidate build passes the failure’s regression set with zero regressions. Unit tests passing is not proof.
- Fix where the failure lives. A downstream outage goes to its owner, not into the prompt. A guardrail that fired correctly is not a bug.
- Honest numbers. Null means no data, never zero. Every rate carries its sample size. A void comparison is never ranked. Two out of three is not a pass.
- A person merges. Neens and your coding agent diagnose, implement and prove. A person decides what ships.
The nine skills
| Skill | Use it when | Loop step |
|---|---|---|
neens-start | You’re not sure where to begin. Checks the connection, reads the loop, routes you to the right skill | Entry point |
instrument-agent | No traces yet, or traces arrive looking empty. Wires OpenTelemetry to Neens and proves a trace arrived with what Neens needs | 1 · Instrument |
triage-failures | ”What’s failing?” Reads failure modes and real traces, sorts each failure by where it lives, records your verdicts | 2 · Triage |
build-regression-set | A failure is confirmed. Freezes its traces, plus passing controls, into an immutable golden set | 3 · Freeze |
write-judge | A failure needs an automated check. Writes a pass/fail judge and measures it against your verdicts | 4 · Judge |
fix-failure | Time to fix. Implements a remediation in your repo, proves it against the failure’s golden set, reports the PR and proof back | 5 · Fix |
gate-release | ”Can we ship this?” Replays every regression set against the candidate and returns GO, NO-GO or NO DATA with the evidence | 6 · Gate |
choose-model | ”Which model?” Runs a pass^k model sweep and reads it honestly | Model choice |
reliability-review | ”How are we doing?” A one-page status report: KPIs, trends, top failure modes, the fix backlog, which fixes held | Reporting |
The usual path for a new agent runs straight down the loop:
instrument-agent → triage-failures → build-regression-set → write-judge → fix-failure → gate-releaseInstall
Connect Neens
The skills drive the Neens MCP tools, so connect first. See Connect for every client.
Add the skills
Inside a Claude Code session:
/plugin marketplace add neens-ai/skills
/plugin install neens@neensThe skills then appear namespaced by the plugin, as neens:triage-failures, neens:fix-failure
and so on.
Start
Ask in plain words. Your coding agent picks the matching skill from its description.
Example prompts
| Prompt | Skill it triggers |
|---|---|
| ”Where do I start with Neens?” | neens-start |
| ”Send this agent’s traces to Neens and check they arrive.” | instrument-agent |
| ”What’s failing in my agent this week?” | triage-failures |
| ”Freeze the refund-policy failures into a regression set.” | build-regression-set |
| ”Write a judge for answers that contradict the tool result.” | write-judge |
| ”Fix the top open Neens remediation and prove it against my preview deploy.” | fix-failure |
| ”Can we ship branch fix/refund-policy?” | gate-release |
| ”Is the smaller model good enough for this agent?” | choose-model |
| ”Give me this week’s reliability review.” | reliability-review |
Requirements
- A Neens MCP connection scoped to the agent you’re working on. Traces are needed for every
skill except
instrument-agent. - An LLM connection in Settings → LLM providers. Judges and remediations run on it; see LLM connections.
- A way to run the candidate for
fix-failure,gate-releaseandchoose-model: your agent reachable over HTTP from Neens (a preview deploy, for example), or a harness of your own that replays golden inputs and sends the traces. Each skill explains both options, and Pre-prod evaluations covers them in full.
The skills never merge and never deploy. fix-failure opens a pull request with its proof
attached, and gate-release returns a verdict with its evidence. The decision stays with you.
Write your own
The nine skills cover the loop every agent team runs. Your agent has its own failure modes, tools and release process, so write skills for those too:
- Start from one of the nine as a template. A skill is a folder with a
SKILL.md: front matter with anameand adescriptionthat says when to use it (and when not to), then the procedure. - Follow the eight rules above, especially evidence before opinion and a person owns the verdict.
- Name only real tools and arguments. The MCP tools reference lists every one; the skills repo includes a checker that validates tool names and arguments against the live tool surface.
Related
- Coding agents overview — how tools, skills and docs fit together.
- Connect — add the Neens MCP server to your client.
- MCP tools — the full tool reference.