Skills

The MCP tools are the raw capability. The Neens skills are the judgment about how to use them: which call comes first, what counts as evidence, what a number means, and when to stop and ask a person. They are plain SKILL.md files, open source at github.com/neens-ai/skills, and work in any coding agent that reads skills.

At a glance

Sourcegithub.com/neens-ai/skills
SkillsNine — one per step of the loop, plus an entry point and a status report
InstallAs a Claude Code plugin, or with npx skills add for any agent that reads SKILL.md
NeedsA Neens MCP connection, and an LLM connection in Settings → LLM providers

The eight rules

Every skill follows the same eight rules. They are also a fair summary of how Neens expects the loop to be run.

  1. Evidence before opinion. Read real traces before naming a failure, writing a judge or proposing a fix. Every claim cites a trace id or a run id.
  2. A person owns the verdict. The agent proposes. Ground-truth labels, “this is a real failure” and “ship it” come from a human.
  3. One failure, one test. A judge checks one failure mode, returns pass or fail, and is checked against a person before it gates anything.
  4. Freeze before you measure. Runs replay immutable golden versions, so results stay comparable over time.
  5. Every fix ships with its proof. A fix is done when a real evaluation of the candidate build passes the failure’s regression set with zero regressions. Unit tests passing is not proof.
  6. Fix where the failure lives. A downstream outage goes to its owner, not into the prompt. A guardrail that fired correctly is not a bug.
  7. Honest numbers. Null means no data, never zero. Every rate carries its sample size. A void comparison is never ranked. Two out of three is not a pass.
  8. A person merges. Neens and your coding agent diagnose, implement and prove. A person decides what ships.

The nine skills

SkillUse it whenLoop step
neens-startYou’re not sure where to begin. Checks the connection, reads the loop, routes you to the right skillEntry point
instrument-agentNo traces yet, or traces arrive looking empty. Wires OpenTelemetry to Neens and proves a trace arrived with what Neens needs1 · Instrument
triage-failures”What’s failing?” Reads failure modes and real traces, sorts each failure by where it lives, records your verdicts2 · Triage
build-regression-setA failure is confirmed. Freezes its traces, plus passing controls, into an immutable golden set3 · Freeze
write-judgeA failure needs an automated check. Writes a pass/fail judge and measures it against your verdicts4 · Judge
fix-failureTime to fix. Implements a remediation in your repo, proves it against the failure’s golden set, reports the PR and proof back5 · Fix
gate-release”Can we ship this?” Replays every regression set against the candidate and returns GO, NO-GO or NO DATA with the evidence6 · Gate
choose-model”Which model?” Runs a pass^k model sweep and reads it honestlyModel choice
reliability-review”How are we doing?” A one-page status report: KPIs, trends, top failure modes, the fix backlog, which fixes heldReporting

The usual path for a new agent runs straight down the loop:

instrument-agent → triage-failures → build-regression-set → write-judge → fix-failure → gate-release

Install

Connect Neens

The skills drive the Neens MCP tools, so connect first. See Connect for every client.

Add the skills

Inside a Claude Code session:

/plugin marketplace add neens-ai/skills
/plugin install neens@neens

The skills then appear namespaced by the plugin, as neens:triage-failures, neens:fix-failure and so on.

Start

Ask in plain words. Your coding agent picks the matching skill from its description.

Example prompts

PromptSkill it triggers
”Where do I start with Neens?”neens-start
”Send this agent’s traces to Neens and check they arrive.”instrument-agent
”What’s failing in my agent this week?”triage-failures
”Freeze the refund-policy failures into a regression set.”build-regression-set
”Write a judge for answers that contradict the tool result.”write-judge
”Fix the top open Neens remediation and prove it against my preview deploy.”fix-failure
”Can we ship branch fix/refund-policy?”gate-release
”Is the smaller model good enough for this agent?”choose-model
”Give me this week’s reliability review.”reliability-review

Requirements

  • A Neens MCP connection scoped to the agent you’re working on. Traces are needed for every skill except instrument-agent.
  • An LLM connection in Settings → LLM providers. Judges and remediations run on it; see LLM connections.
  • A way to run the candidate for fix-failure, gate-release and choose-model: your agent reachable over HTTP from Neens (a preview deploy, for example), or a harness of your own that replays golden inputs and sends the traces. Each skill explains both options, and Pre-prod evaluations covers them in full.

The skills never merge and never deploy. fix-failure opens a pull request with its proof attached, and gate-release returns a verdict with its evidence. The decision stays with you.

Write your own

The nine skills cover the loop every agent team runs. Your agent has its own failure modes, tools and release process, so write skills for those too:

  • Start from one of the nine as a template. A skill is a folder with a SKILL.md: front matter with a name and a description that says when to use it (and when not to), then the procedure.
  • Follow the eight rules above, especially evidence before opinion and a person owns the verdict.
  • Name only real tools and arguments. The MCP tools reference lists every one; the skills repo includes a checker that validates tool names and arguments against the live tool surface.