GuidesWatch mode

Watch mode: run your evals on every save

neens eval watch brings the pre-prod evaluation loop into your editor. It watches your agent’s source files, and every time you save it replays your golden dataset against your local agent, gates the result, and prints what changed since the previous run: which items newly fail (▼), which newly pass (▲), which still fail, the gate verdict, and how the headline numbers moved. You find a regression while you are still editing the prompt, not in CI.

At a glance

Commandneens eval watch (Python) · npx neens-eval watch (TypeScript / Node)
Each iterationA new, normal pre-prod run, the same as neens eval run --create: it appears on the Pre-prod Evals page and is gated the same way
AuthYour nk_live_ agent API key (NEENS_API_KEY), the same as neens eval run
NeedsA dataset with a golden version, enabled judges, and an agent that exports OpenTelemetry traces to Neens (see Pre-prod evaluations)
CostEvery iteration scores every golden item with your judges, on your LLM connection
⚠️

Every save costs judge tokens. Each iteration creates a new run and scores every golden item with your agent’s judges on the LLM connection you configured (Settings → LLM providers). The watcher says so when it starts. Keep the golden set small for local work, narrow what triggers a run with --include, and cap a session with --max-iterations N.

Install

Watch mode ships in both neens-eval SDKs from version 0.4.0, with the same flags and exit codes.

pip install "neens-eval>=0.4.0"   # installs the `neens` console script

Start watching

Point the SDK at Neens

export NEENS_BASE_URL="https://neens.example.com"   # your Neens origin
export NEENS_API_KEY="nk_live_..."                   # your agent API key

Your agent must also export its OpenTelemetry traces to Neens, with the same key:

export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT="$NEENS_BASE_URL/v1/traces"
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer $NEENS_API_KEY"

The watcher tags every trace for you, exactly as neens eval run does.

Find your golden dataset’s id

Use any dataset that has a golden version (Datasets → the dataset → Versions). To list the ids your key can see:

curl -s -H "Authorization: Bearer $NEENS_API_KEY" "$NEENS_BASE_URL/api/datasets"

Each entry in datasets carries its id and name.

Run the watcher

Everything after -- is your agent command. It runs once per golden prompt, with the prompt on stdin and in $NEENS_ITEM_INPUT.

neens eval watch --dataset ds_golden --version-label feat-refunds \
  --include "*.py" --include "prompts/*" \
  -- python -m myagent

The first iteration runs straight away and becomes the baseline for the next diff. It prints 12 item(s), 2 not passing — baseline for the next diff and lists the items that are not passing. After that, each save starts a new iteration. Press Ctrl-C to stop.

Reading the output

Every iteration prints a header, the normal run log, and then the diff against the previous iteration that produced results:

── iteration 3 · feat-refunds-watch.20261003T091530.3 ────
  changed: src/agent.py
  created run ppr_7c1e09a4b2d3f5e6
Running 12 item(s) with concurrency=1 (version='feat-refunds-watch.20261003T091530.3')
  …
  gate: FAILED (was PASSED at #2) · exit 1
  pass rate 83.3% (-8.3 pts) · avg score 0.812 (-0.041) · passed 10 (-1) · failed 2 (+1) · errored 0 (+0) · regressions 1 (+1)
  vs iteration #2:
    ▼ newly failing: 1
        ▼ dvi_4f1a9c  pass → fail  "What's the refund window for opened items?"
    ▲ newly passing: 0
    ✗ still failing: 1
        ✗ dvi_9c2e07  fail → fail  "Can I return a gift without a receipt?"
    = unchanged passing: 10
LineMeaning
gate:This iteration’s verdict (server gate plus your gate policy, if any), and the previous one
metrics linePass rate, average score, passed / failed / errored item counts, and regressions against the run’s baseline, each with its change since the previous iteration
▼ newly failingPassed last time, does not pass now
▲ newly passingDid not pass last time, passes now
✗ still failingDid not pass in either iteration
= unchanged passingHow many items passed both times
only in this run / missing from this runThe item exists on one side only, usually because the dataset’s golden version changed between iterations

Items are matched by their golden-item id, so the diff is item by item, not just a pass-rate delta. An item that errored or was never scored counts as not passing. It is never counted as a pass, and it is never silently dropped from the diff.

What triggers a run

The watcher polls the watched files for changes. It needs no file-system notification support, so it behaves the same on macOS, Linux and in containers.

FlagDefaultEffect
--path Pthe current directoryA file or directory to watch. Repeatable.
--include GLOBevery fileOnly files matching a glob start a run. Repeatable. *.py matches at any depth, and prompts/* matches paths under prompts/.
--exclude GLOBnoneSkip matching files, and whole directories that match. Repeatable.
--debounce S1Seconds of quiet after the last save before a run starts, so a burst of saves (or a formatter rewriting files) becomes one run.

These directories are always skipped, at any depth: .git, .hg, .svn, node_modules, __pycache__, .venv, venv, dist, build, .mypy_cache, .pytest_cache, .ruff_cache, .tox, .next. A tree with more than 5,000 watched files prints a warning. The scan stops at 50,000 files, so narrow a large repository with --path or --include.

Runs never overlap. If you save while an iteration is running, the watcher finishes that iteration first and then runs exactly one follow-up, however many saves you made meanwhile.

Files your agent writes don’t loop. If your agent writes into the watched tree on every run (a log file, a local database, a trace file), the watcher notices that the same file changed during two runs in a row, prints a warning, and stops re-running for that file’s changes made during a run. A save you make to it while the watcher is idle still starts a run. To be explicit, exclude it: --exclude "*.log".

Version labels

Each iteration is stamped with its own version label, <label>-watch.<session>.<n>, for example feat-refunds-watch.20261003T091530.3. <label> is --version-label (default local), <session> is the UTC time the watcher started, and <n> is the iteration number. Every run on the Pre-prod Evals page, and every trace it produced, can be traced back to the save that caused it.

Use it in scripts and git hooks: --once

--once runs a single iteration, without watching anything, and exits with the gate code. That is the same contract as neens eval run, so it works in a pre-commit or pre-push hook:

# .pre-commit-config.yaml
repos:
  - repo: local
    hooks:
      - id: neens-eval
        name: Neens pre-prod gate
        entry: neens eval watch --once --dataset ds_golden --version-label pre-commit -- python -m myagent
        language: system
        pass_filenames: false
        stages: [pre-push]

Install the hook type with pre-commit install --hook-type pre-push. Without pre-commit, a plain .git/hooks/pre-push (or a husky pre-push) works the same way:

#!/bin/sh
npx neens-eval watch --once --dataset ds_golden --version-label pre-push -- node dist/myagent.js

A hook runs on every push, and every push spends judge tokens. Most teams put it on pre-push rather than pre-commit.

Machine-readable output: --json

With --json, the watcher writes one JSON object per iteration on its own line to stdout. Human logs still go to stderr. Each object carries iteration, runId, versionLabel, exitCode, gatePassed, policyBlocked, error, trigger (the changed files), metrics, comparedTo, deltas, and diff (newlyFailing, newlyPassing, stillFailing, unchangedPassing, onlyInCurrent, onlyInPrevious). diff is null on the first iteration.

Stopping, errors and exit codes

SituationWhat happensExit code
--onceOne iteration, then exit0 passed · 1 failed · 2 run error
Ctrl-CYour agent’s running processes are stopped and the in-flight run is cancelled130
--max-iterations N reachedStops after the Nth runthe last iteration’s gate code
One iteration errors (a backend error, scoring timed out)The error is printed, and a run that did not finish is cancelled. The watcher keeps going, and the next diff compares against the last iteration that finished—
The run cannot be created (bad key, no access, unknown dataset, no golden version)Stops: saving a file will not fix it2
All flags for neens eval watch
FlagPurpose
--dataset IDGolden dataset every iteration replays. Required.
--dataset-version-id IDPin an explicit dataset version (defaults to the golden version).
--version-label LABELBase label (default local); see Version labels.
--name NAMERun name (defaults to <label> (watch #N)).
--baseline SPECprod, prod:<range> or run:<preprod_run_id>, the same as neens eval run.
--max-regressions N / --min-pass-rate RServer gate thresholds.
--gate-policy PATH and the policy override flagsThe same gate-as-code options as neens eval run.
--path / --include / --exclude / --debounceSee What triggers a run.
--max-iterations NStop after N runs.
--onceOne iteration, exit with the gate code.
--timeout, --concurrency, --poll-interval, --poll-timeoutThe same as neens eval run. --poll-timeout applies to each iteration.
--no-gateReport the gate but count it as a pass (--once exits 0 unless the run errors, which is 2).
--jsonOne JSON line per iteration on stdout.
--base-url, --api-key, --project-idConnection; env NEENS_BASE_URL, NEENS_API_KEY, NEENS_PROJECT_ID.

Troubleshooting

SymptomCauseFix
Every iteration times out with no items scoredYour agent’s traces are not reaching Neens, so nothing can be scoredCheck the agent exports OTLP traces to Neens (see Send traces)
A save does not start a runThe file is outside --path, does not match --include, matches --exclude, or sits in an always-skipped directoryAdjust the flags; the startup line shows how many files are watched
Runs start while you are still typingYour editor autosavesRaise --debounce, or narrow --include to the files you change on purpose
Watcher exits with 2 at startThe run could not be created: a wrong key, an unknown dataset id, or a dataset with no golden versionFix the key or dataset, or mark a version golden (Datasets → Versions)