Watch mode: run your evals on every save
neens eval watch brings the pre-prod evaluation loop into your
editor. It watches your agent’s source files, and every time you save it replays your golden
dataset against your local agent, gates the result, and prints what changed since the
previous run: which items newly fail (▼), which newly pass (▲), which still fail, the
gate verdict, and how the headline numbers moved. You find a regression while you are still
editing the prompt, not in CI.
At a glance
| Command | neens eval watch (Python) · npx neens-eval watch (TypeScript / Node) |
| Each iteration | A new, normal pre-prod run, the same as neens eval run --create: it appears on the Pre-prod Evals page and is gated the same way |
| Auth | Your nk_live_ agent API key (NEENS_API_KEY), the same as neens eval run |
| Needs | A dataset with a golden version, enabled judges, and an agent that exports OpenTelemetry traces to Neens (see Pre-prod evaluations) |
| Cost | Every iteration scores every golden item with your judges, on your LLM connection |
Every save costs judge tokens. Each iteration creates a new run and scores every golden item
with your agent’s judges on the LLM connection you configured (Settings → LLM providers). The
watcher says so when it starts. Keep the golden set small for local work, narrow what triggers a
run with --include, and cap a session with --max-iterations N.
Install
Watch mode ships in both neens-eval SDKs from version 0.4.0, with the same flags and exit
codes.
pip install "neens-eval>=0.4.0" # installs the `neens` console scriptStart watching
Point the SDK at Neens
export NEENS_BASE_URL="https://neens.example.com" # your Neens origin
export NEENS_API_KEY="nk_live_..." # your agent API keyYour agent must also export its OpenTelemetry traces to Neens, with the same key:
export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT="$NEENS_BASE_URL/v1/traces"
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer $NEENS_API_KEY"The watcher tags every trace for you, exactly as neens eval run does.
Find your golden dataset’s id
Use any dataset that has a golden version (Datasets → the dataset → Versions). To list the ids your key can see:
curl -s -H "Authorization: Bearer $NEENS_API_KEY" "$NEENS_BASE_URL/api/datasets"Each entry in datasets carries its id and name.
Run the watcher
Everything after -- is your agent command. It runs once per golden prompt, with the prompt on
stdin and in $NEENS_ITEM_INPUT.
neens eval watch --dataset ds_golden --version-label feat-refunds \
--include "*.py" --include "prompts/*" \
-- python -m myagentThe first iteration runs straight away and becomes the baseline for the next diff. It prints
12 item(s), 2 not passing — baseline for the next diff and lists the items that are not passing. After that,
each save starts a new iteration. Press Ctrl-C to stop.
Reading the output
Every iteration prints a header, the normal run log, and then the diff against the previous iteration that produced results:
── iteration 3 · feat-refunds-watch.20261003T091530.3 ────
changed: src/agent.py
created run ppr_7c1e09a4b2d3f5e6
Running 12 item(s) with concurrency=1 (version='feat-refunds-watch.20261003T091530.3')
…
gate: FAILED (was PASSED at #2) · exit 1
pass rate 83.3% (-8.3 pts) · avg score 0.812 (-0.041) · passed 10 (-1) · failed 2 (+1) · errored 0 (+0) · regressions 1 (+1)
vs iteration #2:
▼ newly failing: 1
▼ dvi_4f1a9c pass → fail "What's the refund window for opened items?"
▲ newly passing: 0
✗ still failing: 1
✗ dvi_9c2e07 fail → fail "Can I return a gift without a receipt?"
= unchanged passing: 10| Line | Meaning |
|---|---|
gate: | This iteration’s verdict (server gate plus your gate policy, if any), and the previous one |
| metrics line | Pass rate, average score, passed / failed / errored item counts, and regressions against the run’s baseline, each with its change since the previous iteration |
▼ newly failing | Passed last time, does not pass now |
▲ newly passing | Did not pass last time, passes now |
✗ still failing | Did not pass in either iteration |
= unchanged passing | How many items passed both times |
only in this run / missing from this run | The item exists on one side only, usually because the dataset’s golden version changed between iterations |
Items are matched by their golden-item id, so the diff is item by item, not just a pass-rate delta. An item that errored or was never scored counts as not passing. It is never counted as a pass, and it is never silently dropped from the diff.
What triggers a run
The watcher polls the watched files for changes. It needs no file-system notification support, so it behaves the same on macOS, Linux and in containers.
| Flag | Default | Effect |
|---|---|---|
--path P | the current directory | A file or directory to watch. Repeatable. |
--include GLOB | every file | Only files matching a glob start a run. Repeatable. *.py matches at any depth, and prompts/* matches paths under prompts/. |
--exclude GLOB | none | Skip matching files, and whole directories that match. Repeatable. |
--debounce S | 1 | Seconds of quiet after the last save before a run starts, so a burst of saves (or a formatter rewriting files) becomes one run. |
These directories are always skipped, at any depth: .git, .hg, .svn, node_modules,
__pycache__, .venv, venv, dist, build, .mypy_cache, .pytest_cache, .ruff_cache,
.tox, .next. A tree with more than 5,000 watched files prints a warning. The scan stops at
50,000 files, so narrow a large repository with --path or --include.
Runs never overlap. If you save while an iteration is running, the watcher finishes that iteration first and then runs exactly one follow-up, however many saves you made meanwhile.
Files your agent writes don’t loop. If your agent writes into the watched tree on every run (a
log file, a local database, a trace file), the watcher notices that the same file changed during
two runs in a row, prints a warning, and stops re-running for that file’s changes made during a
run. A save you make to it while the watcher is idle still starts a run. To be explicit, exclude it:
--exclude "*.log".
Version labels
Each iteration is stamped with its own version label, <label>-watch.<session>.<n>, for example
feat-refunds-watch.20261003T091530.3. <label> is --version-label (default local),
<session> is the UTC time the watcher started, and <n> is the iteration number. Every run on
the Pre-prod Evals page, and every trace it produced, can be traced back to the save that
caused it.
Use it in scripts and git hooks: --once
--once runs a single iteration, without watching anything, and exits with the gate code. That is
the same contract as neens eval run, so it works in a pre-commit or pre-push hook:
# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: neens-eval
name: Neens pre-prod gate
entry: neens eval watch --once --dataset ds_golden --version-label pre-commit -- python -m myagent
language: system
pass_filenames: false
stages: [pre-push]Install the hook type with pre-commit install --hook-type pre-push. Without pre-commit, a
plain .git/hooks/pre-push (or a husky pre-push) works the same way:
#!/bin/sh
npx neens-eval watch --once --dataset ds_golden --version-label pre-push -- node dist/myagent.jsA hook runs on every push, and every push spends judge tokens. Most teams put it on pre-push
rather than pre-commit.
Machine-readable output: --json
With --json, the watcher writes one JSON object per iteration on its own line to stdout.
Human logs still go to stderr. Each object carries iteration, runId, versionLabel,
exitCode, gatePassed, policyBlocked, error, trigger (the changed files), metrics,
comparedTo, deltas, and diff (newlyFailing, newlyPassing, stillFailing,
unchangedPassing, onlyInCurrent, onlyInPrevious). diff is null on the first iteration.
Stopping, errors and exit codes
| Situation | What happens | Exit code |
|---|---|---|
--once | One iteration, then exit | 0 passed · 1 failed · 2 run error |
| Ctrl-C | Your agent’s running processes are stopped and the in-flight run is cancelled | 130 |
--max-iterations N reached | Stops after the Nth run | the last iteration’s gate code |
| One iteration errors (a backend error, scoring timed out) | The error is printed, and a run that did not finish is cancelled. The watcher keeps going, and the next diff compares against the last iteration that finished | — |
| The run cannot be created (bad key, no access, unknown dataset, no golden version) | Stops: saving a file will not fix it | 2 |
All flags for neens eval watch
| Flag | Purpose |
|---|---|
--dataset ID | Golden dataset every iteration replays. Required. |
--dataset-version-id ID | Pin an explicit dataset version (defaults to the golden version). |
--version-label LABEL | Base label (default local); see Version labels. |
--name NAME | Run name (defaults to <label> (watch #N)). |
--baseline SPEC | prod, prod:<range> or run:<preprod_run_id>, the same as neens eval run. |
--max-regressions N / --min-pass-rate R | Server gate thresholds. |
--gate-policy PATH and the policy override flags | The same gate-as-code options as neens eval run. |
--path / --include / --exclude / --debounce | See What triggers a run. |
--max-iterations N | Stop after N runs. |
--once | One iteration, exit with the gate code. |
--timeout, --concurrency, --poll-interval, --poll-timeout | The same as neens eval run. --poll-timeout applies to each iteration. |
--no-gate | Report the gate but count it as a pass (--once exits 0 unless the run errors, which is 2). |
--json | One JSON line per iteration on stdout. |
--base-url, --api-key, --project-id | Connection; env NEENS_BASE_URL, NEENS_API_KEY, NEENS_PROJECT_ID. |
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| Every iteration times out with no items scored | Your agent’s traces are not reaching Neens, so nothing can be scored | Check the agent exports OTLP traces to Neens (see Send traces) |
| A save does not start a run | The file is outside --path, does not match --include, matches --exclude, or sits in an always-skipped directory | Adjust the flags; the startup line shows how many files are watched |
| Runs start while you are still typing | Your editor autosaves | Raise --debounce, or narrow --include to the files you change on purpose |
Watcher exits with 2 at start | The run could not be created: a wrong key, an unknown dataset id, or a dataset with no golden version | Fix the key or dataset, or mark a version golden (Datasets → Versions) |