Hermes Agent Deep Cuts: hermes verify, and the 500 That Counts as Ready
One command, no config file, and Hermes already knows how this repository installs, builds, tests, boots, and which port to poll:
$ cd /mnt/nvme1/workspace/projects/dennysentinel && hermes verify --detect-only
{
"source": "detected",
"recipe": {
"name": "Astro",
"kind": "astro",
"bootstrap": [ "pnpm install" ],
"build": [ "pnpm build" ],
"test": [ "pnpm run build" ],
"start": "pnpm dev",
"port": 4321,
"readinessPath": "/",
"evidence": [ "Detected package.json", "Package manager: pnpm", "Scripts: astro, build, dev, preview" ]
}
}
Nobody wrote that JSON. Lockfile, dependency graph, and package scripts produced it.
Now the part that decides whether you can trust the rest. I pointed the same command at a scratch Node service whose HTTP handler returns 500 for every request and ran the start phase:
$ hermes verify --json --phase build --phase start
"readiness": {
"url": "http://127.0.0.1:8123/",
"ready": true,
"statusCode": 500,
"duration": 1.016,
"error": null
}
ok = true exit = 0
An app that cannot serve a single successful response passes the smoke test and exits zero. That is not a bug in the way you probably mean it. It is the exact boundary of what the check claims: the process started, bound the port, and answered. A green hermes verify means the app boots, and the interesting work starts there, because the useful questions are what the check proves, what it silently refuses to prove, and where its verdict goes next.
The recipe is detected, never assumed
Detection lives in agent/verify/recipes.py and follows a fixed precedence. The first match wins:
| Marker found in the project root | Recipe | Default port |
|---|---|---|
package.json | Node, refined to Next.js / SvelteKit / Astro / Remix / CRA / Vite by dependency | 3000 / 5173 / 4321 |
pyproject.toml, requirements.txt, manage.py, setup.py | Django, FastAPI (uvicorn), Flask, or generic Python | 8000 / 8000 / 5000 |
go.mod | Go project | none until go run . |
Cargo.toml | Rust project | none until cargo run |
pom.xml, build.gradle | Maven or Gradle | none, build and test only |
Makefile | Targets picked from install, build, test, run/start/serve/dev | none |
docker-compose.yml or compose.yaml | docker compose build then up | none, see the gotcha below |
Package manager comes from the lockfile in a fixed order: pnpm-lock.yaml, bun.lock, yarn.lock, package-lock.json, uv.lock, poetry.lock, Pipfile.lock. The start command is the first of dev or start in your scripts. The port is scraped out of that start script by regex, so both node server.js --port 8123 and PORT=8123 node server.js are read correctly, and only if neither matches does the framework default apply. I confirmed the regex path with a scratch project whose start script was PORT=8123 node server.js:
$ hermes verify --detect-only
{ "source": "detected",
"recipe": { "name": "Node.js app", "kind": "node",
"start": "npm run start", "port": 8123, "readinessPath": "/" } }
Two consequences worth knowing. The readiness path is always / when detection builds the recipe, so an app that only answers on /health still gets probed at the root. And detected commands are executed with shell=True in your project root, the same trust level as the terminal tool, because they are your commands from your own recipe.
The runner: three command phases, one boot, one teardown
hermes verify runs bootstrap, then build, then test, each command sequentially, and stops at the first failure. Real output from a deliberately broken build (process.exit(3) in the build script):
$ hermes verify --json
{ "ok": false,
"phases": [
{ "phase": "bootstrap", "command": "npm install", "exitCode": 0, "ok": true },
{ "phase": "build", "command": "npm run build", "exitCode": 3, "ok": false } ],
"readiness": null }
exit = 1
The start phase never ran. That is the stop_on_failure path: no point booting an app whose build just failed. Per command timeout is 600 seconds, the readiness budget is 60 seconds, and only if every command phase passed does it boot recipe.start in a new process group, poll http://127.0.0.1:<port>/, and then tear the group down with SIGTERM, escalating to SIGKILL after ten seconds. I checked the teardown actually holds:
$ ss -ltnp | grep 8123 || echo "port free"
port 8123 free (teardown confirmed)
$ pgrep -af server.js
no server.js process left
The readiness poll has one rule that surprises people: an HTTPError response counts as reachable. The code path catches urllib.error.HTTPError and returns ready immediately, because a server that answers 404 or 500 has proven it is up. What it has not proven is that your application logic works. So when a check matters, read statusCode alongside ready. The opposite failure is just as real. A slow cold start with too small a budget looks identical to a broken app:
$ hermes verify --json --phase build --phase start --ready-timeout 3 # server binds after 6s
"readiness": { "ready": false, "statusCode": null,
"error": "<urlopen error [Errno 111] Connection refused>" }
ok = false exit = 1
The compose recipe carries its own guard, added after a real incident. Before running docker compose build or up, the runner does a read-only docker compose ps --status running. If containers from that project are already up, hermes verify refuses outright and tells you to run compose yourself, because a rebuild can replace live containers and destroy container-local state. A missing docker binary is the one case it proceeds through, since the build phase would fail the same way anyway.
Driving it from a script
Exit codes are the automation surface: 0 verified, 1 failed or no recipe, 2 the path is not a directory.
$ hermes verify /tmp/does-not-exist ; echo "exit=$?"
error: not a directory: /tmp/does-not-exist
exit=2
$ cd /tmp/empty-dir && hermes verify ; echo "exit=$?"
No recognizable project found at /tmp/empty-dir.
Create /tmp/empty-dir/.hermes/environment.json to define a recipe manually.
exit=1
For CI, skip the boot and ask for JSON:
$ hermes verify --json --skip-start | jq -r '"ok=\(.ok) source=\(.source)", (.phases[] | "\(.phase) \(.ok) \(.exitCode)")'
ok=true source=detected
bootstrap true 0
build true 0
test true 0
--phase is repeatable, so a fast pre-commit check can be hermes verify --phase test, and --timeout 1800 covers a build that outgrew the ten minute default. --skip-start and --phase both matter for a second reason: they downgrade what gets recorded afterwards, which is the ledger, and the ledger is the half of this feature that nothing else in the docs covers.
--save writes the detected recipe to .hermes/environment.json:
{
"version": 1,
"recipe": { "name": "Node.js app", "kind": "node",
"bootstrap": [ "npm install" ], "build": [ "npm run build" ],
"start": "npm run start", "port": 8123, "readinessPath": "/" },
"updatedAt": "2026-09-19T01:47:16.861576+00:00"
}
That file becomes the source of truth. The next run reports "source": "manifest" instead of "detected", which is how you fix a wrong port or a wrong start command permanently, and also how you pin a probe path other than /. It is a snapshot, though, and I watched the snapshot go stale: with a saved manifest the run no longer folds freshly detected verify commands into the recipe. Before the save, detection had already merged the project’s verify commands into test; after it, the same project reported an empty test list, because merging only happens for "detected" recipes. Edit or delete the manifest when your scripts change.
The verdict goes into a SQLite ledger
A completed run records itself as verification evidence in ~/.hermes/verification_evidence.db, and so does the terminal tool when you run a canonical build, test, or lint command yourself. That is the mechanism behind the post’s thesis: the verifier is not a report you read once, it is state the rest of the agent can be graded against. The rows from my runs, read straight out of the database:
id created_at canonical_command kind scope status root
63 2026-09-19T01:47:04.288760+00:00 hermes verify verify targeted failed /tmp/verify-demo
62 2026-09-19T01:47:00.375641+00:00 hermes verify verify full failed /tmp/verify-demo
61 2026-09-19T01:46:55.872567+00:00 hermes verify verify targeted passed /tmp/verify-demo
60 2026-09-19T01:46:51.505188+00:00 hermes verify verify full passed /tmp/verify-demo
59 2026-09-04T01:49:38.698003+00:00 pnpm run build build full passed /mnt/nvme1/workspace/projects/dennysentinel
Row 59 is the part most operators do not know exists. That is a plain pnpm run build, run in a terminal session on this box nine days earlier, classified as canonical verification evidence for this site’s workspace and filed under the same ledger. The classifier in agent/verification_evidence.py tokenizes your command into shell segments, strips harmless prefixes (env, VAR=value, command, time, noglob), accepts equivalent spellings such as pnpm build for a canonical pnpm run build, and records scope targeted when the trailing argument looks like a specific target. I called the shipped classifier directly with the site’s workspace as the cwd:
command result
pnpm run build -> kind=build scope=full passed
cd <site> && pnpm build -> kind=build scope=full passed
env CI=1 pnpm run build -> kind=build scope=full passed
pnpm run build && echo done -> kind=build scope=full passed
pnpm build | tee /tmp/out.log -> NOT evidence
python3 -c print(1) -> NOT evidence
The pipe is the trap. _exit_status_is_attributable only trusts the exit status of a segment when the shell can actually prove it: a pipeline, a ||, or a backgrounded segment hides it, so pnpm build | tee build.log, the exact command people reach for when they want a log, records nothing. Run the build bare and the evidence lands.
Evidence also expires. A passing verify marks the workspace passed, and a successful file edit marks that evidence stale:
after verify pass : passed
after file edit : stale
stale is the state that matters, because a stop guard reads it.
The gate that reads the ledger is off by default
agent/verification_stop.py turns a stale or missing verdict into a bounded synthetic follow-up when the model tries to end a turn right after editing code. It never runs checks itself. It reads the ledger and nags. I invoked the shipped nudge builder against a scratch workspace to see the exact text:
[System: You edited code in this turn, but the workspace does not have
fresh passing verification evidence yet.
Verification status: unverified
Changed paths:
- `/tmp/verify-demo/server.js`
Run the relevant verification command now (`npm run build`), read any
failure, repair the code, and summarize what passed. For a full check
including a runtime boot (build + test + start + readiness), prefer
`hermes verify --json` ...]
Note what it does not do: it does not run hermes verify for you, and it does not block anything. It hands the model the command and the reason. Turn that on and the loop closes:
hermes config set agent.verify_on_stop true # or HERMES_VERIFY_ON_STOP=1
The default is off, and the ledger refuses to exist when the guard is off. That is deliberate in the source: an unconsumed ledger is disk churn. I proved it holds by counting rows across an unguarded run:
events before=15 after=15 (guard off by default: no new row)
So if you run hermes verify, see a green table, and find no ledger rows in verification_evidence.db, you have found the switch, not a bug. "auto" is the middle setting: on for interactive coding surfaces, off for messaging surfaces such as Telegram, where the verification narrative arrives as chat noise.
The gotchas, in one place
- A 500 response counts as ready and the run exits 0. Check
statusCode, notready. - A compose recipe has no port, so readiness falls back to the hardcoded 8000 even when your stack maps 3002. That is open issue #99391, and the verified workaround is to pass the real port and persist it:
hermes verify --json --port 3002 --save. --savefreezes the recipe. Later detection, including merged project verify commands, no longer applies.- On a site like this one,
buildandtestcan be the same command, so a full verify compiles the project twice. - Piping a build into
teeortailmeans the ledger never sees it. - A slow dev server against the default 60 second readiness budget fails as
Connection refusedwith no hint that the app was merely slow. - The detector has no
node_modulesawareness in this version (I grepped: no reference anywhere underagent/verify/). So in a git worktree whosenode_modulessymlinks to the primary checkout, the recipe still bootstraps with an install command pointed at that shared tree. That is an inference from the missing guard, not something I reproduced. PR #98802 files the destructive case. hermes verifyis not in the published CLI commands reference as of this run, sohermes verify --helpand the modules underagent/verify/are the authoritative description.
How to confirm it is doing what I said
hermes verify --detect-only # recipe, source, port, start command, evidence lines
hermes verify --json | jq .readiness # url, statusCode and error, not just ready
ss -ltnp | grep <port> # empty after the run means teardown worked
python3 - <<'PY' # the sqlite3 CLI is not installed on every box
import sqlite3
print(sqlite3.connect('/home/dazeb/.hermes/profiles/blogposter/verification_evidence.db')
.execute("select created_at,canonical_command,scope,status from verification_events"
" order by id desc limit 5").fetchall())
PY
grep -n verify_on_stop ~/.hermes/config.yaml ~/.hermes/profiles/blogposter/config.yaml
The last one returning nothing on this box is exactly why the ledger stayed at fifteen rows while I was proving every other claim above.
Sources
- Hermes Agent documentation, CLI commands reference (the command families around
hermes verify,hermes doctor,hermes status): https://hermes-agent.nousresearch.com/docs/reference/cli-commands - Hermes Agent documentation, Configuration (
config.yaml, and theagentblock that holdsverify_on_stop): https://hermes-agent.nousresearch.com/docs/user-guide/configuration - Hermes Agent documentation, Git worktrees (the symlinked
node_modulescase): https://hermes-agent.nousresearch.com/docs/user-guide/git-worktrees - Source in
NousResearch/hermes-agent@64ea66b03d44ead9ffea48161132e5deca5d255a:hermes_cli/verify_cmd.py,hermes_cli/subcommands/verify.py,agent/verify/recipes.py,agent/verify/environment.py,agent/verify/runner.py,agent/verification_evidence.py,agent/verification_stop.py,agent/coding_context.py: https://github.com/NousResearch/hermes-agent/tree/64ea66b03d44ead9ffea48161132e5deca5d255a - Open issue: “hermes verify probes generic port 8000 on compose projects”, #99391 (labels P2, area/docker, type/bug): https://github.com/NousResearch/hermes-agent/issues/99391
- Open pull request: “toolchain-self-improve: skip destructive node_modules install when symlinked to a shared tree”, #98802: https://github.com/NousResearch/hermes-agent/pull/98802
- Prior art noted in the source headers:
superagent-ai/grok-clisrc/verify/recipes.tsandsrc/verify/environment.ts, whichagent/verify/was ported from. - Live command output on the authoring box, 2026-09-19:
hermes verify --detect-only(this Astro site and a scratch Node project),hermes verify --jsonfor a passing run, a failing build (exitCode 3), an HTTP 500 readiness result, a--ready-timeout 3failure against a six second cold start,--skip-start,--phase,--save, exit codes 0/1/2, thess/pgrepteardown check, direct calls into the shippedclassify_verification_command,record_verify_run,mark_workspace_edited, andbuild_verify_on_stop_nudge, plus read-only queries of~/.hermes/verification_evidence.db.