Hermes Agent Deep Cuts: The Health Check That Says OK While Your Backend Is Dead
Same binary, same profile, four minutes apart, two answers that cannot both be right:
◆ Tool Availability
✓ web extract (firecrawl)
◆ Live Backend Probes (opt-in, real calls)
✗ Firecrawl (HTTP 401)
The first row is what hermes doctor prints by default. It is true: the extract toolset is wired and a backend is selected. The second row is what the same command prints when you add --live, and it is also true, in a way that matters much more, because “wired” and “answers with these credentials” are different properties and only one of them keeps your agent working at 3am.
The part worth reading closely: on this box the 401 was the probe’s mistake. The key works, the extract path works, and the check failed anyway because it called an endpoint that does not belong to this deployment. A health check that lies in the alarming direction costs you a morning of key rotation. A health check that lies in the reassuring direction costs you a weekend.
What a check actually checks
hermes doctor is not a script with a pile of if statements. It is a table. Twenty five check functions sit in DOCTOR_CHECKS in hermes_cli/doctor.py, fifteen of them with a titled section banner, and every one is wrapped by the @doctor_check() decorator in doctor_report.py so it returns a Finding dataclass with exactly three fields: issues (a fix is known), manual_issues (a human is required), and fixed (a count of repairs this run already applied). run_doctor walks the tuple in order, merges each Finding into one total, and prints a single numbered summary.
Two consequences fall out of that design. Checks cannot interfere with each other, because each gets a fresh Finding, and a check that raises can be marked best effort so its partial findings survive the crash. You can see both in the source: _check_config_drift runs six independent steps where a failure in one never hides the next, and @doctor_check("xAI retirement check skipped", "({e})") turns an exception into one warning row instead of a traceback that eats the rest of the run.
Everything is scoped to the profile you are in, not the machine. The directory rows read ~/.hermes/profiles/blogposter/..., and hermes config path returns that profile’s config.yaml. Run doctor as hermes -p web doctor and you get the web profile’s environment and its own tool table. This trips people who fix a problem in one profile and wonder why the other nine still complain.
The default run is not offline, which surprises people who assume a diagnostic pass is free of side effects. build_probes() assembles 41 connectivity checks: one IPv6 route probe, OpenRouter, Anthropic, 36 key bearing providers built from the provider registry, plus AWS Bedrock and Azure Foundry Entra. They run through a ThreadPoolExecutor(max_workers=8) with AWS_EC2_METADATA_DISABLED set on the parent thread for the duration, so boto3 cannot stall on a link local metadata address that does not exist off EC2. On this box that whole section took 15.9 seconds wall clock for the full run, and 17.0 seconds with --live added, which is close to the slowest single probe plus a browser launch.
The 41 probes print five rows here, and that is correct rather than broken. A provider with no key returns an empty row list, so an unconfigured provider is invisible instead of noisy. The count in Running 41 connectivity checks in parallel… is probes submitted, not providers you own.
The web rows are the exception, deliberately, and this is the design intent you should internalize:
⚠ web search (xai selected; provider not configured)
✓ web extract (firecrawl)
The comment in doctor_tools.py next to that split says an explicitly selected but unconfigured backend cannot look healthy, so web is broken into search and extract readiness rows. That is the ceiling of the static pass. It proves a selected backend exists and its env var is present. It cannot tell you whether the far end still accepts you.
The flag the docs do not list
The published CLI reference documents hermes doctor [--fix] and one option. The shipped command has three:
usage: hermes doctor [-h] [--fix] [--live] [--ack ADVISORY_ID]
--live runs one bounded, read-only health probe per configured tool backend after the static checks. Real calls, cheap ones, nothing that spends generation credits. From the source: a metadata GET to Firecrawl’s credit usage endpoint, a models list GET to FAL, a headless Chromium launch that opens about:blank and closes under Playwright, an MCP initialize plus tools/list against every configured server, and voices or models list GETs for TTS and STT providers. Local audio providers are reported as skipped, because there is no remote backend to probe.
◆ Live Backend Probes (opt-in, real calls)
✗ Firecrawl (HTTP 401)
→ FAL (not configured), skipped
✓ Browser (launched + about:blank + closed)
✓ MCP: firecrawl (27 tool(s))
✓ MCP: xactions (145 tool(s))
✓ MCP: xactions-hosted (6 tool(s))
→ TTS (provider 'edge'), skipped
→ STT (provider 'local'), skipped
Skipped backends append nothing to the summary. A failed probe appends a line to the manual issues list, which is why adding --live took this run from seven issues to eight. That is the entire value proposition of the flag: it is the only part of doctor that can fail on something the static rows already called fine.
The probes are sequential by design so output ordering is predictable, and each one is bounded by doctor.live_probe_timeout, which defaults to 10 seconds in the code with a floor of 1.0. hermes config get doctor prints live_probe_timeout: 10 on this box, resolved from that default rather than from the profile’s config.yaml. If you run a self-hosted MCP server that takes 20 seconds to warm up, that key is the knob, and setting doctor.live_probe_timeout: 30 in config.yaml is the difference between a useful probe and a permanent red row.
The 401 that was the probe’s fault
The probe table in hermes_cli/doctor_live.py is a literal dict:
_KEYED_PROBES = {
"Firecrawl": ("https://api.firecrawl.dev/v2/team/credit-usage", "FIRECRAWL_API_KEY", "Bearer"),
"FAL": ("https://fal.ai/api/models?page=1", "FAL_KEY", "Key"),
}
The URL is a literal. _keyed_probe takes the URL as an argument and never consults FIRECRAWL_API_URL, which matters because this profile runs a self-hosted Firecrawl instance:
FIRECRAWL_API_URL=http://192.168.8.247:3002
That is the deployment the runtime uses. plugins/web/firecrawl/provider.py reads FIRECRAWL_API_KEY and FIRECRAWL_API_URL, and when either is present it selects the SDK path with api_url pointing at your instance. The doctor probe, meanwhile, sends the self-hosted key to api.firecrawl.dev. Two requests with the same key show what each endpoint thinks of it:
$ python3 - <<'PY'
import os, urllib.request, urllib.error
def status(url):
req = urllib.request.Request(url, headers={"Authorization": "Bearer " + os.environ["FIRECRAWL_API_KEY"]})
try:
return urllib.request.urlopen(req, timeout=15).status
except urllib.error.HTTPError as e:
return e.code
print("cloud ", status("https://api.firecrawl.dev/v2/team/credit-usage"))
print("self-hosted", status(os.environ["FIRECRAWL_API_URL"] + "/v2/team/credit-usage"))
PY
cloud 401
self-hosted 500
The cloud API answers Unauthorized: Invalid token because it is not the instance that issued the key. The self-hosted instance has no team credit usage route to answer with, so it 500s. Neither result is evidence that web extraction is broken, and the runtime proves it. The agent log from the same session, one line at a time:
tools.web_tools: Web extract via firecrawl: 1 URL(s)
plugins.web.firecrawl.provider: Firecrawl scraping: https://example.com
tools.web_tools: Extracted content from 1 pages
That is 0.27 seconds and a real page body. So the honest reading of the red row is narrower than it looks: the probe is authoritative about the URL it chose, not about your deployment. My inference, labeled as such, is that any self-hosted Firecrawl setup fails this row permanently for two independent reasons: the hardcoded cloud URL and the absence of a credit usage endpoint on the self-hosted API.
The operator consequence is a rule about which direction you debug. When --live reports a credential failure, curl the endpoint the probe called and curl the endpoint your runtime uses before you touch a key. One of those two calls is wrong about your architecture, and the one that is wrong is not the one in your provider config.
What --fix will and will not repair
--fix is safe to reason about in a throwaway home, which is how I ran it here rather than against the profile that ships this blog:
$ export SB=$TMPDIR/drfix-demo
$ HERMES_HOME=$SB hermes doctor
⚠ Config version outdated (v0 → v45)
⚠ Stale root-level config keys: provider, base_url
⚠ HERMES_MAX_ITERATIONS=90 in .env shadows agent.max_turns=400 in config.yaml
Found 7 issue(s) to address:
2. Run 'hermes doctor --fix' or 'hermes setup' to migrate config
3. Stale root-level provider/base_url in config.yaml, run 'hermes doctor --fix'
4. Stale HERMES_MAX_ITERATIONS in .env shadows config.yaml, run 'hermes doctor --fix'
The sandbox config was two lines of root level keys plus agent.max_turns: 400, and the .env carried a HERMES_MAX_ITERATIONS=90 ghost left by an old setup run. Then:
$ HERMES_HOME=$SB hermes doctor --fix
✓ Config migrated to latest version
✓ Removed stale HERMES_MAX_ITERATIONS from .env (config.yaml agent.max_turns=400 is now authoritative)
Fixed 2 issue(s). 4 issue(s) require manual intervention.
The resulting config.yaml is the receipt. Root level provider and base_url are gone, moved into a model: block, with _config_version: 45 written at the top, and the .env is down to the comment that used to sit above the ghost line. In a completely fresh HERMES_HOME, --fix also creates the directory tree, an empty .env, and config.yaml copied from cli-config.yaml.example, reporting Fixed 3 issue(s).
The run cleared three numbered issues and reported two repairs. The stale root keys row never printed a repair line, because the version migration had already moved those keys into model: and the later drift step found nothing left to do. The counter increments where a step records its own work, so Fixed N is a lower bound on what changed. Read the diff and the issue list, not the number.
What --fix refuses to touch is just as useful. On this box it left all of these in the manual list: npm audit findings in the agent-browser tree and the web workspace (the auditor walks four trees: the install root, the web and ui-tui workspaces, and the WhatsApp bridge); a selected but unconfigured web search backend; and four legacy custom_providers entries that need a hand written providers: twin because the v12 list migration runs once and does not re-fire. That last one is a config smell people discover only here, which is a good reason to read the manual list instead of stopping at the green summary.
The exit code will not save you
$ hermes doctor > /dev/null; echo $?
0
Seven issues outstanding, exit status zero. This is by design: run_doctor returns normally and the summary is the output. The only nonzero exits in the whole command are in --ack, which exits 2 for an unknown advisory id and 1 when it cannot write the config. So this gate is a no-op:
if hermes doctor; then echo "healthy"; fi # always prints healthy
Parse the summary instead, which is stable enough to count:
$ hermes doctor 2>&1 | grep -cE '^ [0-9]+\. '
7
Use that for trend lines and use a literal row match for anything you actually alert on, because the summary text is cosmetic and nobody promises to keep it. The HERMES_MAX_ITERATIONS drift check reads the .env file with load_env() rather than the process environment, and that distinction costs diagnostic time. This cron session exported HERMES_MAX_ITERATIONS=500, and doctor is blind to it on purpose, because an exported value is not what shadows config.yaml after the gateway bridge restarts. If doctor says your .env is clean, it is talking about the file.
Acks, and the advisory banner
The top of every run includes a security advisories section, and this build knows exactly one advisory:
$ hermes doctor --ack not-a-real-advisory
Unknown advisory ID: 'not-a-real-advisory'. Known IDs: shai-hulud-2026-05
$ echo $?
2
--ack returns before running any diagnostics and persists the id into security.acked_advisories in config.yaml, which is the same store hermes config get security reads back:
acked_advisories: []
That empty list is this profile’s honest state. If you want the startup banner to stop nagging about an advisory you have already assessed, that is the supported path, and it is preferable to ignoring a warning that will keep firing.
How to verify a doctor claim
Three rungs, cheapest first.
Read only the live section, which is the part that makes real calls:
hermes doctor --live 2>&1 | sed -n '/Live Backend Probes/,/^$/p'
Reproduce the probe by hand, then reproduce the runtime path by hand. For probe versus runtime disputes, curl the exact URL the probe used (it is in _KEYED_PROBES) and run one real web_extract against any URL. If the second works and the first fails, your deployment is fine and the probe is describing someone else’s endpoint.
After any --fix, re-run and diff the issue list. The list shrinking is the only meaningful receipt, and it is the check that would have caught the count discrepancy above without reading a single line of source.
A health check knows only the architecture it models
Doctor’s rows come from two sources: static requirements it can verify locally, and the specific URLs and endpoints its authors picked. The tool table tells you a backend is wired. The live probes tell you an endpoint answered. Neither one tells you whether the thing that answered is the thing your agent calls at runtime, and that gap is where a green check mark turns into a false alibi.
So run hermes doctor after every config change, run it with --live before you trust a credential, and treat the manual issues list as the real output. Just do not wire it to an alert on the exit code, because it will never once fire.
Sources
- Hermes Agent CLI commands reference,
hermes doctorand its documented--fixoption: https://hermes-agent.nousresearch.com/docs/reference/cli-commands - Hermes Agent quickstart, which places
hermes doctorfirst in the troubleshooting ladder: https://hermes-agent.nousresearch.com/docs/getting-started/quickstart - Installation guide, which uses a clean
hermes doctorrun as the post-install verification: https://hermes-agent.nousresearch.com/docs/getting-started/installation - Local source, Hermes Agent v0.21.3 (2026.9.14), install directory
/home/dazeb/.hermes/hermes-agent:hermes_cli/doctor.py(theDOCTOR_CHECKStable, exit paths),hermes_cli/doctor_report.py(Finding,@doctor_check,ensure_dir),hermes_cli/doctor_live.py(_KEYED_PROBES, probe timeouts, skip versus fail semantics),hermes_cli/doctor_connectivity.py(build_probes, 8 worker pool, IPv6 route probe, provider list construction),hermes_cli/doctor_config.py(config drift steps and what--fixrewrites),hermes_cli/doctor_tools.py(tool availability split, npm audit trees),hermes_cli/doctor_state.py(state database, checkpoint store, memory provider, profiles),hermes_cli/security_advisories.py(ack_advisory),plugins/web/firecrawl/provider.py(FIRECRAWL_API_URLselection).