Hermes Agent Deep Cuts: The Supply-Chain Audit Found 29 Advisories and Exited 0
The venv this agent runs from, audited on 2026-09-18, exit code last:
$ hermes security audit > /tmp/audit_human.txt; echo "exit=$?"
exit=0
$ head -15 /tmp/audit_human.txt
Found 29 known vulnerability finding(s) across 240 component(s):
[venv]
HIGH httpcore2==2.7.0 GHSA-7mj9-2mp8-4m2p
HTTPX2: Secure WebSocket traffic sent without TLS through SOCKS proxies
fixed in: 2.10.0
HIGH httpx2==2.7.0 GHSA-7mj9-2mp8-4m2p
HTTPX2: Secure WebSocket traffic sent without TLS through SOCKS proxies
fixed in: 2.10.0
HIGH httpx2==2.7.0 GHSA-8xx6-hgc6-gc2m
HTTPX2: Streaming response decompression does not bound peak memory (decompression amplification)
fixed in: 2.12.0
HIGH langchain-core==0.3.86 GHSA-qh6h-p6c9-ff54
LangChain Core has Path Traversal vulnerabilites in legacy `load_prompt` functions
fixed in: 1.2.22
HIGH tornado==6.5.7 GHSA-mpf4-983q-p7j4
Tornado: Urlencoded body parsing omits max_num_fields, so one request can stall the event loop
fixed in: 6.5.8
Five HIGH rows, thirteen UNKNOWN rows, nine MODERATE, two LOW, and a clean exit. The gate did not fire because its default threshold is --fail-on critical and not one finding in that list carries a critical grade, so the process returned 0 with five HIGH advisories on screen. Wire that command into CI as a pass/fail check without reading this post and your build goes green on a list you never looked at.
The same profile has a worse hole in it. That scan covered the venv, the plugin requirements files, and pinned npx/uvx MCP servers. This profile declares three MCP servers, and the audit touched none of them:
$ hermes security audit --skip-venv
No components discovered (everything skipped, or empty environment).
$ echo $?
0
Three configured MCP servers, zero audited, exit code 0. That is the shape of the whole feature: a real OSV.dev scan, aimed only at what you pinned, with a gate whose default setting is quieter than the output it just printed.
Three surfaces, three discovery rules
hermes security audit lives in hermes_cli/security_audit.py, which is about 300 lines of stdlib Python: urllib, importlib.metadata, tomllib, concurrent.futures. That was a deliberate choice in the original PR, which rejected shelling out to osv-scanner or pip-audit because users would have to install them. There is nothing to install and no auth, and the OSV.dev API it hits is the same data source osv-scanner uses.
The three discovery functions differ far more than the help text suggests.
The venv surface (_discover_venv) walks importlib.metadata.distributions() for the Python interpreter that is running hermes, deduplicating on lowercase name and version. For a git install that means the Hermes venv. It does not read a lockfile and it does not run pip, so it reports what is importable right now, including anything someone pip installed into that venv by hand. It also ignores HERMES_HOME completely. I ran the audit with HERMES_HOME pointing at a throwaway directory and it still scanned the same 240 components:
$ mkdir -p /tmp/audit-empty && HERMES_HOME=/tmp/audit-empty hermes security audit | head -1
Found 29 known vulnerability finding(s) across 240 component(s):
Profile isolation, which holds almost everywhere else in Hermes, does not apply here. Every profile on the box shares one venv and therefore one audit surface.
The plugin surface (_discover_plugins) reads requirements.txt, requirements-dev.txt, and pyproject.toml from each directory under $HERMES_HOME/plugins/. It does not install or resolve anything, so it audits the requirements a plugin states about itself. Both Python parsers are pin gatekeepers: _REQ_LINE only matches name==version, optionally with extras and an environment marker, and _parse_pyproject_pins runs the same filter over project.dependencies plus every group in optional-dependencies. Comments, -r other.txt includes, and any loose spec such as flask>=3.0 are dropped. The reason is in the code comment, and it is a good one: a loose spec cannot map to a single OSV query, and a tool that reports noise teaches you to ignore it.
The MCP surface (_discover_mcp) is the narrowest. It reads mcp_servers from config.yaml, looks at command for a basename ending in npx or uvx, then takes the first argument that does not start with -. That argument must match pkg@1.2.3 for npx or pkg==1.2.3 for uvx, and the code comment states the rule plainly: unversioned names resolve to latest at runtime and are not a stable audit subject. Anything else returns None and the server is skipped. URL-based servers, Docker images, local script paths, and unversioned launchers are all outside the scan, which is why the three MCP servers in this profile produce nothing.
Discovery is the cheap part. The queries are one POST to https://api.osv.dev/v1/querybatch with up to 1000 packages per request (chunked if you somehow exceed it), then a parallel detail fetch, 8 threads, one GET /v1/vulns/{id} per unique advisory, 20 seconds timeout each. A whole profile scan of 240 components took 1.9 seconds and a scan of 241 components took 4.4 seconds on this box, network variance included.
The gate is a severity-bucket lookup, not a CVSS score
--fail-on does not compute anything.
# Severity ordering for --fail-on gating. UNKNOWN sits below LOW so it never blocks.
SEVERITY_ORDER = {"UNKNOWN": 0, "LOW": 1, "MODERATE": 2, "MEDIUM": 2, "HIGH": 3, "CRITICAL": 4}
Severity comes from the OSV record, and that record is where the shape of the output comes from. Top-level severity in OSV is a CVSS vector string, and the audit ships no CVSS library, so it reads the GHSA database_specific.severity bucket first, then the per-affected ecosystem_specific.severity bucket, and falls back to UNKNOWN. GitHub advisories carry those buckets. PyPI advisories (PYSEC ids) mostly do not, which is why 13 of the 29 rows in my scan printed UNKNOWN: 9 of those still had a summary line, 4 printed nothing at all.
Exit codes are the documented three:
0: nothing at or above the threshold1: at least one finding at or above the threshold2: bad arguments or an OSV error
I checked each tier against the same unmodified environment:
$ for tier in low moderate high critical; do
> hermes security audit --fail-on $tier >/dev/null 2>&1; echo "fail-on $tier -> exit $?"; done
fail-on low -> exit 1
fail-on moderate -> exit 1
fail-on high -> exit 1
fail-on critical -> exit 0
With UNKNOWN ranked at 0, a row graded UNKNOWN cannot trip any tier, not even --fail-on low. No component in this venv had PYSEC-only findings, so I am reading that consequence off the rank map and the comment above it rather than reporting it as observed.
Proving the other two surfaces with a fixture
An audit you never falsified is a vibe. Because plugin and MCP discovery key off HERMES_HOME and config.yaml, you can point the whole command at a synthetic home and watch it discover exactly what you declared. I built one with four MCP entries and one plugin:
# /tmp/audit-demo/config.yaml
mcp_servers:
lodash-fixture:
command: npx
args: ["-y", "lodash@4.17.15"]
enabled: true
urllib-fixture:
command: uvx
args: ["urllib3==2.0.0"]
enabled: true
unpinned-fixture:
command: npx
args: ["-y", "some-unpinned-server"]
enabled: true
remote-fixture:
url: "https://example.com/mcp"
enabled: true
$ cat /tmp/audit-demo/plugins/demo/requirements.txt
requests==2.31.0 ; python_version >= "3.8"
flask>=3.0
# comment
-r other.txt
$ HERMES_HOME=/tmp/audit-demo hermes security audit --skip-venv
Found 28 known vulnerability finding(s) across 3 component(s):
[mcp:lodash-fixture]
HIGH lodash==4.17.15 GHSA-35jh-r3h4-6jhm
Command Injection in lodash
fixed in: 4.17.21
...
[mcp:urllib-fixture]
HIGH urllib3==2.0.0 GHSA-2xpw-w6gg-jr37
urllib3 streaming API improperly handles highly compressed data
fixed in: 2.6.0
...
[plugin:demo]
MODERATE requests==2.31.0 GHSA-9hjg-9r4m-mvj7
Requests vulnerable to .netrc credentials leak via malicious URLs
fixed in: 2.32.4
Three components found from five declarations: the pinned npm package, the pinned PyPI package, and the plugin’s exact pin. The unpinned npx server and the URL server produced nothing, and the loose flask>=3.0 line plus the -r other.txt include were dropped. Note the exit code on that run: 0, even with HIGH rows, because the default tier is critical. Add the venv back in and the same fixture grows the total to 241 components and 31 findings.
The fixture is also the fastest way to sanity check a config change. Pin an MCP server to a version you know is dirty, run the audit with --skip-venv, then unpin it and confirm the component count drops. That is a real read on whether the version pin you just added is in the audit path or just decoration.
What the counts do and do not mean
Twenty-nine findings are not twenty-nine problems. Grouping by (package, version, fixed-version list), 12 of the 13 PYSEC rows sit next to a GHSA row for the same package, the same version, and the same fix. The same urllib3 cookie-header bug appears once as GHSA-v845-jxx5-vc9f (MODERATE) and once as PYSEC-2023-192 (UNKNOWN, empty summary). Read the JSON and you can see the pairing directly:
$ hermes security audit --json | python3 -c "import json,sys,collections; d=json.load(sys.stdin); print(d['total_components_scanned'], d['finding_count']); print(collections.Counter(f['severity'] for f in d['findings']))"
240 29
Counter({'UNKNOWN': 13, 'MODERATE': 9, 'HIGH': 5, 'LOW': 2})
Two more properties worth knowing before you paste this into a dashboard. First, the human renderer truncates each advisory summary at 100 characters, appending an ellipsis, while --json carries the full string. The pydantic-settings row in my scan printed as 97 characters plus ... and came out of the JSON as 144 characters, including the words “bypassing secrets_dir_max_size” that the terminal view never showed. Second, the detail fetch degrades quietly: _fetch_one catches URL errors per advisory and returns a Vulnerability with UNKNOWN severity and no summary, so a flaky request turns a real HIGH into an UNKNOWN row that cannot trip a tier. Only the batch query raises.
The three MCP servers nothing audits
Back to the profile, because this is the part that generalizes. Here is what it actually declares:
mcp_servers:
xactions:
command: npx
args: ["-y", "xactions-mcp"] # no version pin
firecrawl:
command: npx
args: ["-y", "firecrawl-mcp"] # no version pin
xactions-hosted:
url: "https://xactions.app/mcp" # URL transport, no process at all
Every one of those is outside the scan for a different reason, and two of them are the exact npx -y <package> form that most MCP setup instructions tell you to paste. That install is real supply chain: npx resolves the name at launch, which means the code your agent executes is whatever the registry serves that day. hermes security audit will not tell you about it, because the audit subject is a version and there is no version here. Pinning costs one @x.y.z and is the only way the row can ever appear:
mcp_servers:
firecrawl:
command: npx
args: ["-y", "firecrawl-mcp@1.2.3"] # the pin is the whole difference
That is also the shape hermes mcp add firecrawl --command npx --args -y firecrawl-mcp@1.2.3 writes into config.yaml (its help text notes --args must come last), and it is the only shape the audit reads. Add it and hermes security audit --skip-venv grows a [mcp:firecrawl] section; leave it unpinned and the server stays invisible no matter how many advisories the registry has on it.
The same logic explains why the plugin surface matters more than it looks. A plugin typically does not install into the Hermes venv, so its Python dependencies are invisible to the venv scan. Its requirements.txt is the only place those pins exist, and only exact == pins are read. A plugin that ships aiohttp>=3 is audited as zero components, and the “No components discovered” line will be your only warning.
Wire it into a gate without lying to yourself
The failure mode that bites in CI is the exit code, not the vulnerability. A dead network and a vulnerable venv both produce a non-zero exit, and a gate that treats all non-zero as fail will one day page you about [Errno 111] Connection refused. I forced that case with a dead proxy:
$ HTTPS_PROXY=http://127.0.0.1:9 HERMES_HOME=/tmp/audit-demo3 hermes security audit --skip-venv; echo "exit=$?"
audit failed: OSV batch query failed: <urlopen error [Errno 111] Connection refused>
exit=2
Exit 2 is “the audit could not run”, a fact about your network, not about your packages. Here is a wrapper that keeps the two apart and dumps the rows you actually care about:
#!/usr/bin/env bash
set -uo pipefail
rc=0
report=$(hermes security audit --json --fail-on high) || rc=$?
case "$rc" in
0) echo "gate: PASS (no finding at/above high)" ;;
2) echo "gate: ERROR (audit could not run)" >&2; exit 2 ;;
*) echo "gate: FAIL (findings at/above high)"
printf '%s' "$report" | python3 -c 'import json,sys; d=json.load(sys.stdin); [print(" ",f["severity"],f["package"],f["version"],f["vuln_id"]) for f in d["findings"] if f["severity"] in ("HIGH","CRITICAL")]'
exit 1 ;;
esac
Run against this box, unmodified:
$ ./gate.sh; echo "wrapper_exit=$?"
gate: FAIL (findings at/above high)
HIGH httpcore2 2.7.0 GHSA-7mj9-2mp8-4m2p
HIGH httpx2 2.7.0 GHSA-7mj9-2mp8-4m2p
HIGH httpx2 2.7.0 GHSA-8xx6-hgc6-gc2m
HIGH langchain-core 0.3.86 GHSA-qh6h-p6c9-ff54
HIGH tornado 6.5.7 GHSA-mpf4-983q-p7j4
wrapper_exit=1
The wrapper earns its keep in three specific choices. --fail-on high instead of the default critical means a HIGH cannot slip through under a green build. Separating 2 from 1 keeps an OSV outage from masquerading as a finding. And because the JSON comes from the run that produced the exit code, the rows printed are the rows the gate saw, not a second scan against a registry that may have changed in between.
One more trap for anyone scheduling this: do not run it daily. The PR is explicit that the tool is on-demand by design, because daily scans become noise you train yourself to ignore. Run it when you are paying attention, or when a dependency changed, not on a timer.
How to verify it is working at all
Three checks, each one falsifiable, all of them cheap:
- The scan found something.
hermes security audit --json | python3 -c "import json,sys; print(json.load(sys.stdin)['total_components_scanned'])"printed240for the Hermes venv. A zero there means the interpreter’s import path is empty, which is a broken install, not a clean one. - The narrow surfaces work. Build the fixture config above, run with
HERMES_HOMEpointed at it and--skip-venv, and confirm the component count matches your pinned declarations exactly. Then unpin one and watch the count drop. - The gate responds.
--fail-on lowand--fail-on criticalagainst the same environment returned 1 and 0 for me. If every tier returns the same code, you are looking at exit 2 (an error) or an empty scan, and the JSON will tell you which.
There is a sibling feature worth knowing about while you are here, because it covers the case the audit cannot: a package that is already compromised rather than merely vulnerable. hermes_cli/security_advisories.py holds a catalog of known-bad versions and runs a version check at startup, silent unless one is installed, and hermes doctor --ack <advisory-id> acknowledges a specific advisory so its startup banner stops nagging. On this box that startup path is live and it logs, which I confirmed in the agent log: the SSH check fired at 14:50:24 today with [security 1/1] SSH password authentication is ENABLED. That audit checks posture rather than packages: running as root, password auth on sshd, a container whose HERMES_HOME is not on a persistent volume, and a network-accessible API server with no API_SERVER_KEY. Same family, different question.
The audit is only as wide as your pins
The uncomfortable truth about this command is that its output is a function of your discipline, not of your risk. Everything that showed up in that scan came from the one surface that is always enumerable, the venv Hermes installs into. The three MCP servers stayed out because nobody wrote a version number, and a plugin’s whole dependency tree disappears the same way the moment it uses >= instead of ==. The scan is honest about this, and the honesty is buried in “No components discovered” and in UNKNOWN rows that can never trip a gate.
Treat it as a verifier, not a scanner. Its value comes from the boundary you hand it: pin your MCP servers, pin your plugin requirements, and the audit becomes a real check on the code your agent executes. Leave them unpinned and you have a scanner pointed at the one surface that was already enumerable, printing rows whose severity buckets decide whether your gate ever fires.
Sources
- Hermes Agent documentation, CLI Commands Reference (
hermes security audit, flags, and exit codes): https://hermes-agent.nousresearch.com/docs/reference/cli-commands#hermes-security - Hermes Agent documentation, security guide: https://hermes-agent.nousresearch.com/docs/user-guide/security
feat(security): on-demand supply-chain audit via OSV.dev, PR #31460 (NousResearch/hermes-agent), including the “on-demand, not daily” and stdlib-only rationale and the exit-code contract: https://github.com/NousResearch/hermes-agent/pull/31460- OSV.dev API (querybatch and vulns endpoints): https://google.github.io/osv.dev/api/
- Source in
NousResearch/hermes-agent@ 64ea66b0:hermes_cli/security_audit.py(discovery,SEVERITY_ORDER, exit codes, renderers),hermes_cli/security_audit_startup.py(startup posture checks),hermes_cli/security_advisories.py(advisory catalog): https://github.com/NousResearch/hermes-agent/tree/64ea66b03d44ead9ffea48161132e5deca5d255a - Live command output on the authoring box, 2026-09-18:
hermes security audit(human and--json), each--fail-ontier,--skip-venvwith and without a syntheticHERMES_HOME, a dead-proxy run for the exit-2 path, and the CI wrapper above. Thelodash/urllib3/requests/setuptoolsfindings came from a syntheticHERMES_HOMEfixture created for this post, not from installed production dependencies.