Hermes Agent Deep Cuts: The Process Registry Is the Part You Never See
Spawn a bounded background job in Hermes with background=true and nothing else, then keep working. The job runs, exits clean, and you learn nothing. That is not a bug. It is the design, and it is the single most expensive mistake you can make with the process tool, because the completion was captured all along. Every byte of its output sits in a JSON receipt in logs/process-results/, redacted, time-stamped, and reachable by the same session_id you were handed. The hard part is not getting the process to run. It is knowing how to get its result back, and that is where everyone who copies the happy path gets burned.
I reproduced the whole surface live this session, from the spawn to the on-disk receipt, so every snippet below ran. The gap between “background=true prints a session_id” and “the registry makes one background job survive a crash, inherit a different attacker’s kill semantics, and refuse to deliver its own completion to a child” is the entire post.
The mechanism: one thread, a ring buffer, and a shelf
The file that owns all of this is tools/process_registry.py, 2,211 lines of “in-memory registry for background processes spawned via terminal(background=true)”. Its own docstring names the three things it does: print status, poll, wait, kill; keep a rolling output buffer; and park results on disk. The constants at the top define the operating budget:
MAX_OUTPUT_CHARS = 200_000 # rolling output buffer
FINISHED_TTL_SECONDS = 1800 # keep finished processes 30 minutes
MAX_PROCESSES = 64 # max tracked processes (LRU pruning)
The first line is the one that matters for how results come back. A background job does not print to a terminal you are watching. subprocess.Popen runs with stdout=PIPE, stderr=STDOUT, stdin=DEVNULL, and a single _reader_loop thread drains that pipe into a per-session in-memory buffer capped at 200,000 characters. When the buffer would overflow, the oldest characters are dropped, so a talkative process stays retrievable but early output can vanish. That is why on a long build you poll early: the tail is always there, the head is not.
On exit the session is moved to a finished shelf in memory with a 30-minute soft TTL, and the result is written to disk under logs/process-results/<session_id>.json in the active profile’s Hermes home, so it survives a restart. The retention contract, from the Background Process Management docs: the newest 64 completed results are kept for up to 7 days after completion, and each receipt holds at most the rolling 200,000-character tail, with secret redaction always applied even when live redaction is disabled. There is also a crash-recovery checkpoint at <profiles>/processes.json, and a kill_started_since path in the gateway that reaps any process launched after a baseline when a session is abandoned, so nothing your session spawned is left running after a reset.
This session’s own receipt tells the story. After I ran a five-tick loop in the background and polled it to exit, this file existed on disk:
{
"id": "proc_f72cc7ff89ad",
"command": "for i in 1 2 3 4 5; do echo \"tick $i\"; sleep 1; done; echo \"JOB-FINISHED rc=0\"",
"task_id": "default",
"owner_task_id": "cron:5032b7d71ac3:e106741e2b2545e98d10b414d5d86ddc",
"exit_code": 0,
"completion_reason": "exited",
"termination_source": "",
"output": ".autocomplete__key-bindings:39: terminfo[kcbt]: parameter not set\ntick 1\ntick 2\ntick 3\ntick 4\ntick 5\nJOB-FINISHED rc=0\n"
}
The owner_task_id field is the lock. Results are retrievable by session_id only from the conversation that launched the process or its compressed continuation. Unrelated conversations cannot read them “even with an exact process handle”, per the docs. That is an ownership boundary, not a cache key: a receipt is nobody’s business except the session that owns it.
The advanced surface: poll, wait, log, write, submit, close, kill, handoff
process_manage exposes nine verbs. The two that trip people are write versus submit and wait’s timeout. The schema is blunt about the first: “submit appends Enter; use it to answer prompts; write sends raw bytes, no newline.” A lone newline sent through write on a Windows PTY is not Enter. That is a real trap from issue #95681 that the schema authors called out in the description.
I drove a PTY job live to prove the stdin loop works. Spawn bash -c 'echo READY; read -r line; echo "YOU SAID: $line"; exit 0' with pty=true, then submit "hello from pty":
READY
hello from pty
YOU SAID: hello from pty
wait behaves differently from poll. poll returns a snapshot; wait blocks until exit or timeout, and on timeout it returns partial output with a note, not an error. I ran wait with a 3-second window against a 10-second job:
{"status":"timeout","command":"echo start; sleep 10; echo end",
"output":".autocomplete__key-bindings:39: terminfo[kcbt]: parameter not set\nstart\n",
"process_running":true,
"timeout_note":"[elided here: states the wait is not an error, reports uptime 9s, and suggests notify_on_complete for next time]"}
The timeout returns whatever the reader loop had drained so far. process_running: true tells you to come back. The same handle, polled after, returned the full output with exit code 0.
kill has the most interesting ordering and I saw it fire. I killed a sleep 300 process and the snapshot contained a line the process should never have printed, never reached, because _terminate_host_pid uses psutil to SIGTERM children before the parent. The shell got the signal, ran its pending echo, and only then died. A kill in this registry is not “send signal to pid”. It is a tree teardown, children first, then the parent, with a SIGKILL escalation after a configurable grace period. The exit code is recorded as -15 (SIGTERM), completion_reason is killed, and termination_source is process.kill.
The last verb, handoff, is the one that saves a subagent’s work. A process you start inside a delegated child is killed when the child finishes, and its completion notice never reaches the parent. To prevent the loss you call process_manage(action='handoff', session_id=..., data='<purpose>'). That flips ProcessSession.owner_task_id to the parent under the registry lock, capped at three handoffs per child per the code.
The gotcha: silent background, and the oneshot downgrade that cancels notify
Two failure modes matter for anyone who runs Hermes daily. The first is the silent-background footgun. A bounded task launched with background=true and no notify_on_complete runs to completion and tells no one. The tool returns a hint field on that exact code path telling you so, and the background tool’s own header says “Almost always pair with notify_on_complete=true”. For a bounded job, a build, a test suite, a deploy, you almost certainly wanted the notification. For a real daemon, silence is correct and the hint is one cheap ignored read. The schema ships a second, narrower hint for gh pr checks | jq shaped CI polls, which look like status tracking but quietly never print JSON.
The second failure mode is the odd one, and it is easy to hit wrong. In a one-shot session, a hermes -z, a cron job, a Kanban worker, or a stateless HTTP endpoint, notify=true does not deliver a notification when the job exits, because there is no turn left to accept one. The spawn returns a field you will not see on the happy path:
{"output":"Background process started","session_id":"proc_97d0fc63232d","pid":3582068,
"notify_on_complete":false,
"notify_unsupported":"notify_on_complete / watch_patterns are not available in this session: it cannot receive an async completion after the turn ends (a one-shot runner such as `hermes -z`, a cron job, a Kanban worker, or a stateless HTTP endpoint). The process is running in the background; retrieve its result with process action poll or wait. [full text reproduced verbatim in this session's tool output]"}
That is not a bug and not a silent drop. _apply_async_support in tools/terminal_tool_background.py detects the session cannot route a completion back, stamps notify_on_complete: false, and tells you to poll instead. In a cron job you have to poll, or the work is lost the moment the turn ends.
The most surprising demo, though, is process(action='list') in this same session. I started a 20-second job, confirmed it was mid-flight, then listed:
{"processes":[]}
Empty, with a process running under my own task id. list is scoped to the current task or conversation and is not a global ps. If you are in a oneshot or across a boundary that changes task identity, list shows you nothing while the work is alive, and you are back to tracking session_id yourself. The docs are honest that list includes retained results for the current task or conversation, but they do not lead with the part that matters: retain the handles you spawn, and do not trust list to be your source of truth in a finite session.
How to verify it is working
The most direct check needs no polling loop. Spawn a bounded job, wait for it to exit, then read the receipt off disk:
ls -la <profile-home>/logs/process-results/proc_<id>.json
python3 -c "import json;print(json.load(open('<profile-home>/logs/process-results/proc_<id>.json'))['exit_code'])"
The exit code, the completion reason, and the full output tail are all there, and the file is mode 600. A second verification is rehydration across a boundary: in a normal interactive session a background build that finishes while you are on another tab is woken by notify_on_complete=true, and the completion arrives as a new turn. In a oneshot session, that same flag comes back as notify_unsupported, and the only correct move is process(action='wait'|'poll') before the turn ends. If you see the hint on a bounded job, you are one flag short. If you see notify_unsupported, you are in a session that cannot take the notification at all, and polling is the whole answer.
The registry is the unglamorous infrastructure that makes a CLI feel like it can actually run work for you. Most people never think about it, because a foreground terminal call just returns when it returns. The moment you background anything, though, you are buying into a ring buffer, an ownership lock, a oneshot downgrade, and a tree-kill you did not ask for, and the only way to get your result back is to speak its language. That is the whole difference between a command runner and a process system, and it is the difference between an agent that silently loses a six-minute green CI and one that tells you when it is done.
Sources
- Built-in tools reference (background process management, terminal backends): https://hermes-agent.nousresearch.com/docs/user-guide/features/tools
- Process registry, reader loop, kill ordering, retention constants: https://github.com/NousResearch/hermes-agent/blob/main/tools/process_registry.py
- Background spawn, silent hint, oneshot async downgrade, subagent note: https://github.com/NousResearch/hermes-agent/blob/main/tools/terminal_tool_background.py
process_manageschema, write-vs-submit note: https://github.com/NousResearch/hermes-agent/blob/main/tools/process_registry.py- Silent-background warning behavior: https://github.com/NousResearch/hermes-agent/pull/31289
All outputs shown above were produced in this session on 2026-09-12 against the hermes-agent tree at commit 1e7d29a081 (2026-09-10), version 0.21.1. Receipts were inspected from <profile-home>/logs/process-results/. Everything shown ran live; no result was reconstructed from memory.