Hermes Agent Deep Cuts: Read the Restart Worklist Before You Update
$ hermes update --plan
Update plan:
Install: git (v0.21.5 @ 6ec05205)
Profiles: default, blogposter, developer, finances, gemma, growers-postbot, marketing, politics-blog, redteam, video
Running services to restart (3):
The full output on this machine names three different restart owners. The gateway is systemd-managed. The dashboard was launched manually. Desktop owns the serve backend. A one-line “update available” check could not tell me any of that.
If you run Hermes across profiles or launch a dashboard outside the gateway service, this is the useful command to run before touching the checkout:
hermes update --plan
It is a read-only preview. It does not apply an update or restart those processes. The same inventory is captured into the receipt during a real update, where it becomes a worklist the updater reconciles against its restart actions.
The plan is an inventory, not a guess
The CLI calls collect_runtime_inventory(). That builds an UpdatePlan with the install method, current code identity, profile names, and runtime records. Each runtime record includes the service kind, profile, PID, supervisor, running code version or SHA when available, and a restart-mechanism identifier.
The collector gathers live gateway processes and then adds runtimes found in Hermes’ service ledger. That second pass matters. The agent should not infer what a PID is from its name alone, or assume one gateway per installation. Profiles share the code checkout, while processes can be managed by systemd, launched directly, or owned by the Desktop app. Those cases need different restart handling.
The printed list is therefore operationally specific. In my output, the manually launched dashboard is scheduled to stop before the code swap and relaunch with its saved launch arguments. The serve backend is left to Desktop. Those are not two spellings of the same restart. If Desktop is closed or a manually launched serve process is outside the expected inventory, the preview cannot magically restart it for you.
The implementation is in update_inventory.py. The CLI handles --plan before the mutating update path in main.py. During an actual run, update_cmd.py records the pre-update plan, and update_cmd_fleet.py checks the restart outcome against it.
Use it as a preflight, then choose the recovery point
For a normal source install, I would inspect the worklist first, then check whether the selected channel has an update:
hermes update --plan
hermes update --check
Those answer different questions. --plan describes the local install and the running services that need handling. --check compares the checkout with its update target. Neither is the update itself.
For a production profile or a shared host, decide how much state you need to recover before running the apply step. The default is a quick snapshot. A full backup also zips HERMES_HOME, including configuration, credentials, sessions, and skills, subject to the documented backup exclusions. You can request that for one update:
hermes update --backup
Or set it as the ongoing policy in ~/.hermes/config.yaml:
updates:
pre_update_backup: full
quick, full, and off are the accepted modes. --no-backup opts out for one run, and it wins over --backup if both flags are supplied. The quick snapshot skips individual files larger than 1 GiB. Full archives can take minutes when the Hermes home contains large data, so do not turn it on blindly for a large installation.
One more flag matters when the update is itself running inside the gateway that it would normally restart:
hermes update --no-gateway-restart
That updates the code and dependencies while deferring the fleet restart, so the updater does not kill its own process. Hermes keeps the restart obligation for a later catch-up. Use this only when a separate, deliberate restart will follow. Otherwise you have changed the checkout and left the live gateway on the old code.
The gotcha: a green plan is not a green update
--plan tells you what Hermes can currently identify and how it intends to treat each runtime. It does not test the download, dependency preparation, backup, code swap, or restart. It cannot prove that a manually owned process will come back healthy after an update.
There is a second boundary in the update path: a service manager reporting “active” is not enough. After a restart, Hermes compares the runtime’s recorded code identity with the new checkout and reconciles the planned services with the actions actually taken. If a gateway may still be running stale code, or the restart phase is incomplete, the updater fails rather than quietly calling that fleet current. The plan matters because it gives that check something concrete to compare against.
The backup is not a transaction either. A quick-snapshot failure is reported, but by design it does not block the code update. If recovery is a hard requirement, inspect the output for the snapshot result and use --backup for the full archive. Do not treat a successful process exit as proof that a usable recovery point exists.
Verify the inventory before and after
Before applying an update, run the plan and look for a service whose restart owner is outside the command you are about to run. Manual serve and dashboard processes are the usual place to slow down. On package-managed or image-managed installs, the plan can also name the external update mechanism instead of suggesting an in-place code swap.
After the update, run the plan again:
hermes update --plan
Compare the running code SHA or version with the installed code identity and confirm that the expected services are present under the right profile and supervisor. For an update launched from inside the gateway, confirm that the deferred restart has happened before treating the fleet as current. A second plan is an inventory check, not an end-to-end agent health test. Send a test message through the gateway or use the health check appropriate to the service you operate.
The distinction is easy to miss: the plan is not a fancy way to say “update available.” It tells you which running processes will be affected, who owns their restart, and what evidence the updater will later use to decide whether the fleet caught up. If the host has more than one Hermes process, that is the part worth reading before you pull code.
Sources
- Hermes Agent: Updating and uninstalling, including fleet preview, update phases, pre-update backups, and restart behavior.
- Hermes Agent source:
update_inventory.py. - Hermes Agent source:
main.py. - Hermes Agent source:
update_cmd.py. - Hermes Agent source:
update_cmd_fleet.py. - Hermes Agent source:
update_cmd_maint.py.