Hermes Agent Deep Cuts: hermes pause Is Not a Process Kill, It Is a Sentinel
Part of the Hermes Agent: Deep Cuts series

Hermes Agent Deep Cuts: hermes pause Is Not a Process Kill, It Is a Sentinel

Your cron fan-out starts doubling every tick and you are looking at maybe two hundred jobs a day instead of the four you scheduled. The natural reflex is to kill something. kill -9 the gateway, reboot the box, chase the PID. Stop. Run this instead:

$ hermes pause --reason "runaway fan-out, freeze dispatch now"
⏸️ Hermes paused — reason: runaway fan-out, freeze dispatch now
 sentinel: /home/dazeb/.hermes/ESTOP
 Cron dispatch, kanban dispatch, and new gateway turns are on hold.
 In-flight work keeps running. Run `hermes resume` to lift the pause.

That is the whole emergency stop. One command wrote one small JSON file and every scheduler on the box froze new work at the next tick. No process died. Nothing in flight was lost. This is the control-loop lever most operators never touch because they assume a pause means something expensive like draining a queue or stopping a service. It means one os.stat per tick. And because it is a file, not a signal, it works across processes, across profiles, and even after a reboot you left in place.

The entire mechanism is one file and two stat calls

The heart is agent/estop.py, a module so small it fits in your head. hermes pause calls engage(reason) which writes a sentinel at $HERMES_HOME/ESTOP:

{
  "engaged_at": "2026-09-18T01:48:18.093824+00:00",
  "reason": "deploy window, do not dispatch new crons"
}

The check side is is_engaged(), and it is deliberately about as cheap as a check can be: one or two uncached os.stat calls, one for the active profile home and one for the fleet root when they differ. There is no central registry, no IPC, no daemon to ask. Any process that wants to know whether work is allowed just stats a filepath. Cron, kanban, and the gateway all call into the same check_paused(component, logger) helper, which returns True while any candidate sentinel exists.

The fail-safe direction is the interesting part. If the stat itself raises an OSError, the check returns True, meaning paused. That is not a bug, it is the point. A filesystem that has become unreadable should stop accepting new autonomous work, not keep firing jobs you can no longer observe. Same spirit applies to the sentinel body. I tested a corrupt non-empty sentinel, plain text not-json{, and the result was still engaged:

$ python3 -c "from agent.estop import is_engaged, get_state, paused_reply; print(is_engaged()); print(get_state()); print(paused_reply())"
True
{'reason': None, 'engaged_at': None}
⏸️ Hermes is paused. New work is on hold; run `hermes resume` to pick things back up.

An empty file is the documented equivalent of a bare touch ~/.hermes/ESTOP. Both pause. So the feature does not trust the JSON, it trusts the file’s existence.

Where the gate actually sits

The same two-line gate appears at four separate control points, and they are not the same check, they are four gates that log independently. Reading where they sit tells you what a pause is and is not.

The cron ticker, cron/scheduler_tick.py, checks the gate at the top of _tick_admitted, after it takes the tick lock:

with contextlib.suppress(ImportError):
    from agent.estop import check_paused as _estop_check_paused
    if _estop_check_paused("cron", _sched.logger):
        return 0

Returning before get_due_jobs means no new job is dispatched. The misfire catch-up sweep, in cron/scheduler_provider.py, has its own gate with its own component name, cron-misfire, so it can log separately even though it reuses the same helper. The managed-cron webhook path, gateway/platforms/api_server.py, is the one that returns a real HTTP status instead of silently skipping. It answers the external fire request with a 503 and a Retry-After: 60 header, so the upstream scheduler in the managed-cron service reschedules rather than dropping the job. That is an elegant piece of design: an emergency stop does not lose deliveries, it defers them, and the network protocol lets the caller be a good citizen.

The kanban dispatcher, gateway/kanban_watchers_common.py, is the one people forget. _kanban_dispatch_allowed() returns False while the stop is engaged, also checked every tick before spawning so in-flight workers are never touched.

The gateway turn gate is the richest. In gateway/run_inbound.py, a new external turn is refused with the pause notice, but not every turn. _hm_estop_turn_allowed lets through: recognized slash commands including /pause off (so a messaging-only operator is never locked out of lifting it), replies owned by already-running work, pending dangerous-command approvals, and pending slash confirms. Internal events representing in-flight work completions bypass the gate entirely. The docstring is exact about it: the pause blocks NEW agent turns, never running work or control traffic.

That boundary is the whole thesis. If you pause and then the agent sends a status update about the work it was already doing, that is allowed. If a human DMs a brand new task, it is blocked with a notice to run hermes resume. Slash commands and control survive, new model turns do not. That asymmetry is what makes the lever safe to pull mid-incident.

The gateway has its own lever in band

If you are looking at this from a chat client, you do not need a shell. The gateway registers a /pause slash command with a busy_policy of dispatch, which the test suite checks explicitly, so it works even while an agent session is mid-run. agent/estop.py and gateway/run_busy.py share the same store, so the CLI and the slash command are two frontends to one sentinel.

/pause deploy window
⏸️ Paused (reason: deploy window). New cron/kanban/gateway work is on hold; in-flight work finishes normally. Use `/pause off` to resume.
/pause off
▶️ Resumed — new work is accepted again.

The reason string is optional but it pays for itself fast. It is stored in the sentinel, shown to users who touch a paused gateway, and surfaced in hermes status. Without it, a gateway reply that was supposed to be a quick status check just says paused and you have no memory of why.

How to verify it is actually engaged

hermes status renders a one-line banner in yellow when the stop is set:

⏸️ PAUSED (global emergency stop — reason: deploy window; `hermes resume` to lift)

That banner is driven by the same get_state() as everything else, so if it shows, every gate on the box sees the same sentinel. The truly independent check is the file itself. The reason field is written to the sentinel, so cat the path hermes pause printed and you confirm both that it is set and why.

The gotcha that will actually bite

The failure mode most operators hit is a pause that appears to lift but does not, because is_engaged() checks two candidate paths and resume removes both, but one of them fails to unlink. The candidate list is the active profile home first, then the canonical Hermes root if it is a different directory. The reason for two candidates is fleet semantics: a profile gateway running with HERMES_HOME=~/.hermes/profiles/<name> must still honor an operator-level ~/.hermes/ESTOP. I reproduced the profile case by monkeypatching the root resolver to a throwaway tree and pointing HERMES_HOME at a nested profile dir. With the root sentinel present and only the profile home active, is_engaged() correctly returned True. So a single hermes pause at the root silences an entire fleet of profile gateways.

The trap: a sentinel can also exist where you did not put it. You (or another operator, or a leftover script) might have touched a file at the fleet root, and your resume in the current profile removes the profile copy but the root copy keeps the whole fleet paused. disengage() tries to unlink every visible candidate, but an unlinked root path, from a permission error or a different mount, silently leaves the other in place. The practical rule: if work stays frozen after hermes resume says resumed, check both paths, ~/.hermes/ESTOP and ~/.hermes/profiles/<profile>/ESTOP, not just the one you typed.

The second gotcha requires no file at all. A fresh pause does not stop a job that is already running, and that is a feature. If what you actually need is to abort the current run, pause is the wrong tool; you want per-session controls like /stop, not the global lever. Pulling the global stop and expecting the in-flight job to halt is the classic mismatch, the one that makes people conclude the feature is broken.

Operating the global lever

Keep the essentials in your head:

hero=$(printf 'runaway fan-out: do not dispatch new crons')
hermes pause --reason "$hero"     # freeze new work everywhere
hermes status                     # yellow PAUSED banner confirms
hermes resume                     # lift on the next tick, no restart
# emergency, no CLI available:
touch ~/.hermes/ESTOP             # bare file counts as engaged (fail safe)
rm ~/.hermes/ESTOP                # bare file lifts it, empty body is fine

The bare-file fallback is the part worth internalizing for an automation operator. Because the gate is a file existence check that fails safe, you can pause from any environment that can create a file, no hermes binary on the path, no agent running. A monitoring script, a systemd unit, an SSH one-liner from a jump box can all pull the lever. That is the property of a control loop worth building: a single, observable, reversible artifact that every worker honors, that no one needs a shell session with a running agent to write.

The uncomfortable truth is that most agent control loops lack exactly this. They have approval gates for individual commands but no operator-level circuit breaker that halts autonomous dispatch without killing state. hermes pause is that breaker, and it is nothing more than a file and two stat calls. Once you see the sentinel, the CLI stops being a command and becomes a mechanism you can wire into your own automation: watch your error rate, touch the file, and the fleet stops itself before the blast radius grows.

Sources

Keep reading