Hermes Agent Deep Cuts: A Handoff Is a Capsule, Not a Transcript
Part of the Hermes Agent: Deep Cuts series

Hermes Agent Deep Cuts: A Handoff Is a Capsule, Not a Transcript

$ python3 demo_build.py
  [runner] plugin shelled out to: /home/dazeb/.hermes/hermes-agent/venv/bin/hermes
           sessions export --session-id demo-session-0001 --format jsonl -
  [writer] prompt chars=1683
== build result ==
{"status": "ok", "lane": "demo-lane-not-a-real-conversation", "messages": 6,
 "jev": "not_used", "chars": 1666}

== capsule on disk ==
## Working on
The writer model did not return a usable handoff, so this is the filtered transcript
instead. Lines marked KEEP VERBATIM are the ones that matter.

## State
user: The nightly mirror job is failing on the second host. I want it to keep going
when one host is unreachable and report which one it skipped.

Status ok, and a capsule-shaped file on disk, with no summary inside it. The transcript in that run is fabricated for the demo; the code path, the writer model, the prompt and the file writes are the real ones.

That is the design working, not a bug. The handoff plugin on this box writes a capsule only if the writer model returns something capsule-shaped, and the writer is the profile’s auxiliary compression model, a different model from the one you are talking to. In the run above, every auxiliary candidate answered 401 (stale/unrefreshable credential) and no fallback_chain was declared, so the validator rejected the empty reply and the plugin fell back to the transcript, which its own source calls “worse to read but loses nothing”.

The valuable part is that the fallback kept the file, kept the message count, and kept the Recovery section. What it did not do is tell you the summary is missing anywhere except inside the file. A cron watchdog that greps for the marker file sees a handoff. An operator who opens it sees a transcript wearing a capsule’s clothing.

What a handoff actually is here

“Handoff” in Hermes means session continuity, and the mechanism is deliberately small: close a session on purpose, carry forward a short note, and start the next session light instead of dragging the transcript forever.

The hermes-handoff plugin (version 0.4.1 in this profile, enabled in plugins.enabled at line 458 of the profile’s config.yaml) claims three seams:

SeamRegistered asJob
pre_gateway_dispatchhookthe bare word handoff starts a capsule for this conversation
/wrapupslash commandthe same build, on demand
pre_llm_callhookthe next session’s first turn receives the capsule as context

The gateway contract for the first seam is narrow and worth knowing before you write anything on top of it: the hook is called with event=, gateway= and session_store= and nothing else, it fires once per inbound MessageEvent before auth and pairing, and a returned {"action": "skip"} drops the message. The plugin returns None on every path, because skipping would mean you type handoff and get silence, and on a host that also rotates its session on that word a skip would pre-empt the half that already works.

The build path is where the interesting failures live. build() exports the session through the CLI rather than the session store:

hermes sessions export --session-id <id> --format jsonl -

The comment in the source explains the choice: the CLI is a contract that survives upgrades, the store’s schema is not. That command returns one JSON row per session with every message inside it, and on this box a real export of a five-message session returned a row carrying roles ['user', 'assistant', 'tool', 'assistant', 'session_meta']. The plugin keeps only user and assistant text, which means a value that appeared in a tool result is not in the capsule unless a turn mentioned it. Add role_filter="user,assistant,tool" when you search that session later, which is exactly what the Recovery block tells you to do.

Then the writer gets the whole dialogue, up to 300,000 characters, with a prompt asking for five headings in this order: Working on, State, Decisions, Pointers, Next. The size limits are the design:

  • 1,200 words for the capsule (400 in a confidential lane)
  • 12,000 characters hard cap on what gets stored
  • 240 turns maximum put in front of the Jev pre-pass, judged 40 turns per request
  • 30 second timeout on the export, because “an 800-message session exports in about half a second”

Nothing in that list is a guess about what makes a good summary. It is what the author measured. The scorecard in the plugin repository, over 104 recall questions on seven real sessions:

Versionrecall alonewith one search
Jev digest of the last 24,000 characters, 400 words37.5%68.3%
whole dialogue, 1,200 words, Recovery section58.7%75.0%

Question by question the new configuration won 26 and lost 4. A capsule with no way back to the session answered 59% alone and 75% with one search, while no capsule plus one search answered 57%, which is why the plugin stopped trying to make the note self-sufficient and started appending a pointer to the transcript instead.

A lane is a conversation, not a session

Capsules are keyed per lane, not per session id. The key is built from platform, chat id and thread id, and the reasoning is in the source: one capsule per conversation, because the session id changes every time the conversation restarts and the capsule is supposed to survive exactly that.

Long conversation ids get a hash suffix rather than a truncation, because two Microsoft Teams conversations run to 131 characters and a plain cut would make two customers share one lane. I ran the function directly against two ids that agree for their first 120 characters:

139-char id A: 19:abc-def-1234-guid-less:topic:aaa...aaa-8b67025add22d7bc7516
139-char id B: 19:abc-def-1234-guid-less:topic:bbb...bbb-04d4acf9287a05853625
identical prefix through char 120, keys differ: True

The files that pair up in a lane are <home>/handoffs/handoff-<lane>.md and <home>/handoffs/pending-<lane>.json. The second one is the part that is easy to miss, and it is what makes the capsule arrive exactly once. I ran the consume path against a scratch home with a hand-written capsule:

capsule file: handoff-telegram:1033877751.md | exists: True
pending file: pending-telegram:1033877751.json | exists: True
1) asking as the OLD session (the one being summarised): None
   pending still there? True
2) asking as a NEW session: # Handoff (demo)  ## Working on build the mirror fix ...
   pending still there? False
3) asking again: None

Three rules are packed into that output. The session that was just summarised never receives its own capsule, so a conversation cannot eat its own summary. The marker is deleted on handover, so the capsule is not injected forever. And a marker older than 36 hours is deleted unread, so last week’s stale plan cannot walk into this morning’s session claiming to be context you established.

The injection itself lands in the user message, not the system prompt. pre_llm_call returns {"context": ...} and Hermes appends it to the current turn’s user message, which keeps the system prompt byte-identical across turns so the cache still hits. The plugin wraps the capsule in a short label that says this is history you already established, not a new instruction, and tells the model not to greet you again or re-ask what the note settles. That label is doing real work: it is the difference between a fresh session that continues and a fresh session that introduces itself.

Turning parts off, and the confidential lane

Two settings matter in daily use. Injection can be disabled through the plugin settings subtree, which the plugin reads as plugins.entries.<plugin-id>.settings.<key>:

plugins:
  entries:
    hermes-handoff:
      settings:
        inject: false   # still writes capsules, stops injecting them

Capsule writing keeps working, which is what you want while you audit what the writer is producing. The Jev pre-pass is a separate switch, HANDOFF_JEV=1, and it is off by default for a measured reason rather than a taste reason: capsules written from the Jev-digested transcript recalled less than capsules written from the plain text of the same size (4 questions won, 15 lost), while Jev’s marks themselves beat marks assigned by recency 11 to 4. The judgement is good. The digest built around it cost more than the judgement earned, so the default is the path that measured best and sends nothing anywhere.

The confidential lane is the one place the feature refuses to be helpful. Switch it on with HANDOFF_CONFIDENTIAL=1 or a file named CONFIDENTIAL in the profile’s handoffs directory, and the capsule becomes a 400-word breadcrumb with no Recovery section, because a pointer into a transcript full of customer detail is what that contract exists to prevent. The writer sees a narrower read and a prompt that forbids names, addresses, file names and account numbers, and a mechanical scrub runs over the output even when the writer behaved. Ask for confidentiality on a host that cannot supply the prompt builder and you get nothing at all:

confidential with no scrubber: {'status': 'confidential_unsupported', 'reason': 'no_scrub'}

That refusal is correct and it is also a sharp edge. A thin handoff is a bad morning; a transcript on disk is a broken promise, and the plugin picks the bad morning.

The gotcha: three features share the word

Type /handoff expecting a capsule and you get something else entirely. Hermes ships a built-in /handoff <platform> that transfers your CLI session to a messaging platform, and it is a state machine, not a summariser:

None  ->  "pending"  ->  "running"  ->  ("completed" | "failed")

The CLI writes pending and handoff_platform into your session row, then block-polls. If the gateway does not claim it within 60 seconds the CLI compare-and-swaps the row to failed so you can retry. Once claimed, the CLI waits up to 900 seconds with a heartbeat every 30 seconds, and on that timeout it deliberately leaves the row alone: failing it there was a split-brain bug, because the gateway still owned the transfer. The session row on this box carries handoff_state, handoff_platform and handoff_error columns among its 59, and across 267 sessions and 42 still-open ones, exactly zero have ever used them. That is the shape of a feature people do not know is there.

So the plugin could not register /handoff. Its source says the registration was refused on every load and the command never existed, and that a hyphen was also out because Telegram rejects one in a command name. The command is /wrapup. Three things, one word:

  • bare handoff in a chat, captured by the dispatch hook, closes the stretch and writes a capsule
  • /handoff telegram moves the live CLI session to Telegram and exits the CLI when the transfer completes
  • /wrapup is the capsule from a command rather than from a keyword

The trigger is strict on purpose. is_trigger strips whitespace, full stops and exclamation marks, then requires an exact match against handoff, hand off or hand-off:

'handoff'                        -> True
'Handoff.'                       -> True
'hand off'                       -> True
'handoff now please'             -> False
'can you do a handoff for me?'   -> False
'/handoff telegram'              -> False

A sentence that mentions a handoff is a normal message. That matters in shared chats, and so does the next part: pre_gateway_dispatch runs before the gateway’s own authorization, so a plugin that acts on that hook is spending the owner’s model calls for whoever can post in the room. This one re-checks authorization itself and refuses on an explicit no, and it ignores proactive plugin events that carry text nobody typed. If you build on this hook, both checks are yours to make.

The trim you never see

Capsules get cut from the middle, not the end. A plain truncation at the 12,000-character cap removes the last sections first, which are Pointers and Next, the two sections the next session cannot reconstruct. I ran the trim function on a 12,564-character capsule with all five headings present:

input chars=12564  cap=12000
output chars=12000
keeps '## Pointers'? True   keeps '## Next'? True
keeps '## Working on'? True   has trim marker? True

The cut lands in the middle of the note and leaves the tail intact. That is why a long handoff can lose its State section and still be worth reading, and why you should not assume a capsule that looks thin near the top was written badly. Check whether the trim marker is in it before you blame the writer.

The other silent path is the one from the opening. When the writer fails or returns something that is not capsule-shaped, valid catches it, and the stored text says so in the first section:

## Working on
The writer model did not return a usable handoff, so this is the filtered transcript
instead. Lines marked KEEP VERBATIM are the ones that matter.

So the verification for a real capsule is one grep for five headings, not a check for the file’s existence:

$ grep -c '^## ' <home>/handoffs/handoff-<lane>.md
6        # five capsule headings plus a Recovery section, all present

And if the auxiliary model is the thing that is broken, fix that first. The capsule writer has no credential of its own, on purpose: it borrows whatever the profile already uses for compression. If your aux provider is stale and no fallback chain is declared, every handoff writes a transcript, reports ok, and the only signal is the first paragraph of the file.

The nightly version, and where it stops

The same repository ships a nightly closer, scripts/nightly-handoff.py, and its guard rails are worth copying even if you never run it. It only considers sessions with at least 6 messages, only conversations active within 3 days, and caps a night at 40 profiles-worth of model calls. It writes the capsule before it closes the session, on the theory that losing the thread is worse than a large context. It takes one capsule per conversation from that conversation’s most recent session and closes older open sessions alongside rather than summarising over the top of them.

It also excludes cron sessions by name. Every cron job’s session lands on one platform-wide lane, so a capsule written from one stale job would be handed to whichever job ran next as context you established.

That exclusion points at the limit of the whole design: a handoff assumes a conversation that will come back. In a gateway, the capsule is built on a worker thread while the person keeps typing, and the next session picks it up on its first turn. In a one-shot scheduled run there is no next session in the same lane, so treat the trigger as a gateway and CLI feature and do not expect a cron job to hand itself off. Inferring more than that would need a test this run did not do.

What to check when you do not trust it

Every claim above is checkable without reading the plugin. The files are in <profile home>/handoffs/. A completed build logs one info line naming the lane, the session and the message count, and a refusal logs a warning with the status instead, so the dispatch path leaves evidence even though it never replies. A pending marker that is still sitting in the directory after the next session’s first turn means the injection did not happen, and the lane key did not match, which usually means the platform or chat id changed shape.

The note is the cheap half of a handoff, and the measurements say so: a capsule alone answered 59% of the recall exam, one search of the old session took it to 75%, and no capsule with one search managed 57%. What earns its place at the bottom of the file is the Recovery block, the one no model writes: the session id, the message count, and the two session_search calls that work. There is a trap in that footer worth knowing before you use it, because the obvious call is wrong. Passing query together with session_id does not search that session, it reads it from the top and ignores the query, since session_id wins the dispatch at tools/session_search_tool.py line 598 and only session_id plus around_message_id scrolls. Discovery and recall are two different calls on purpose:

session_search(query="2 to 5 keywords")                         -> session_id + match_message_id
session_search(session_id="<id>", around_message_id=<match>)    -> the messages around the hit

A session that never ends is not memory. It is a context bill that grows every turn, and 2,255 messages in one conversation is the kind of thing nobody chooses and everybody ends up with. A capsule is twelve kilobytes, one file, one injection, and a way back to the transcript it came from. The summary’s quality is not what makes it work. The payoff is that the session can end at all.

Sources

Keep reading