Hermes Agent Deep Cuts: /compress Is a Rewrite, and It Fails Closed
$ python3 probe.py # read-only copy of this box's session store
== sessions holding every compacted row ==
id=20260803_082413_6001150a
started=2026-08-03 07:24:13Z end_reason=None compacted=7010 total=7185
id=cron_266e70983203_20260831_110042
started=2026-08-31 10:00:45Z end_reason=cron_complete compacted=6365 total=7994
sessions whose end_reason == 'compression': 0
sessions with a parent_session_id: 0
total message rows: 27474
compacted=1 rows: 13375
7,010 of the 7,185 message rows in the oldest session on this box sit behind a summary. So do 6,365 of 7,994 in a cron session. Across 264 sessions and 27,474 stored rows, not one session has a parent_session_id and not one ended with end_reason = 'compression'.
That combination is the point of the feature. Compaction rewrote 13,375 rows, in place, and the conversations kept the same identity the whole time.
The command that does it is /compress, which most operators meet as a status line in a long chat and never read the help for. What is actually there is the only sanctioned mutation of a transcript in Hermes, and it has a grammar, a protection list, a summarizer model, a persisted failure ladder, and three separate ways of doing nothing while keeping your history intact. The failure modes are more interesting than the happy path, because every one of them chooses conversation over guesswork.
Two compressors, and you only asked for one
Compression runs at two layers. The gateway has a pre-agent pass called session hygiene, fixed at 85% of the model’s context window, that catches sessions which grew between turns. The in-loop compressor inside the agent is the primary one, firing at compression.threshold, which defaults to 0.50.
The numbers for a 200,000-token model at defaults:
context_length = 200,000
threshold_tokens = 200,000 × 0.50 = 100,000
tail_token_budget = 100,000 × 0.20 = 20,000
max_summary_tokens = min(200,000 × 0.05, 12,000) = 10,000
threshold_tokens here is derived from the main model’s window, never the summarizer’s, and compression.threshold_tokens: 256000 caps it so a 1M-token model cannot silently defer compaction to half a million tokens.
ContextCompressor.compress() then runs four phases. Phase one prunes tool results longer than 200 characters that sit outside the protected tail, replacing them with [Old tool output cleared to save context space]. No model call, and on tool-heavy sessions it does a lot of the work before the summarizer is involved at all.
Phase two draws the boundaries. The first protect_first_n messages (default 3) are pinned so your original goal survives every pass, the middle is what gets summarized, and the tail is protected by a token budget that walks backward from the end, falling back to protect_last_n (default 20) if the budget would protect fewer messages. Boundaries are aligned backward past consecutive tool results so a tool call never gets separated from its result.
Phase three sends the middle to the auxiliary compression model with a fixed template: Goal, Constraints and Preferences, Progress split into Done / In Progress / Blocked, Key Decisions, Relevant Files, Next Steps, Critical Context. The summary budget scales at 20% of the compressed content, floored at 2,000 tokens and capped at 12,000.
Phase four assembles head, summary, tail and sanitizes orphaned tool pairs, dropping results whose calls were removed and injecting stub results for calls whose results vanished. On a repeat compaction the previous summary goes back to the model with instructions to update it rather than start over, so a long session accumulates one running picture of its own work instead of a fresh lossy summary every time.
The grammar you never typed
The command line after /compress is parsed by hermes_cli/partial_compress.py and agent/conversation_compression_manual.py. Running the shipped parser against realistic input:
arg after /compress partial keep_last focus preview aggressive
----------------------------------------------------------------------------------------
/compress False 2 None False False
/compress here True 2 None False False
/compress here 5 True 5 None False False
/compress up to here 3 True 3 None False False
/compress --keep=7 True 7 None False False
/compress -k 2 --dry-run True 2 None True False
/compress --preview here 4 True 4 None True False
/compress fix the auth bug False 2 fix the auth bug False False
/compress here 9999 True 100 None False False
/compress --aggressive False 2 None False True
Four things in that table are worth internalizing.
here [N] is the boundary form, and it counts user turns, not messages. The split walks backward to the start of the Nth most recent user message, so the verbatim tail always begins on a user turn and the rejoined history keeps the alternation providers validate. On a 17-message transcript with four user turns, here 2 produces a 9-message head to summarize and an 8-message tail ending up ['user', 'assistant', 'tool', 'assistant', 'user', 'assistant', 'tool', 'assistant'].
The aliases are undocumented in the help text: up to here 3, --keep 7, -k 7 and --keep=7 all take the partial path.
The count is clamped to a maximum of 100. /compress here 9999 does not no-op and does not error; it keeps the last 100 exchanges and compresses the rest. The floor is 1, and a fat-fingered here 0 lands on 1.
Anything that is not a flag or a boundary keyword is a focus topic. /compress fix the auth bug sends that phrase into the summary prompt, which is the answer to a session that compacted away the thread you were actually working on half an hour ago. If the boundary split would leave nothing to summarize, the whole thing falls back to full compression rather than rotating the session for a no-op.
Then there is --preview, otherwise --dry-run. It reports the numbers a real run would use without calling any model and without writing anything:
Would compress 5 of 17 message(s) (~41,250 tokens currently in context).
Boundary: keeping the last 3 exchange(s) (12 message(s)) verbatim.
That output came from calling the shipped summarize_compress_preview() on a synthetic transcript, so the counts are small. On your own session it is the cheapest way to answer “what is about to disappear”.
The refusals, verbatim
--aggressive is refused rather than parsed as a focus topic, because no surface implements an LLM-free hard truncate:
--aggressive is not supported; use '/compress here [N]' to keep only recent exchanges, or /undo to drop turns.
Below four messages you get Not enough conversation to compress (need at least 4 messages). If another compression already holds this session’s lock, you get a holder string rather than a false start, because a failed lock acquire is not proof that another compression is running and the two cases are worded differently.
A compression that changes nothing reports it. The shipped user-facing headlines distinguish outcomes that look identical in a log:
Compressed: 137 → 26 messagesNo changes from compression: 137 messagesCompression refused (summary would grow the conversation): 137 messages preserved
An aborted run says so in the same register: no messages dropped, conversation unchanged, and the auxiliary compression model named as the thing to check. Nothing pretends to have worked.
That last one is real and it is worth reading twice. When the generated summary would be larger than the block it replaces, Hermes refuses to commit it. There is a matching note for the opposite surprise: fewer messages can still raise the token estimate, because a dense handoff summary costs more tokens than the twelve short turns it replaced.
The gotcha: the summarizer’s window
The obvious cost optimization is pointing auxiliary.compression.model at a small cheap model. The documentation is blunt about what that costs you: the summary model must have a context window at least as large as the main agent model’s, because the entire middle section goes to it in one call. A smaller window returns a context-length error, and the docs name that as the most common cause of degraded compaction quality.
What Hermes does with a failed summary is the part that matters. Some failures always abort and preserve everything: access or quota errors, network errors, a summary truncated by the output cap, empty content, and provider overload (overloaded, at capacity, HTTP 529). Everything else lands on compression.abort_on_summary_failure, which is false by default and commits a deterministic fallback summary instead: old tool results pruned, and a static handoff in place of the summarized middle. That is a real rewrite with a real placeholder, and the earlier context is no longer in the request. Set abort_on_summary_failure: true if you would rather the conversation freeze until you run /compress or /new than accept a summary nobody wrote.
My read on that asymmetry: the dangerous summarizer is not the one that errors out, it is the one that returns something plausible and wrong, because that result gets committed.
Why it looks broken: the cooldown
A stall arms a per-session cooldown persisted in state.db, escalating 60s, 300s, 900s, never shorter than compression.context_timeout_seconds. While it is armed, ordinary threshold-triggered compaction is deferred, so a broken summary backend does not re-fire on every turn. Three paths still run a real attempt anyway, and one of them is yours: manual /compress passes force=True, which bypasses the summary-failure cooldown.
The store on this box has exactly one such row:
id: cron_266e70983203_20260829_110004
started: 2026-08-29 10:00:06Z ended: 2026-08-29 12:44:31Z
end_reason: cron_incomplete_no_output
cooldown_until: 2026-08-29 12:06:55Z
error: host compress_context timeout (no summary progress)
compacted rows in this session: 0 of 547
Zero compacted rows in a session that hit a summary stall. The transcript was never touched. I am not claiming the stall caused that cron run to produce no output (the store does not establish that either way), only that the stall left the session’s history exactly where it was.
What a committed compaction costs you
Two costs that do not show up in the token line.
A committed compaction rewrites history that the provider has already cached, so the prompt-cache prefix is invalidated. That is why the opt-in proactive prune (compression.proactive_prune_tokens, off at 0) refuses to commit unless it reclaims at least proactive_prune_min_reclaim_tokens (4,096): the gate exists to keep cache breaks episodic instead of firing on every tool iteration. Micro-compaction, compression.micro_compact: true, goes the other way and folds one old exchange into a running summary after every turn, which is smoother but breaks the cached prefix every turn. The docs are explicit that for some setups that cost exceeds the benefit.
The second cost is re-reading. Compression resets the agent’s file-read deduplication, deliberately, so it can read a file again after the content was summarized away. If your agent suddenly re-opens files it “already had”, that is the compaction boundary, not a loop.
How to confirm all of this on your own box
/compress here 3 --preview # numbers only, no model call, no writes
hermes config get compression # the effective knobs, not the docs' defaults
hermes config get compression on this box returns keys the configuration reference’s compression block does not list, which is worth knowing before you assume a setting is unset. An excerpt of the live output:
enabled: true
threshold: 0.5
threshold_tokens: 256000
target_ratio: 0.2
tail_mode: lean
protect_last_n: 20
min_tail_user_messages: 1
proactive_prune_tokens: 0
proactive_prune_min_reclaim_tokens: 4096
hygiene_hard_message_limit: 5000
context_timeout_seconds: 120
protect_first_n: 3
abort_on_summary_failure: false
micro_compact: false
To prove a compaction landed, do not read the reply text. Read the archive markers. Copy the database first, because the live one is held open by the running agent:
cp ~/.hermes/profiles/<profile>/state.db /tmp/probe.db
python3 - <<'PY'
import sqlite3
c = sqlite3.connect("file:/tmp/probe.db?mode=ro", uri=True)
print("compacted rows:", c.execute("SELECT COUNT(*) FROM messages WHERE compacted=1").fetchone()[0])
print("rotated sessions:", c.execute("SELECT COUNT(*) FROM sessions WHERE end_reason='compression'").fetchone()[0])
print("cooldown rows:", c.execute(
"SELECT id, compression_failure_error FROM sessions WHERE compression_failure_cooldown_until > 0").fetchall())
PY
On this profile that prints 13,375 compacted rows, 0 rotated sessions, and the one timeout row from August 29. Two facts, one command: compaction is rewriting history, and it is doing it in place.
The soft archive is why in_place: true (the default) is the right call for most operators. Pre-compaction turns stay under the same session id, marked inactive and compacted, and session_search still finds them. Setting in_place: false restores the older behavior where each compaction rotates to a new session id linked by parent_session_id, which is the mode to pick only if you have tooling that needs the lineage.
The settings worth changing
compression.progress_notices: trueif you run on Telegram or Discord and want to see routine compaction, because automatic compaction is silent on chat surfaces by default and only failures and manual runs are always visible.auxiliary.compression.modelpointed at a model whose window matches your main model, or you are opting into the fallback path on purpose.abort_on_summary_failure: truewhen a wrong summary is worse than a frozen session.idle_compact_after_seconds: 1800for a long-lived chat you return to hours later, so the first turn does not re-read stale context.proactive_prune_tokens: 48000on large-window models, where the 50% threshold rarely fires and bulky tool output otherwise rides along on every request.protect_first_n: 0for rolling sessions whose opening turn is no longer relevant.
Editing any compression.* key on a running gateway takes effect on the next message, no restart required.
The boundary is the contract
/compress is not a summarizer you invoke. It is a rewrite with a defined boundary, a pinned head, a token-budgeted tail, a refuse-to-grow check, a persisted cooldown ladder, and a rule that a failed summary either aborts and preserves everything or commits a placeholder you were told about. The 13,375 archived rows sitting behind a session that never changed its name are the argument: the transcript is negotiable, the session is not.
Sources
- Hermes Agent documentation, Context Compression (
compression.*reference, thresholds,in_place,abort_on_summary_failure, hygiene budgets, hot reload): https://hermes-agent.nousresearch.com/docs/user-guide/configuration/ - Hermes Agent documentation, Context Compression and Caching (dual compression system, the four-phase algorithm, computed values for a 200K model, summary-template fields, iterative re-compression, the summarizer-window requirement, the cooldown rungs): https://hermes-agent.nousresearch.com/docs/developer-guide/context-compression-and-caching/
- Hermes Agent documentation, Micro-compaction (off by default, per-turn passes, the cached-prefix trade-off): https://hermes-agent.nousresearch.com/docs/developer-guide/micro-compaction
- Hermes Agent documentation, Messaging (
/compressin the slash-command table): https://hermes-agent.nousresearch.com/docs/user-guide/messaging/ - Source in
NousResearch/hermes-agent@c1488ac947c9bc33fd65ec464548dc9d8edd6122:hermes_cli/partial_compress.py(argument grammar, clamps, split, preview),agent/conversation_compression_manual.py(parse_compress_args,MIN_MESSAGES, the--aggressiverefusal),agent/manual_compression_feedback.py(headlines and notes),agent/context_compressor.py(phases,_abort_on_summary_failure,_TERMINAL_SUMMARY_FAILURES, the compaction note),agent/compression_facade.py(force=Trueand the cooldown bypass),locales/en.yaml(gateway.compress.*strings): https://github.com/NousResearch/hermes-agent/tree/c1488ac947c9bc33fd65ec464548dc9d8edd6122 - Live output on the authoring box, 2026-09-21, Hermes Agent v0.21.3: read-only queries against a copied
state.dbfor the blogposter profile (264 sessions, 27,474 message rows, 13,375compacted=1, 0parent_session_id, 0end_reason='compression', the August 29 timeout row),hermes config get compression, and the shipped parsersparse_compress_args,split_history_for_partial_compressandsummarize_compress_previewexecuted against a synthetic 17-message transcript. The transcript in that parser run is synthetic; the store counts are not.