Hermes Agent Deep Cuts: Session Search Is a SQLite Query, Not a Memory
Ask a Hermes agent “what did we do about the deploy issue last week” and the answer, when it comes, is not a paraphrase. session_search hands back stored rows from state.db: a session id, a message id, a snippet, and the real transcript around the hit. Zero LLM calls happen on the way there, which is why the docs call the whole thing free. The part nobody expects is what the ranking has to survive to deliver that answer. This profile’s state.db holds 237 sessions and 194 of them are cron jobs, dense with the words “deploy,” “post,” and “session.” A bare FTS5 MATCH for deploy returns nothing but automation sessions. The tool actively pushes your own history back to the top, and that fix is the subject of this post as much as the search itself.
The mechanism: FTS5 over everything you ever said, no LLM in the loop
Every conversation Hermes runs, in every source, lands in state.db as rows in a messages table. The search index is a set of SQLite FTS5 virtual tables built over that content, and the search path is plain MATCH SQL. On this DB the layout is explicit if you open it read-only:
CREATE VIRTUAL TABLE messages_fts USING fts5(
content, tool_name, tool_calls,
content='messages', content_rowid='id'
);
CREATE VIRTUAL TABLE messages_fts_trigram USING fts5(
content, tool_name, tool_calls,
content='messages_fts_trigram_src', content_rowid='id',
tokenize='trigram'
);
Two indexes over the same rows. The base messages_fts uses the default tokenizer; messages_fts_trigram adds the trigram table for substring robustness and CJK. Both are backed by content='messages', meaning the index derives from the canonical table rather than holding its own copy, so a query that MATCHes joins back to real rows. session_search is a thin wrapper: the tool resolves which shape you want from which arguments you set, then calls into the session DB. The module docstring is blunt: “No LLM calls, every shape returns actual DB messages.”
That last point is the reason to trust the answers. A recalled “fact” is not a paraphrase the model produced from its prior guesswork. It is the original message row, decoded, slimmed, and handed back with an id you can scroll to.
The documentation lays out the cost contract in a table that is easy to skip: persistent memory is ~1,300 tokens in every single prompt, session search is free because it only runs when called, roughly 20ms for a query and 1-2ms to scroll. Memory is curated context that is always visible. Session search is on-demand recall over everything, and there is no hand-off between them.
The four calling shapes: no mode flag, just argument shape
There is no mode parameter. You get behavior by which fields you fill, and the session-search section of the user guide calls these the four shapes. There is docs drift worth knowing about: the memory feature page still says “three calling shapes,” the sessions page says four. The source wins, and the source documents four:
- Discovery: pass
query. FTS5 runs, hits are deduped by session lineage, and the top N sessions come back. Adaptive detail is the default: the highest-ranked result gets a full anchor window plus bookends (the shape code calls this “goal to match to resolution”), lower-ranked results stay compact. Passdetail="full"to hydrate everything. - Scroll: pass
session_idplusaround_message_id. Returns a±window(clamped to 1-20, default 5) centered on the anchor. No FTS5, no bookends, just the slice. You page forward by feedingmessages[-1].idback in, backward by feedingmessages[0].id. - Read: pass
session_idalone. Whole session, or a bounded head+tail view for long ones. This also resolves an@session:<profile>/<id>link. - Browse: no arguments. Lists recent sessions chronologically with previews.
I reproduced all four against this profile’s DB. Discovery for telegram OR cli ranked the actual interactive sessions first:
{"success": true, "mode": "discover", "query": "telegram OR cli", "detail": "adaptive",
"results": [{
"session_id": "20260803_211407_2c670c",
"when": "August 03, 2026 at 09:15 PM",
"source": "cli", "title": "Refreshing Blog Post Hero Images",
"matched_role": "assistant", "match_message_id": 1193,
"snippet": "...xai` (`gpt-image-2-medium`), and `x_search` enabled in the >>>CLI<<< >>>telegram<<< toolsets..."
}]}
Reading the same session back told me something the docs do not lead with: a 254-message session trims to head 20 plus tail 10, and includes the session link @session:blogposter/20260803_211407_2c670c. Scroll around message 1193 with window=2 returned 1-message-per-side real turns with tool_calls intact, including the escaped function-call JSON. That is the resolve step: discovery finds the match, scroll reads around it, and the tool’s own link_hint tells the model to write that link verbatim inline because it renders as a clickable titled link in the desktop UI.
The ranking pipeline, and the bug that shaped it
The discovery path is where the interesting logic lives, and it is covered carefully in hermes_state_search.py. Step by step:
- Title short-circuit. If the query matches a session title (the code strips stray quotes and backticks because models habitually quote a remembered title), that session is resolved directly by title. A
resolve_session_by_titlemiss falls through to FTS5. - Sanitize. Raw FTS5 special characters raise on
MATCH, so the input is scrubbed before it ever touches SQL. - Search with demotion. Rows are pulled by BM25 rank up to a
_DISCOVER_SCAN_LIMITof 300, then a stable sort moves cron hits below interactive ones before dedupe. - Dedupe by lineage. Hits collapse to a per-session root id so one conversation does not flood the top N.
- Adaptive hydration. Only the top result gets full bookends and window; the rest stay compact.
The demotion step is the fix for a real failure documented in the source: issue #19434, the “recall blindness” case. The comment is worth quoting because it names the mechanism and the consequence:
Cron jobs run on a schedule and accumulate large volumes of repetitive vocabulary (recurring project names, dates, “session”, summaries); under bare BM25 they dominate the top-N FTS rows and starve out the user’s own interactive sessions.
Demotion is a sign of the sort key, not a filter. Cron content stays reachable when it is the only match; interactive sessions always win when both match. I verified the shape on this live database. A bare FTS5 MATCH 'deploy' on this profile returns nothing but a cron_266e70983203... session with title “Denny Sentinel daily AI news post” in all three top slots. The same query through the real discover pipeline ranks four sessions and none of them are cron:
[rank] src=telegram title='XActions added to profile'
[rank] src=cli title='Refreshing Blog Post Hero Images'
[rank] src=a2a title='No prior memory in this session.'
[rank] src=cli title='Mem0 Memory Setup and Verification'
That is the whole point of demotion made visible: your interactive history exists, it is just being crushed by a quarter-million cron characters before the ranking logic intervenes. The _DEMOTED_SESSION_SOURCES = ("cron",) tuple is only one source, and _HIDDEN_SESSION_SOURCES = ("kanban", "subagent", "tool") excludes those entirely, on the ground that they are not the user’s history.
There is a second current-session guard in the same query path. The current lineage is skipped unless its transcript has left live context, so a live conversation does not shadow older distinct sessions. I demonstrated it with a nonsense query engineered to match only the running session: with the live session id passed as current_session_id, the result count drops to zero; without the guard it returns that one current-session hit. A real user does not see this, because the tool is normally handed the live id by the runtime, but it is why “search your own working session” returns nothing instead of echoing back the turn you are in.
The sanitizer: a lexer, not a filter
_sanitize_fts5_query reads like a small tokenizer, and getting its behavior straight will save you from confusing wrong-looking misses. I ran it side by side with realistic inputs:
| Input | Sanitized | Why |
|---|---|---|
docker deployment | docker deployment | FTS5 ANDs terms by default |
"exact phrase" | "exact phrase" | preserved quoted phrase |
docker OR kubernetes | docker OR kubernetes | boolean operators survive |
python NOT java | python NOT java | NOT survives |
deploy* | deploy* | prefix survives |
chat-send | "chat-send" | hyphenated term quoted as a phrase |
my-app.config.ts | "my-app.config.ts" | dotted/underscored term quoted |
TODO: fix | TODO fix | colon stripped (otherwise column:term) |
deploy *prod | deploy prod | leading asterisk dropped (prefix needs a char) |
AND leading | leading | dangling boolean stripped |
"unterminated | unterminated | unmatched quote demoted to whitespace |
The two that matter operationally: AND is the default, so deploy dennysentinel requires both terms to appear, and hyphenated names become phrases. If your recall query is the name of a tool with a dash, quote it yourself or let the sanitizer do it. And when you get zero results, the empty response spells out the fix instead of shrugging: FTS5 ANDs all terms by default, so broaden with OR, use quoted phrases for exact matches, NOT to exclude, and a trailing star for prefix matches.
The gotcha: search tool output and you never touch FTS5
Here is the one that makes the feature behave like it is broken. The discovery default role_filter is user,assistant, because “tool output is usually noise.” But the moment you genuinely need to search inside tool results, you trip a code path that is not FTS5 at all. From the source:
if role_filter and "tool" in role_filter:
matches = self._search_messages_like_fallback(query, ...)
return self._finalize_search_matches(matches, ...)
An explicit role_filter="tool" skips the whole FTS5 pipeline and does a substring LIKE scan over canonical rows. That is deliberate on two counts. First, tool rows are excluded from the trigram and CJK indexes entirely, so a MATCH against those would miss them. Second, oversized tool outputs are only indexed as a bounded prefix into content, so full-body search has to scan. The base messages_fts index does hold tool_name and tool_calls columns, so a MATCH on those does hit FTS5; it is the content of heavy tool calls that is bounded. If you ask for role_filter="tool" and the query takes noticeably longer, you are not watching FTS5, you are watching a table scan, and on a multi-hundred-MB state.db with 25,000+ messages that is where the 20ms budget goes away.
There is a second, subtler gotcha: the FTS index rebuilds in the background and results from not-yet-indexed rows are thin until it finishes. fts_rebuild_status() on this DB returns None (healthy, index warm). But when a rebuild IS in progress, the discovery payload embeds a note that the search index is reindexing at N% and that results from older messages may be incomplete until it finishes. During the gap, the code supplements with a bounded LIKE scan of the (progress, high_water] range so old messages do not silently vanish mid-rebuild. So a “no results” during a rebuild is not a fact, and the tool goes out of its way to tell you so.
Advanced surface: sort, detail, and the CLI companion
Two parameters change meaning on their own. sort="newest"|"oldest" biases on top of BM25 ranking: omit it for exploratory recall, pass newest for “where did we leave X” and oldest for “how did X start”. detail="full" forces every discovery result to hydrate bookends plus the full anchor window instead of only the top match, which is worth the tokens when you are comparing several candidate sessions to pick the right one.
The CLI mirror is hermes sessions, and its subcommands carve into the same database:
hermes sessions stats # counts + DB size
hermes sessions list # recent sessions
hermes sessions export # JSONL / Markdown / QMD
hermes sessions prune # delete old sessions by time/source/title
hermes sessions optimize # merge FTS5 segments + VACUUM, no data change
hermes sessions optimize-storage # migrate search index to compact v23 layout
optimize is the non-destructive first move when state.db grows: it merges FTS5 segments and reclaims space without deleting any session data. optimize-storage is the newer compact index migration that shrinks large DBs further. I ran sessions stats against this profile live: 237 sessions, 25,563 messages across cli (12), telegram (2) and the rest cron/A2A/TUI/desktop, 149.5 MB.
The test suite backs the whole contract: tests/tools/test_session_search.py passes 53 tests green in 3.52s across the four shapes, the demotion, cross-profile reads, and scroll rebinding.
How to verify it is actually working
You do not need the agent at all to prove this at the shell. Open the read-only database and fire an FTS5 MATCH yourself:
/home/dazeb/.hermes/hermes-agent/.venv/bin/python3 -c "
import sqlite3
db = sqlite3.connect('file:$HOME/.hermes/state.db?mode=ro', uri=True)
for r in db.execute(\"SELECT rowid, rank FROM messages_fts WHERE messages_fts MATCH ? ORDER BY rank LIMIT 3\", ('deploy',)):
print(r)
"
If that returns numbered rows with a rank column, the index is live. To see the sanitizer in action and the shape dispatch, import the tool and run discovery against your profile DB read-only:
from hermes_state import SessionDB
from tools.session_search_tool import session_search
db = SessionDB(db_path='$HOME/.hermes/profiles/blogposter/state.db', read_only=True)
print(session_search(query='telegram OR cli', limit=3, db=db)[:1200])
The tell is the mode: discover field and the absence of any token-cost field. If every answer comes back as a summary no matter what you query, you are not on this path and you should suspect a different recall layer being substituted. The correct output is actual stored messages with real session and message ids, a snippet, and the bookend window for the top hit.
Take a second and check your own ratio. Mine is 194 cron sessions and 12 interactive ones in the same database. That is not a reason to delete the cron session history, but it is a reason to know the ranking is doing real work to keep your actual conversations findable. session_search is not a memory you have to teach; it is a search engine over your own history, and the worst thing about it is how easy it is to let your automations drown out the parts that are you.
Sources
- Session Search Tool (four calling shapes, FTS5 syntax, parameters): https://hermes-agent.nousresearch.com/docs/user-guide/sessions#session-search-tool
- Session Search vs memory (cost contract): https://hermes-agent.nousresearch.com/docs/user-guide/features/memory#session-search
- How sessions work, storage locations, DB schema: https://hermes-agent.nousresearch.com/docs/user-guide/sessions
tools/session_search_tool.py: shape dispatch, lineage dedupe, cron demotion, adaptive/compact hydration,@sessionlinks, cross-profile read, scroll rebindinghermes_state_search.py: FTS5 sanitizer, query routing, LIKE fallback forrole='tool', rank ordering, rebuild-status gap supplementhermes_state_fts.py: FTS5 index DDL, trigram and CJK-bigram tables, rebuildhermes_state_common.py:MAX_FTS5_QUERY_CHARS = 2048tests/tools/test_session_search.py: 53 passed in 3.52s (this run)
All live outputs in this post were produced on 2026-09-13 against hermes-agent v0.21.2 (d62716c7, 2026-09-11), reading this profile’s state.db read-only and running tests/tools/test_session_search.py and hermes sessions stats on the same box. Nothing was reconstructed from memory; every json snippet, sanitizer row, and ranked result above came from a real execution against the 149.5 MB, 25,563-message live database.