Hermes Agent Deep Cuts: The Grep That Guards the Loop
Run the exact same search_files call four times in a row. The first three return matches. The fourth returns this:
{"error":"BLOCKED: You have run this exact search 4 times in a row. The results have NOT changed. You already have this information. STOP re-searching and proceed with your task.","pattern":"dispatch","already_searched":4}
A search tool just refused to search. Not because the disk changed, not because of permissions. Because the loop decided you were stuck and that repeating yourself burns context you already paid for. The tool that the schema describes as “ripgrep-backed, faster than shell equivalents” is a control-loop component wearing a grep costume, and the search itself is the least interesting thing it does.
The mechanism: a shell pipeline, then a session layer
The file toolset runs through a per-task file-operations host. search_files dispatches by target: target="files" goes to rg --files (or a bounded find), target="content" to rg (or grep). Every content search is one shell command, built in file_operations_search.py:
set -o pipefail; rg --line-number --no-heading --with-filename \
[--max-columns 2000 --max-columns-preview] [-C N] [--glob G] [-l|-c] \
-- <pattern> <path> 2>/dev/null | head -n <limit + offset [+200 if context]>
The output is parsed back into matches, and each match’s content is clamped to 500 characters. Then three layers run that have nothing to do with grep.
Redaction runs first. Every matched line goes through redact_sensitive_text(..., file_read=True), the same scrubber that guards read_file output, so a search that happens to match a line next to an API key does not carry the key into context.
Filtering runs second. Results whose paths hit the read-denylist in file_safety.py are removed in place: project-local .env files anywhere on disk, .ssh keys, .netrc, .pgpass, .git-credentials, Hermes’ own auth.json, mcp-tokens/, the Bitwarden cache. The result JSON then carries "_omitted": "N result(s) omitted because they target credential, token, cache, or secret-bearing environment files." The model sees the note, never the file.
Formatting runs third, and it is the first behavior that looks like a bug until you see the code. When a search returns five or more matches, to_dict(densify=True) stops emitting the JSON matches array and switches to a lossless path-grouped text block: each path once, then indented line: content rows underneath. Fewer than five matches, and you get the structured array. Same call, no flag for it, and the output schema silently changes at a threshold of five (_DENSIFY_MIN_MATCHES = 5 in file_operations_common.py). The payoff is token economy: one path header instead of repeating a 60-character path on every row. The real result from this session carried nine rows across seven files; here is the shape, abridged by dropping the rows for config.py, guide.md, README.md, events.log, and a ~500-character single-line row from big.py:
total_count: 9
matches_format: "path-grouped: each file path on its own line,
followed by indented '<line>: <content>' rows"
/tmp/dc-search-demo/proj/src/app.py
2: def dispatch(task):
3: print("dispatch", task)
7: return dispatch
/tmp/dc-search-demo/proj/src/utils.py
4: x = dispatch("route")
The wrapper then applies its own pagination: offset/limit slice the parsed rows, and because the pipeline feeds through head, total_count is only the number of rows that survived the pipe. When results are truncated the JSON says so explicitly, and the tool appends a hint line (verbatim from a limit-2 run):
[Hint: Results truncated. Use offset=2 to see more, or narrow with a more specific pattern or file_glob.]
Two more registrations matter. search_files is capped at 100,000 result characters in the registry, and the whole result is a JSON string, so a huge hit list can never blow past the read budget in one call.
The loop guard
The repeat blocker lives in file_tools_read_tracking.py and the calling logic in file_tools.py. Every search builds a key:
search_key = ("search", pattern, target, str(path), file_glob or "", limit, offset, order)
That key is compared against the last key seen for the task. Same key again bumps a counter: the third consecutive identical search gets a _warning appended to the results, and the fourth returns the hard BLOCKED error instead of matches. The key has an omission worth naming: output_mode is not in it, and neither is context. In this session I hit the block with a content search, switched to output_mode="count" expecting a fresh start, and got blocked at “5 times in a row” because the key never changed. Only a change in pattern, target, path, file_glob, limit, offset, or order resets the count. So does almost any other tool call: model_tools.py resets the counter on every dispatch that is not read_file or search_files, which are the two tools that share the guard family.
The reset is easy to verify and easy to forget. After the block, a single intervening terminal call was enough, and the identical search ran normally again. I did exactly that in this session. The guard’s assumption is blunt: if nothing about the query changed and the files did not change, the agent already holds the answer, and the cheapest thing the loop can do is refuse to spend another round trip re-deriving it.
The zero-match probe ladder
A bare “0 matches” gives a model nothing to act on, so the tool refuses to return one. When a content search comes back empty, it runs up to three extra count-only rg probes, in order, and attaches the first one that hits as a warning (the ladder is _ZERO_MATCH_PROBES in the source):
rg -i: your casing may be wrong. Real output from this session:
{"total_count":0,"warning":"0 exact matches, but 9 case-insensitive match(es) in 6 file(s): /tmp/dc-search-demo/proj/src/app.py, /tmp/dc-search-demo/proj/src/utils.py, /tmp/dc-search-demo/proj/docs/guide.md, /tmp/dc-search-demo/proj/src/config.py, /tmp/dc-search-demo/proj/src/big.py (+1 more) — the pattern's casing may be wrong."}
-
rg --hidden --no-ignore: the match is in a file ripgrep skips by default. This probe prunes heavyweight trees first (theSEARCH_PRUNE_DIR_NAMESpolicy insearch_policy.py:node_modules,venv,.git,dist,target, caches, backup dirs), because it is about to walk exactly the places the default search refuses to. -
rg -F: only when your pattern contains regex metacharacters, in case you meant a literal substring.
The hidden-file probe is where the security story gets concrete. The needle dispatch_secret lived only in a .env file in my corpus. The search returned zero visible matches and this warning:
{"total_count":0,"warning":"0 matches in visible files, but 1 match(es) in 1 hidden or gitignored file(s): /tmp/dc-search-demo/.env — these are excluded by default."}
Read that carefully. The probe found the match, counted it, and reported the file’s path, but it will not show you the line. The credential filter from earlier makes sure of that. You asked the tool to grep for a secret and it told you the secret exists in a file you cannot read through it. That is the correct answer for an agent that must not leak .env contents into its own context.
What you inherit from ripgrep, including the parts you forgot
Two rg defaults shape every search and neither is visible in the schema. Hidden files and directories are always skipped, with no parameter to turn that on. And .gitignore is only honored when the search root is inside a git repository. I demonstrated the second one in this session with one git init: before it, a search for dispatch_log matched logs/events.log even though logs/ was in .gitignore. After git init, the same search returned zero and the probe correctly reported “1 match(es) in 1 hidden or gitignored file(s)”. Ripgrep does not apply ignore files outside a repo, so neither does the tool. If your scratch directory is not a git repo, your ignore rules silently do not exist, and searches will walk into build/ and dist/ trees you thought you had excluded.
macOS adds a fourth default: broad searches rooted above the home directory prune Desktop, Documents, Downloads, Library, Movies, Music, and Pictures before they can trigger an unattended TCC privacy prompt, and the result carries a warning naming what was skipped. That logic is darwin-only and env-local, so I could not exercise it on this Linux box; it is in the source, gated on sys.platform == "darwin".
There is also a version floor hiding in the file-search path. order="modified" sorts by mtime using rg --sortr=modified, which only exists in ripgrep 14 or newer. The tool parses rg --version against a full SemVer regex, and on an older rg it fails with a precise error instead of silently returning unsorted results: “Exact modification-time order requires ripgrep 14 or newer; upgrade ripgrep or use order=‘discovery’.” The find fallback needs GNU find’s -printf for the same feature. This box runs ripgrep 15.1.0, so order="modified" works and returns files newest-first:
{"total_count":5,"files":["/tmp/dc-search-demo/proj/src/app.py","/tmp/dc-search-demo/proj/src/multi.py","/tmp/dc-search-demo/proj/src/big.py","/tmp/dc-search-demo/proj/src/utils.py","/tmp/dc-search-demo/proj/src/config.py"]}
Advanced usage that the schema does not advertise
Multi-root search in one call. The tool is built to recover when a model stuffs several paths into the single path parameter, and you can exploit that deliberately. Pass a comma-separated list and it probes each entry, searches the ones that exist, and merges the results:
{"warning":"path contained 2 entries; searched 2 that exist"}
Add a root that does not exist and it tells you exactly what it skipped: “searched 1 that exist; skipped missing: /tmp/dc-search-demo/proj/nope”. File-name searches do one global traversal across all roots, so order="modified" stays exact across them. Content searches merge per-root and slice, which is worth knowing if you page: each root is searched with the same limit before the merge slices, so pagination across multi-root content results is approximate by design.
Typo recovery. A missing root is not a bare error. The tool lists the parent directory and suggests siblings whose names are close to what you asked for. Search for big.pyy and it points at big.py.
Multiline regexes work, with a receipt. A literal newline in the pattern (or a \n escape with an odd number of backslashes) auto-enables rg -U, and the tool says so in the result:
{"total_count":2,"matches":[...],"warning":"Pattern contains \\n — multiline mode (-U) was enabled automatically so the regex can match across line boundaries."}
The grep fallback cannot do this, so on a box without rg, a newline pattern errors and the tool explains that line-oriented mode rejects it.
Globs are includes, patterns are regexes. file_glob becomes rg --glob, so it filters by file name while pattern stays a regex over content. A bare file-search pattern like app.py is wrapped to *app.py* so it matches at any depth; the wrapping is skipped once the pattern contains a slash or already starts with *. Dash-prefixed roots are protected: rg -- terminates options and find gets a ./ prefix so a directory named -cache is never parsed as a flag.
The gotchas, collected
The happy path fails in four ways that look like bugs.
The block. Four identical searches in a row and the tool screams at you. It is not broken, and switching output_mode will not help, because the mode is not part of the guard key. Change the path or pattern, run any non-search tool, or page with a different offset. The offset being in the key is deliberate: paging through truncated results is legitimate repetition, so the guard lets it through.
The silent .gitignore gap. Outside a git repo your ignore rules do not apply, so a “clean” search can light up dist/ and node_modules (if unignored) or, after a git init, suddenly stop matching files it matched minutes earlier. The tool inherits this from rg; the probes will at least label the cause when the count drops to zero.
The content clamp. Each match line is cut at 500 characters, and rg itself is told --max-columns 2000 --max-columns-preview so a match inside a multi-megabyte single-line dump does not make head -n emit the whole file. A minified bundle can still produce rows of near-500 filler characters per match. The clamp is why the big.py row in my corpus output ends mid-string of x’s. If you need the full line, that is what read_file with an offset is for.
The hidden-file wall. There is no parameter to include hidden or gitignored files, ever. The tool’s own probes use --hidden --no-ignore internally, but the main search never will, and the credential filter would strip .env results anyway. When you genuinely need to search .git/ or a dotfile tree, the escape hatch is the terminal tool and raw rg --hidden --no-ignore. Know which tool you are in: one is a guarded member of the loop, the other is a raw shell.
How to verify it is working
All of this is checkable in a minute. Confirm the engine and its version: rg --version must report 14 or newer for order="modified", and command -v rg should resolve. Reproduce a content search by hand and compare totals, since the tool is only a pipeline plus a parser:
set -o pipefail; rg --line-number --no-heading --with-filename \
--max-columns 2000 --max-columns-preview 'dispatch' /tmp/dc-search-demo \
2>/dev/null | head -n 50 | wc -l
That returns 8 here, matching the tool’s total_count exactly. The same search returned 9 earlier in this session, before the corpus became a git repo; once git init ran, logs/events.log fell under the logs/ rule in .gitignore and dropped out of both the raw pipeline and the tool. That is the ignore-semantics gotcha from above, demonstrated live. Trigger the guard on purpose: run one identical search four times and watch the warning appear on the third and the block on the fourth, then run any unrelated tool call and confirm the same search passes again. If you manage a fleet of machines, remember the version floor: the same prompt that returns sorted files on rg 15 fails cleanly on rg 13, and now you know the error is a feature, not a malfunction.
The search was never the point. An agent loop needs someone to remember what the agent already knows, to explain empty results instead of guessing at them, and to keep credential files out of context even when the agent asks for them. Hermes put that job inside the grep. Replace search_files with a raw rg because the output format annoyed you, and you are not removing a wrapper. You are removing the guard, the probes, the redaction, and the loop’s memory of what it already saw.
Sources
- Built-in Tools Reference, file toolset: https://hermes-agent.nousresearch.com/docs/reference/tools-reference
- Search engine and pipeline implementation: https://github.com/NousResearch/hermes-agent/blob/main/tools/file_operations_search.py
- Tool handler, guard calls, result shaping: https://github.com/NousResearch/hermes-agent/blob/main/tools/file_tools.py
- Repeat-guard and not-found tracking: https://github.com/NousResearch/hermes-agent/blob/main/tools/file_tools_read_tracking.py
- Densify threshold and pagination normalization: https://github.com/NousResearch/hermes-agent/blob/main/tools/file_operations_common.py
- Search tool reset on other tool calls: https://github.com/NousResearch/hermes-agent/blob/main/model_tools.py
- Heavyweight-tree prune policy: https://github.com/NousResearch/hermes-agent/blob/main/agent/search_policy.py
- Read-denylist and credential classification: https://github.com/NousResearch/hermes-agent/blob/main/agent/file_safety.py
All command outputs shown above were produced in this session against a disposable corpus under /tmp/dc-search-demo (ripgrep 15.1.0). Version-gated and darwin-only behaviors are labeled from source where they could not be exercised locally.