Hermes Agent Deep Cuts: The 100,000-Character Wall Inside read_file
Part of the Hermes Agent: Deep Cuts series

Hermes Agent Deep Cuts: The 100,000-Character Wall Inside read_file

Here is a real read_file result, cut down to the fields that matter, from a file I created this session with 60 lines of 1,900 characters each:

"truncated": true,
"truncated_by": "bytes",
"next_offset": 53,
"hint": "Output truncated at the 100,000-char read budget after 52 line(s)
(showing lines 1-52 of 60). Use offset=53 to continue."

The file is only 60 lines long. You passed no limit, so the default of 2,000 lines was nowhere near exhausted. And the tool still cut you off at line 52, because the real limit was never the line count. It was a 100,000-character budget you never set, measured on the formatted output (line-number prefixes and all), and it hands you a next_offset cursor instead of an error. That is the part of read_file nobody reads the docs for, and it is the whole point of the tool.

read_file is not “cat with line numbers.” It is a context-window governor. Every day a Hermes operator runs it a hundred times without thinking, and every run is being filtered through three independent layers: a character budget, a read cache, and a set of guards that will flat-out refuse to read things. If you only know the happy path (read_file path), you have been leaving tokens, loop protection, and silent data-loss detection on the table.

The mechanism: three layers, in the order they run

The model-facing entry point is read_file_tool() in tools/file_tools.py, which wraps FileReader.read_file() in tools/file_operations.py. The order of operations matters, because each layer can short-circuit the next.

First, pagination is normalized. normalize_read_pagination() clamps offset to at least 1 and limit to tool_output.max_lines (default 2,000). The comment in the source is honest about why this exists: the schema declares min/max values, “but not every caller or provider enforces schemas before dispatch,” so the clamp stops an invalid value like offset=0 from leaking into a sed -n '0,-1p' range.

Second, the actual read is a shell pipeline, not a Python file handle:

sed -n 'OFFSET,END_LINE p' PATH | cut -b1-(4 * max_line_length + 1)

Two things hide in there. The cut -b byte clamp is 4 * max_line_length + 1 because UTF-8 uses up to four bytes per character, so the byte clamp has to be four times the character clamp to avoid splitting a multibyte character into mojibake. Then a second, Python-side clamp in _add_line_numbers() cuts any line longer than max_line_length (default 2,000) and appends ... [truncated]. So a line is clamped twice: once at the byte level by cut, once at the character level by Python.

Third, the line numbers are added with a compact <n>| gutter, not a fixed-width padded one. This is not an aesthetic choice. The source documents an A/B test (Sonnet 4.6, two passes): a zero-padded gutter cost roughly 48% more tokens than bare content and about 16% more than the compact form, because the leading spaces and zeros tokenize into extra tokens on every single line. Dropping the numbers entirely regressed line-referencing tasks because the model hand-counted and was off by one. So the compact gutter is a measured token optimization, not a style preference.

Only after the content is formatted does the character budget kick in. _get_max_read_chars() reads file_read_max_chars from config.yaml (default 100,000), and _truncate_to_char_budget() trims to the last complete line that fits, then returns a next_offset so the model can paginate forward. The key design decision: Hermes used to hard-reject an oversized read and force the model to guess a smaller limit, burning a round-trip and returning nothing. Now it returns what fits plus a cursor.

The guards: what read_file refuses to read

The budget is the polite layer. The guards are the rude ones, and they run before, after, and around the read.

Device and special files. There is a hardcoded frozenset, _BLOCKED_DEVICE_PATHS, that covers the infinite-output and blocking devices: /dev/zero, /dev/random, /dev/urandom, /dev/full, /dev/stdin, /dev/tty, /dev/console, /dev/stdout, /dev/stderr, and the /dev/fd/* aliases. That blocklist is name-based. There is a second, stat-based guard, _special_file_kind(), that catches the whole class: any FIFO, socket, character device, or block device, wherever it lives. A socket at logs/live.pipe hangs read_file just as hard as /dev/zero, and the name blocklist would never see it. The message is blunt for a reason, naming a self-shipped DoS: reading a FIFO blocks until the exec timeout.

Binary files. has_binary_extension() blocks by extension before any I/O, and the error names the escape hatch: “Use vision_analyze for images, or terminal to inspect binary files.” The comment notes the extension is only a claim, and the content-sniffing path names the actual magic-byte type for extension-less or lying files.

Credential stores. get_read_block_error() in agent/file_safety.py maintains a read deny-list: auth.json, auth.lock, .env, .git-credentials, .anthropic_oauth.json, webhook_subscriptions.json, auth/google_oauth.json, and friends, resolved against both the active HERMES_HOME and the global Hermes root so a profile session cannot slip past it. The source is explicit about what this is not: a security boundary. The agent can still cat ~/.hermes/auth.json in the terminal. The deny exists to make an attempted credential read obvious in the logs and to push the model toward the correct internal channel.

The negative-result cache. If a read already discovered a path does not exist (within a TTL), the cached “File not found” is returned without re-spawning the subprocess and re-walking the parent directory. write_file and patch clear it for the path they touch.

The cache and the tripwire

This is the part that changes how you drive an agent. Two independent mechanisms both try to stop the model from burning context on redundant reads.

The first is deduplication, keyed on (resolved_path, offset, limit) plus the file’s mtime. Read a file, read it again without it changing, and the second call returns a stub instead of the content:

{"status": "unchanged", "message": "File unchanged since last read. The content
 from the earlier read_file result in this conversation is still current...",
 "path": "/tmp/hermes-small.txt", "dedup": true, "content_returned": false}

I triggered that live. Two consecutive reads of the same small file returned the stub on the second call. If you are scripting against read_file and expect content every time, the content_returned: false field is your signal that the model is expected to reuse what it already has.

The second is the loop tripwire. The tracker counts consecutive reads of the exact same (path, offset, limit) region. The third read gets a warning. The fourth gets a hard block:

BLOCKED: You have read this exact file region 4 times in a row. The content
has NOT changed. You already have this information. STOP re-reading and
proceed with your task.

This is the gotcha that makes the happy path fail for a stuck agent. A model that keeps re-reading a file to “double-check” gets cut off. The counter is deliberately scoped to truly consecutive reads: notify_other_tool_call() resets it whenever any tool other than read_file or search_files runs. So a genuine interleaving of read-then-act-then-read never trips it, but a tight re-read loop does. There is a second escalation path: repeated dedup stubs (a “weak tool-follower” that ignores the “refer to earlier result” hint) hit a hard block after two stubs for the same key.

One consequence worth knowing: reset_file_dedup() is called after context compression, because the original read content has been summarized away, so the model genuinely needs the full text again. Without that reset, a post-compression read would return an “unchanged” stub pointing at content that no longer exists in context.

Document extraction, and the silent data loss it catches

read_file also extracts structured documents to text before the binary guard would reject them. EXTRACTABLE_EXTENSIONS covers .ipynb, .docx, and .xlsx using only the standard library (zipfile plus XML). When the optional firecrawl-anydoc package is present, coverage widens to .doc, .ppt, .xls, OpenDocument, RTF, EPUB, and PDF, converted to Markdown by its Rust core. The import is lazy and set to prompt=False, so a read never blocks on an install prompt, and a failed load is retried only after a 300-second cooldown rather than hammering pip on every call. Documents larger than 50 MB are refused before conversion.

The interesting failure mode is PDF. anydoc, like every text-layer extractor, returns nothing for scanned or image-only pages, and it emits no page markers, so a mostly-scanned PDF converts “successfully” into a few section headers with empty bodies. That is silent data loss the model cannot see. Hermes detects it with a second pass: it runs pdftotext (form-feed page separators), counts characters per page, and treats any page under 20 characters as empty. When at least two pages are empty and they cross a ratio threshold, it appends an EXTRACTION COVERAGE WARNING to the output listing each empty range and labeling it with the last text extracted before it, so the agent can decide which gaps it actually needs instead of OCRing everything.

The recovery instruction in that footer is exact: render just the relevant range with pdftoppm -jpeg -r 150 -f <first> -l <last> 'file.pdf' /tmp/page and inspect each with vision_analyze, or use the marker-pdf skill for bulk OCR.

There is one more guard on the far side of extraction, and it is the kind of thing that only exists because someone got bitten. read_file extracts a .docx to text, so the model plausibly believes it holds the document and tries to write the edited text back with write_file or patch. A plain-text write can never produce a valid OOXML or OLE container, so _check_binary_document_write() rejects the write instead of silently destroying the original.

Advanced usage: the two config knobs

There are exactly two places to bend the behavior, and both are config.yaml, not tool flags.

# raise the per-read character budget (default 100000)
file_read_max_chars: 200000

# raise the line-count cap and per-line truncation cap
tool_output:
  max_lines: 5000          # read_file pagination + truncation cap (default 2000)
  max_line_length: 4000    # per-line clamp before "... [truncated]" (default 2000)

file_read_max_chars is read on first use and cached for the process lifetime, so changing it requires a restart, not just a config save. The tool_output block is shared: max_lines gates read_file pagination, and max_line_length gates the per-line clamp. The same module also carries max_bytes for terminal output. All three fall back to defaults on any malformed value, so a bad config degrades quietly instead of breaking the tool.

The practical takeaway for an operator: if you routinely read large generated files (logs, wide CSVs, minified JSON, lockfiles), you have two independent ceilings. A file with thousands of short lines hits the max_lines ceiling first and returns a line-count hint. A file with a few very long lines hits the max_line_length clamp first and silently truncates each line. A file with many medium lines hits the 100,000-character budget and returns next_offset. You have to know which ceiling you are up against to raise the right knob.

The gotchas

A single line longer than the budget is unrecoverable. If even the first line of a file exceeds the character budget, _truncate_to_char_budget() clamps that one line at the budget boundary and the hint appends the warning: “its remainder is not retrievable via offset.” Offset pagination is line-based, so there is no cursor that reaches the tail of a line that was cut mid-line. The only recovery is terminal (cut, sed, dd) or reading the raw bytes another way.

“truncated: false” does not mean you saw everything. I read a file containing a 160,000-character line and got back that line clamped to 2,000 characters with ... [truncated] appended, and the result still reported "truncated": false. The truncated field only reflects the line-count ceiling, not per-line clamping. Trust the ... [truncated] marker on individual lines, not the top-level flag.

The fourth read cuts you off. If an agent re-reads the same region four times in a row, it is hard-blocked. This is a feature, but it will surface as a confusing error if you are driving the agent manually and did not realize the earlier reads are still in context. The escape is to run any other tool, which resets the consecutive counter.

Scanned PDFs lie to you. A scanned document “converts” fine and drops whole pages. The coverage warning is the only signal. Read to the bottom of any PDF extraction and check for it before trusting the text.

How to verify it is actually doing what you think

Every claim above came from the installed v0.20.6 source plus live runs in this session. You can reproduce the whole thing in a minute.

# 1. The character budget, not the line count, is the real ceiling
$ python3 -c "open('/tmp/many.txt','w').write('\n'.join(['y'*1900]*60)+'\n')"
$ read_file /tmp/many.txt        # 60 lines, but truncated at line 52
#   -> truncated_by: "bytes", next_offset: 53, "Use offset=53 to continue"

# 2. A single long line is clamped per-line and the tail is gone
$ python3 -c "open('/tmp/one.txt','w').write('x'*160000)"
$ read_file /tmp/one.txt         # clamped to ~2000 chars + "... [truncated]"

# 3. Dedup: read twice, get a stub the second time
$ read_file /tmp/small.txt       # full content
$ read_file /tmp/small.txt       # {"status":"unchanged","dedup":true,"content_returned":false}

# 4. The loop tripwire: read the same region four times in a row -> BLOCKED

# 5. Device and binary guards
$ read_file /dev/zero            # "device file that would block or produce infinite output"
$ read_file photo.jpg            # "Cannot read binary file... use vision_analyze"

# 6. Config knobs
$ grep -n "file_read_max_chars\|tool_output" config.yaml

The one line to remember: read_file will happily tell you it read a file while trimming it three different ways. The budget field, the next_offset cursor, the content_returned flag, and the ... [truncated] marker are the parts that tell you what you actually got. Everything else is just line numbers.

Facts, inference, and open questions

Observed (installed v0.20.6 source plus live runs in this session on 2026-08-28): read_file_tool() in tools/file_tools.py wraps FileReader.read_file(); normalize_read_pagination() clamps offset to at least 1 and limit to tool_output.max_lines (default 2,000); the shell read is sed -n 'O,E p' | cut -b1-(4*max_line_length+1), with a Python-side len(line) > max_line_length clamp appending ... [truncated]; _add_line_numbers() uses a compact <n>| gutter after a documented token A/B test; _DEFAULT_MAX_READ_CHARS is 100,000, configurable via file_read_max_chars, and _truncate_to_char_budget() trims to the last complete line and returns next_offset; a first-line-over-budget read is clamped mid-line with “remainder not retrievable via offset”; total_lines comes from wc -l so a no-trailing-newline file undercounts by one; the device path blocklist and stat-based _special_file_kind() refuse non-regular files; get_read_block_error() denies credential stores by path and is documented as a soft guard, not a security boundary; the dedup cache is keyed on (path, offset, limit) plus mtime and returns a content_returned: false stub; the loop tripwire warns at three consecutive reads and hard-blocks at four, with notify_other_tool_call() resetting on any non-read/search tool; EXTRACTABLE_EXTENSIONS covers .ipynb/.docx/.xlsx via stdlib and firecrawl-anydoc adds PDF and legacy formats; the PDF coverage check uses pdftotext with a 20-character-per-page emptiness threshold and emits an EXTRACTION COVERAGE WARNING with a per-gap map; _check_binary_document_write() rejects text writes to extracted binary documents. Live: the char-budget next_offset: 53 result, the single-line clamp, the dedup stub, the /dev/zero device block, and the .jpg binary block.

Inference: the three-layer design (budget, cache, tripwire) exists because context is the scarcest resource in agent execution, and the model cannot be trusted to self-limit. The dedup and loop guards are cheap to run and save real tokens, but they trade against a subtle failure mode: a stub or block that fires when the model genuinely needs fresh content, which is why reset_file_dedup() is hooked into context compression.

Open questions: I verified the char budget, per-line clamp, dedup stub, device guard, and binary guard live. I did not exercise the PDF coverage warning end to end against a real scanned document (no scanned PDF was at hand), nor the firecrawl-anydoc lazy-install path, nor the binary-document write-back rejection. Those are described from the v0.20.6 source, not from a live run.

Sources

Keep reading