Hermes Agent Deep Cuts: The @ Reference Is an Injection Pipeline, Not a Shortcut
Watch what actually happens when you type @file:AGENTS.md and hit enter. The file’s text does not appear in your terminal as it streams. It is already in the message, cut and fenced, before the model’s first token is generated. Run the expansion yourself against a repo and the returned string carries a note, a token count, and the whole file wrapped in a markdown fence, appended under --- Attached Context ---. The @ reference is a preprocessing stage with its own token budget and its own security gate, and the sooner you treat it as one, the fewer turns you lose to a surprise refusal.
The mechanism: parse, expand concurrently, budget, reassemble
The whole feature lives in one module, agent/context_references.py. Reading it end to end is quick once you know the four stages, because the source comments name the failure each stage exists to prevent.
Stage one is parse_context_references. It runs a regex over your message and pulls out every reference into a ContextReference that records the kind, the target, and an optional 1-indexed line range. The regex matters because of what it does not match. A bare word preceded by a non-word or slash character is what triggers matching, so my@file:path stays literal but @file:path expands. Values are quoted when they contain spaces, and trailing punctuation like commas and periods is stripped so (see @file:x.py) does not swallow the closing parenthesis.
Stage two is where the interesting behavior lives. _expand_reference fans every reference out with asyncio.gather, because each one is independent and several @url: refs would otherwise serialize web fetches one after another. The gather preserves order, so the attached blocks arrive in the order you typed them. Each kind dispatches to its own expander: @file: and @folder: go to _expand_path_reference, the git kinds go to _expand_git_reference, @url: goes to the web extractor.
Stage three is the token budget, computed once after all expansion finishes. hard_limit is 50% of your context_length, soft_limit is 25%. If the total inlined tokens cross the hard limit the whole turn is refused and your message is returned unchanged with a warning. Crossing the soft limit appends a warning but proceeds. Both numbers come from the context length you pass in, and the call sites pass the real current context, so the limit moves with your model.
Stage four reassembles. Attached blocks go under --- Attached Context ---. Warnings go under --- Context Warnings ---. This is the part people get wrong: the @file:AGENTS.md token stays in the message text. The source comment calls the token the reference, because the desktop renders it as an inline chip, and stripping it would leave a hole in the sentence. So your message ships with the chip still in place and the content duplicated underneath. That is by design, and it has a cost: the chip counts as tokens too, on top of the content it points at.
The git layer is hardened, and the reason is an RCE class
@diff and @git:N do not just shell out to git. They shell out through two guards that a repo you clone can attack. The first is harden_git_argv, which inserts a set of no-driver diff-rendering flags right after any diff subcommand. The second is noninteractive_git_env, which rebuilds the environment so the invocation can never prompt, hang, or load a hook.
Both exist because a malicious repository can own your git config. The source cites GHSA-7x36-8jrh-v4pw, the class where a repo’s .gitattributes names an attacker-controlled external diff driver and git runs it when you render a diff. Run that with @diff and a naive git diff call and the clone executes your code. The argv hardening closes it by forcing built-in drivers, and the environment layer disables the rest.
The environment guard is the one I would have missed. noninteractive_git_env sets GIT_TERMINAL_PROMPT=0, forces GCM_INTERACTIVE=Never, drops any inherited GIT_CONFIG_* injection so ambient values cannot re-enable hooks or pagers, and pins core.sshCommand to ssh -o BatchMode=yes. That last one exists because a background git fetch can otherwise steal a prompt. The source points at issue #104591: ssh ignores stdin=DEVNULL and opens /dev/tty directly, so a private remote that needs credentials hangs forever waiting on a terminal nobody is watching. BatchMode turns that into a fast, readable failure. An agent-authenticated ssh still succeeds, and an explicit GIT_SSH_COMMAND env var in your session still wins over the config pin.
I verified the clamping on this box. @git:0 expands to git log -1 -p, @git:200 expands to git log -10 -p. The range is clamped to one through ten no matter what you type, because an eleven-commit -p dump is the kind of thing that blows the budget on its own.
The budget is where the semantics surprise you
The token accounting is not per-file, it is per-message. injected_tokens is the sum over every block that actually got inlined, and the hard check compares that sum to 50% of context. Two files that each fit comfortably can fail together when their combined inlining crosses the line, and the failure is all or nothing. Your message comes back unchanged.
The reproduction is clean. With a context length of 1000, so a 500-token hard limit, and two small Python files whose inlining sums to 924 tokens, expansion returned blocked: True with this warning:
@ context injection refused: 924 tokens exceeds the 50% hard limit (500).
Drop the same pair under a context that lets them sit below 50% and they inline. The soft limit, at 25%, adds @ context injection warning: 727 tokens exceeds the 25% soft limit (500). and proceeds anyway. So the two thresholds are not warnings versus errors by severity. One is keep going and the other is this turn will not happen.
There is a reason the tokens are counted only for what was inlined, and it is a fix that changed the failure mode you will actually hit. Older code let one oversized file poison the aggregate check and refuse the whole turn (issue #61987). Now an oversized text file does not inline at all. It becomes a note block that names the file, gives its approximate token count and size, and points the model at read_file with a narrow range. I ran a 25000-token file against a 5000-token context and got this instead of a refusal: the inliner printed a note that the file was too large to inline safely, that it stayed on disk at its full path, and that the agent should use read_file with a narrow line range or search_files to inspect only what it needed. So the hard refusal you can actually trip is not an oversized single file. It is the sum of several files that each looked reasonable alone. If your prompt comes back unexpanded and unchanged with no context attached, check the total, not the individual sizes.
The security gate fails closed, and the binary case is the interesting one
Every @file: path goes through _ensure_reference_path_allowed. It blocks exact credential files like ~/.ssh/id_ed25519 and ~/.netrc, blocks whole sensitive directories under your home (.ssh, .aws, .gnupg, .kube, .docker, .azure, .config/gh), and then falls through to the canonical read deny-list through agent/file_safety.get_read_block_error. The last check is the one that grows over time, so a path neither list mentions gets verified against it before attachment.
The interesting bit is that the module treats a failed security lookup as a block, not a pass. The source comment is blunt about why: a spurious block is recoverable, a leaked credential is not. If the deny-list cannot be checked, the attachment is refused. I ran @file:~/.ssh/id_ed25519 with the workspace widened to the home directory and got the exact rejection the docstring predicts:
@file:~/.ssh/id_ed25519: path is a sensitive credential file and cannot be attached
Path traversal is handled the same way, by defaulting the allowed root to the working directory. _resolve_path applies resolve() to whatever you type and then checks it against the allowed root, so an absolute path like /etc/passwd is rejected unless a caller explicitly widened the root. The default posture is that @ cannot escape the workspace you are in.
The binary-file path is the one that looks like a bug until you read the comment that built it. A binary file does not warn and does not refuse. It emits a note block that names the file, gives its MIME type and byte size, and tells the model the file is available on disk with a nudge to use its own tools on it. The source comment explains the original design failed: a bare not-supported warning was a dead end, the model gave up, so the block carries an actionable path instead. Here is the live output for a fake PNG blob:
note: /tmp/fake-blob.bin (application/octet-stream, 66 B) - binary file, not inlined as text.
It is available on disk at `/tmp/fake-blob.bin`. Use your tools to work with it.
That is the pattern to internalize. Every refusal in this feature is shaped to keep the model moving. The hard limit returns the message unchanged, but every per-file case hands it a path and a next step rather than a dead end.
The gotchas that make it look broken
The one that costs you the most turns is also the least documented. Context references are a CLI feature. On messaging platforms the @ syntax is not expanded by the gateway at all; the message passes through as-is, and the agent only sees files if it reaches for read_file, search_files, or web_extract itself. If you spend all day driving Hermes through Telegram, @file: is not doing what you think. The autocomplete, the line ranges, and the pre-render expansion are the interactive CLI’s behavior, and you should read those features as belonging to the terminal, not to the agent.
Line ranges are silently ignored when they are invalid. @file:src/main.py:10-25 slices lines 10 through 25 inclusive, 1-indexed, and I confirmed :4-6 on a ten-line file inlines exactly line4, line5, line6. But an out-of-range or malformed range is dropped and the full file is returned, with no warning. Ask for line 200 of a 50-line file and you get the whole file in context, which is the opposite of what you asked for. If the inlined token count looks wrong, your range silently widened to everything.
The expanded flag has a meaning you can trip on. expanded=True means the message accumulated blocks or warnings, not that everything fully inlined. An oversized file, a binary file, and a real inline all set it, and a failed reference that only produced a warning sets it too. Do not read expanded as “all references resolved.”
How to verify it is actually working
The feature is self-contained enough to prove from a script with no agent involved. Against a git repo, run the preprocessing directly and read the result:
from agent.context_references import preprocess_context_references
r = preprocess_context_references(
"Review @file:AGENTS.md and @diff",
cwd="/path/to/repo", context_length=300000,
)
print(r.expanded, r.blocked, r.injected_tokens)
print(r.message[:800])
A working expansion returns expanded=True, blocked=False, a positive injected_tokens, and a message that carries both a fenced file block and a git diff block. To see the refusal path, shrink context_length until the injected sum crosses half of it, and the same call returns blocked=True with the message unchanged.
You can also run the module’s own suite, which I did on this box and got 17 plus 15 tests green in about a second:
python -m pytest tests/agent/test_context_references.py tests/agent/test_plugin_context_references.py -q
The plugin tests matter because they prove the extension point. register_context_reference_provider lets a plugin claim any prefix the built-ins do not, six names are reserved, and a custom provider registers autocomplete and an expand function under @<name>:. That is how the feature stops being a fixed list and becomes a hook a plugin can hang its own @issue: or @channel: prefixes on.
What changes if you start thinking in stages
Once you see it as a preprocessing pipeline, the rules stop being a list of gotchas and become consequences of one design. The budget is a single per-message check, so watch the sum. The security gate fails closed, so expect refusals around credentials and outside the workspace, and expect paths rather than dead ends. The git layer is hostile-repo aware, so know that @diff is safe to run in a clone, but that safety is earned by the argv and environment hardening, not by git being trustworthy. And the CLI-only boundary means the feature is a terminal habit, not a cross-platform one.
The next time @git: returns ten commits when you asked for two hundred, you will know it was never a bug. It is the budget deciding, before the model sees a single token, that the thing you asked for would have sunk the whole turn.
Sources
- Context references documentation (syntax, line ranges, size limits, security, platform availability): https://hermes-agent.nousresearch.com/docs/user-guide/features/context-references
- CLI autocomplete and the full reference table: https://hermes-agent.nousresearch.com/docs/user-guide/features/context-references
- The feature spec and design issue: https://github.com/NousResearch/hermes-agent/issues/682
agent/context_references.py: parse, concurrent expansion, token budget, path sandboxing, git hardening, plugin provider registry, binary and oversized-file blockshermes_cli/_subprocess_compat.py:harden_git_argvandnoninteractive_git_env(GHSA-7x36-8jrh-v4pw, issue #104591)- GHSA-7x36-8jrh-v4pw: https://github.com/advisories/GHSA-7x36-8jrh-v4pw
tests/agent/test_context_references.pyandtests/agent/test_plugin_context_references.py: 32 tests green on this run
All live outputs in this post were produced on 2026-09-14 against hermes-agent commit d62716c7 (2026-09-12), with the local agent/context_references.py verified to match main exactly. Every warning text, refusal, clamping result, line-range slice, and binary nudge above came from a real call to preprocess_context_references on this box; nothing was reconstructed from memory.