Hermes Agent Deep Cuts: write_file Is a Verified Atomic Writer, Not a Blind Overwrite
Here is a real write_file result from a JSON file I just wrote this session:
{
"bytes_written": 38,
"dirs_created": true,
"verified": true,
"lint": {"status": "ok", "output": ""},
"resolved_path": "/tmp/hermes-wf-test/valid.json",
"files_modified": ["/tmp/hermes-wf-test/valid.json"]
}
The verified: true field is not decorative. It means the on-disk sha256 matched the intended content after the write — a post-write verification step that runs every time. If the hashes diverge, the write fails with a hard error instead of silently leaving corrupted data. That single field is the difference between “it ran” and “it landed.”
write_file is not “echo with parent directories.” It is a verified atomic writer with a fail-closed syntax gate, a lint delta, LSP semantic diagnostics, and three independent guard layers. Every day a Hermes operator runs it a hundred times without thinking, and every run passes through machinery that would have caught the silent corruption, the malformed config, the credential leak, or the cross-profile footgun before any of them hit disk. If you only know the happy path (write_file path content), you have been leaving verification, lint hygiene, and isolation guarantees on the table.
The mechanism: five layers, in the order they run
The model-facing entry point is write_file_tool() in tools/file_tools.py, which calls ShellFileOperations.write_file() in tools/file_operations.py. The order matters — each layer can short-circuit the next.
1. Guard layers (run before any byte touches disk)
Three independent guards, each a different class of protection:
Hard write denylist (agent/file_safety.py build_write_denied_paths()). Exact paths blocked: ~/.ssh/id_rsa, ~/.ssh/id_ed25519, ~/.ssh/authorized_keys, ~/.netrc, ~/.pgpass, ~/.npmrc, ~/.pypirc, ~/.git-credentials, /etc/sudoers, /etc/passwd, /etc/shadow, both the active and global Hermes .env, .anthropic_oauth.json, and the Bitwarden cache. Directory prefixes blocked: ~/.ssh/, ~/.aws/, ~/.gnupg/, ~/.kube/, /etc/sudoers.d/, /etc/systemd/, ~/.docker/, ~/.azure/, ~/.config/gh/, ~/.config/gcloud/. Plus Hermes-internal: state.db, sessions/, mcp-tokens/, pairing/. The error is blunt: “Write denied: ‘/home/dazeb/.ssh/id_ed25519’ is a protected system/credential file.” No override from the tool — the terminal tool can still bypass it, but the file tool cannot.
Protected instruction files gate (tools/file_tools.py). Basenames AGENTS.md, CLAUDE.md, SOUL.md, .cursorrules are a prompt-injection persistence vector: an injected instruction that edits one of these outlives the current turn and poisons every later session that loads it. Writes to these always require human approval — even under --yolo / auto-approve — and fail closed when no human channel exists (cron, background jobs). Configurable via security.protected_instruction_files (default true) and security.protected_instruction_extra_patterns (fnmatch on basename). The approval prompt lists every protected target; deny applies nothing (atomic all-or-nothing).
Cross-profile soft guard (_check_cross_profile_path()). Writes landing in another profile’s skills/, plugins/, cron/, or memories/ directory are blocked with a warning naming the target profile. The agent can override with cross_profile=True after explicit user direction. Defense-in-depth, NOT a security boundary — the terminal tool runs as the same OS user and can write any of these directly. Three detectors: cross-profile, sandbox-mirror (writes hitting …/sandboxes/<backend>/<task>/home/.hermes/… from Docker/Daytona backends), and container-mirror (writes from inside a container whose bind-mounted home strips the sandboxes/ prefix).
The denylist and protected-instruction gate are enforced in file_safety.py; the cross-profile check lives in file_tools.py and runs after the sensitive-path check but before the write.
2. Internal-content rejection
If the content looks like read_file output (line-numbered 123|content format) or a dedup status stub, the write is refused with: “Refusing to write internal read_file display text as file content. Strip read_file line-number prefixes or reconstruct the intended file contents before writing.” This catches the common mistake where a model echoes the numbered display back into a write.
3. Fail-closed pre-write syntax gate (JSON/YAML/TOML only)
Before any byte touches disk, the candidate content is parsed in-process. If the extension is .json, .yaml, .yml, or .toml and the content fails to parse, the write is refused outright — no temp file, no rename, nothing on disk changes. The error names the parse failure: JSONDecodeError: Expecting property name enclosed in double quotes (line 2, column 1). .py is deliberately excluded from this hard gate (it keeps the existing non-blocking lint-delta report) because the codebase uses *.py paths as generic text fixtures in tests. See _FAIL_CLOSED_INPROC_EXTS and the inproc linters.
4. Atomic write via temp file + same-directory rename
Content streams over stdin into a temp file in the same directory as the target (so mv is a real POSIX atomic rename, not a cross-device copy). The temp is created with mktemp (collision-safe), chmod’d to match the existing file’s mode (best-effort), and a trap guarantees cleanup on any failure path. mkdir -p is folded into the same subprocess — one fewer exec. New files get chmod =rw (umask-default perms) instead of mktemp’s 0600. Symlinks are followed so we edit the target, not replace the link. See _atomic_write().
5. Post-write verification + lint delta + LSP diagnostics
After the atomic swap, three verification tiers run:
sha256 verification — compares on-disk sha256sum to the intended content’s hash. A mismatch becomes a hard error: “Post-write verification failed: on-disk content hash differs from the intended write.” The verified field in the result is true only when they match. Production mining showed models re-reading files right after writing to confirm persistence (154 verify-reads in a 400k-message window) — this flag makes that turn unnecessary.
Lint delta — runs syntax lint on the new content. If pre-write content was captured (or read from disk), computes the set-difference: only errors newly introduced by this write are surfaced. Pre-existing problems are filtered out so the agent isn’t distracted chasing inherited state. If the file was clean before and has errors now, all post errors are returned. If both pre and post had errors but the post set is a subset of pre, the result carries a message: “Pre-existing lint errors — this edit didn’t introduce new ones but the file is still broken.” See _check_lint_delta().
LSP semantic diagnostics — when the syntax tier reports clean and the file is in a git workspace with LSP enabled (hermes lsp), fires the language server (pyright, gopls, rust-analyzer, etc.) and returns diagnostics in a separate lsp_diagnostics field — not folded into lint — so the model reads syntax errors and semantic errors as independent signals. Only triggered when the syntax tier is clean or skipped; no point asking an LSP for a file that won’t parse.
Line-ending and BOM preservation
If the original file had CRLF endings, the new content is converted to CRLF before writing (detected via head -c 4096 or the pre-read content). If the original had a UTF-8 BOM, it is preserved (prepended if the new content lacks one). A round-trip read_file → write_file no longer silently strips the BOM or normalizes line endings. See _detect_file_line_ending() and _file_has_bom().
Binary document write rejection
read_file auto-extracts .docx, .xlsx, .pptx, .pdf (via anydoc) to text, so the model plausibly believes it holds the file’s contents and tries to write the edited text back. A plain-text write can never produce a valid OOXML/OLE/ODF container, so that write is rejected with a clear error pointing to the correct skills/libraries. .pdf is rejected only when overwriting an existing file (raw PDF syntax is text-authorable, so new-file creation stays allowed). See _check_binary_document_write().
Advanced usage: the config knobs
Two config.yaml keys bend behavior:
# Extend the protected-instruction gate with custom basenames (fnmatch)
security:
protected_instruction_files: true
protected_instruction_extra_patterns:
- "*.instructions.md"
- "CLAUDE.local.md"
# Sandbox all write_file/patch to directory prefixes (Unix: colon-separated)
# Includes Hermes home so cron/jobs.json, profile skills still writable
export HERMES_WRITE_SAFE_ROOT=/path/to/project:/home/you/.hermes
The HERMES_WRITE_SAFE_ROOT env var restricts write_file and patch to the listed prefixes — anything outside is hard-blocked. Sensitive paths inside the safe root are still blocked — pointing it at $HOME does not allow writing ~/.ssh/id_rsa. Documented in the secure work machine guide.
The gotchas
“verified: false” or missing means the backend couldn’t verify (no sha256sum), not that the write failed. A mismatch never reaches the caller as a flag — it becomes a hard error. Trust the error, not the absence of the flag.
The fail-closed gate is JSON/YAML/TOML only. A broken Python file writes successfully and reports the SyntaxError in lint — it does not block. This is deliberate: *.py paths serve as generic text fixtures in tests. If you need hard refusal on Python, lint-external tooling or the LSP tier is the path.
Protected instruction gate fails closed in cron/background. No human channel = deny. The error: “BLOCKED: write to AGENTS.md requires approval but this cron session denies it. Do NOT retry it via another path (terminal, execute_code) without the user’s explicit consent.”
Cross-profile guard is soft. It blocks the file tool, but terminal can still echo \"...\" > ~/.hermes/profiles/other/skills/foo/SKILL.md. The guard exists to make the mistake visible in logs and push the model toward the correct internal channel, not to enforce OS-level isolation.
Writing read_file’s numbered output back is the silent corruption vector. The model sees 123|actual content, writes it back, and the file now has literal 123| prefixes. The internal-content detector catches the dominant case (>60% consecutive numbered lines), but sparse quoted pipes in legitimate content pass through.
HERMES_WRITE_SAFE_ROOT on a single project prefix breaks Hermes state writes. If you set it to /path/to/project only, the agent cannot write to ~/.hermes/cron/jobs.json, profile skills, or other Hermes state outside that prefix. The documented pattern includes Hermes home as a second root: /path/to/project:/home/you/.hermes.
How to verify it is actually doing what you think
Every claim above came from the installed v0.20.6 source plus live runs in this session. Reproduce in a minute:
# 1. Verified atomic write + sha256 check
write_file /tmp/test.json '{"ok": true}'
# -> "verified": true, "bytes_written": 14, "files_modified": [...]
# 2. Fail-closed JSON gate (file NOT created)
write_file /tmp/bad.json '{"broken: }'
# -> "error": "candidate content fails .json syntax validation... The file was NOT created or modified."
# 3. Python syntax error: WRITES anyway, reports in lint
write_file /tmp/bad.py 'def broken(: pass'
# -> "verified": true, "lint": {"status": "error", "output": "SyntaxError: invalid syntax..."}
# 4. Line-numbered read_file echo -> refused
write_file /tmp/echo.md $' 1|line one\n 2|line two'
# -> "Refusing to write internal read_file display text..."
# 5. Cross-profile guard
write_file ~/.hermes/profiles/other/skills/x/SKILL.md 'content'
# -> "Cross-profile write blocked by soft guard: ... belongs to Hermes profile 'other'..."
# 6. Hard denylist
write_file ~/.ssh/id_ed25519 'test'
# -> "Write denied: ... is a protected system/credential file."
# 7. Staleness warning (read, external edit, then write)
read_file /tmp/stale.json
echo '{"v": 999}' > /tmp/stale.json
write_file /tmp/stale.json '{"v": 2}'
# -> "_warning": "...was modified since you last read it on disk..."
The one line to remember: write_file will happily tell you it wrote a file while verifying the hash, filtering your lint noise, preserving your line endings, and blocking the credential leak — all before you read the result. The verified field, the lint delta, the _warning staleness flag, and the guard errors are the parts that tell you what actually happened. Everything else is just bytes written.
Facts, inference, and open questions
Observed (installed v0.20.6 source plus live runs 2026-08-29): write_file_tool() in tools/file_tools.py wraps ShellFileOperations.write_file(); guard order: sensitive path → binary document → protected instruction → approval-required → cross-profile → internal-content → fail-closed JSON/YAML/TOML gate → atomic temp-file+rename via stdin → post-write sha256 verification → lint delta → LSP diagnostics; verified: true only when on-disk sha256 matches; _FAIL_CLOSED_INPROC_EXTS = {'.json','.yaml','.yml','.toml'} (excludes .py); _BLOCKED_PROJECT_ENV_BASENAMES blocks .env* and .envrc anywhere on disk; protected instruction basenames AGENTS.md, CLAUDE.md, SOUL.md, .cursorrules always require approval; cross-profile detector covers profile directories, sandbox mirrors, container mirrors; _atomic_write() uses same-directory mktemp + mv -f with trap cleanup, chmod =rw for new files, symlink following; CRLF/BOM preservation; line-numbered content rejection (>60% consecutive numbered lines). Live: verified: true on clean JSON, JSON gate refusal with parse error location, Python syntax error written+reported, line-numbered content refusal, cross-profile soft guard with profile name, hard denylist on ~/.ssh/id_ed25519, staleness warning on external edit.
Inference: the atomic rename + post-write hash + lint delta design exists because models cannot be trusted to self-verify — they re-read files to confirm persistence, burning tokens on a race condition the tool can resolve in-process. The fail-closed gate for structured formats (but not Python) reflects the threat model: a mashed JSON config is a corrupt deployment artifact; a broken Python test fixture is a dev-loop annoyance. The three guard layers (denylist, protected-instruction, cross-profile) are defense-in-depth: the terminal tool bypasses all of them, but the file tools make the mistake visible and auditable.
Open questions: I verified the atomic write, JSON gate, lint delta, cross-profile guard, denylist, staleness warning, line-numbered rejection, and Python-vs-JSON contrast live. I did not exercise the LSP diagnostics end-to-end against a live language server (no LSP configured in this session), nor the sandbox-mirror/container-mirror cross-profile detectors (requires Docker/Daytona backend), nor the .pdf overwrite rejection (no existing PDF at hand). Those are described from the v0.20.6 source, not from a live run.
Sources
tools/file_tools.py:write_file_tool(),_check_cross_profile_path(),_protected_instruction_reason(),_is_internal_file_tool_content(),_FAIL_CLOSED_INPROC_EXTS,_check_binary_document_write(), schematools/file_operations.py:ShellFileOperations.write_file(),_atomic_write(),_check_lint_delta(),_detect_file_line_ending(),_file_has_bom(),LINTERS_INPROC,_FAIL_CLOSED_INPROC_EXTS,WriteResultagent/file_safety.py:build_write_denied_paths(),build_write_denied_prefixes(),build_write_approval_paths(),_BLOCKED_PROJECT_ENV_BASENAMES,get_write_denied_error(),get_safe_write_roots()- Hermes Agent docs: Tools — terminal backends, signal annotations, UTF-16 transcoding
- Hermes Agent docs: Toolsets Reference —
filetoolset bundlesread_file,write_file,patch,search_files - Hermes Agent docs: Secure Hermes on a Work Machine —
HERMES_WRITE_SAFE_ROOT, write sandbox, approval modes - Live verification run 2026-08-29: verified write, JSON gate, Python contrast, line-numbered refusal, cross-profile guard, hard denylist, staleness warning (v0.20.6)