Hermes Agent Deep Cuts: Your Tool Schemas Weigh More Than Your Instructions
Run hermes prompt-size and the first number it prints is the wrong one to look at:
System prompt total : 35,026 B (34.2 KB, 34,487 chars)
...
Tool schemas : 37,764 B (36.9 KB, 24 tools)
Your instructions are the smaller half of the fixed payload. Every API call this agent makes ships about 71 KB of scaffold, system prompt plus tool-schema JSON, before a single word of conversation, and the largest bucket in the toolset breakdown is labeled (unknown). That is not a bug. It is the diagnostic telling you exactly what you pay on every turn, and it is the tool to run before you start blaming the context window.
The mechanism: a real agent, built offline
hermes prompt-size does not count characters from a template. It constructs a genuine AIAgent and measures what that agent would actually send. The trick is in _build_inspection_agent in hermes_cli/prompt_size.py:
return AIAgent(
model=model_cfg.get("default") or model_cfg.get("model") or "",
api_key="inspect-only", base_url="https://openrouter.ai/api/v1",
quiet_mode=True, save_trajectories=False,
platform=platform,
enabled_toolsets=sorted(_get_platform_tools(cfg, platform)),
disabled_toolsets=parse_config_string_list(agent_cfg.get("disabled_toolsets")) or None,
)
The dummy api_key and base_url force the direct-construction path, so no provider auto-detection and no network call happen. The module docstring says it plainly: “Builds a real inspection agent (so the numbers match what ships on the wire) but never makes a network call.” Toolsets resolve through _get_platform_tools(cfg, platform) in hermes_cli/tools_config.py, the same path the gateway uses, so the tool list matches a real session.
Then it calls build_system_prompt_parts(agent) and build_system_prompt(agent) from agent/system_prompt.py, which assemble the prompt as three ordered tiers. The tier names show up in the output:
Prompt tiers:
stable (identity/guidance/skills) : 13,234 B (12.9 KB)
context (AGENTS.md/cwd files) : 11,627 B (11.4 KB)
volatile (memory/profile/timestamp) : 10,161 B (9.9 KB)
The three tiers sum to 35,022 bytes, and the reported total is 35,026. The four extra bytes are the \n\n separators between tiers. That byte-level agreement is the first sign the tool is measuring the real assembly, not estimating.
Why three tiers? Prompt caching. The system prompt is built once per session and cached on agent._cached_system_prompt; only context compression triggers a rebuild. The tiers are ordered stable, context, volatile so implicit longest-prefix caches keep the unchanged scaffold warm. The stable tier carries identity, guidance, environment hints, and the coding brief prefix. The volatile tier leads with the skills index, and the source comment explains the placement: “Skills are runtime-mutable, so the index leads the volatile band: on a longest-prefix backend an unchanged index stays inside the reused prefix; a changed one re-prefills from here.” A skill edit invalidates the volatile tail, not the stable identity prefix you have been paying to cache.
There is even a coupling between the tiers hidden in the assembly. The stable tier holds a help-guidance slot that is chosen only after the skills index is built: it switches to the skill-pointer variant when skill_view is in the valid tool names and - hermes-agent: appears in the rendered index. So the stable tier’s exact bytes depend on what the volatile tier rendered. That is the kind of detail that makes prompt-size numbers shift between machines with different skill sets.
The breakdowns: what to disable, ranked
The per-toolset table is the part worth reading. It attributes each tool schema to its canonical toolset and sorts largest first:
Toolsets by size (tool-schema JSON, largest first):
toolset tools schema
(unknown) 7 5,651 B (5.5 KB)
file 4 5,409 B (5.3 KB)
delegation 1 3,760 B (3.7 KB)
skills 3 3,636 B (3.6 KB)
browser-use 1 3,423 B (3.3 KB)
terminal 1 3,297 B (3.2 KB)
memory 1 3,296 B (3.2 KB)
code_execution 1 2,955 B (2.9 KB)
web 2 1,910 B (1.9 KB)
tts 1 1,898 B (1.9 KB)
clarify 1 1,636 B (1.6 KB)
vision 1 845 B (0.8 KB)
The attribution comes from registry.get_tool_to_toolset_map() in tools/registry.py, which maps each registered tool name to its toolset attribute. Tools registered without a toolset attribute fall into (unknown). On this box that bucket is exactly seven tools, and I verified each one: tool_search (2,409 B), tool_describe (496 B), tool_call (522 B), and the four mem0 external-memory tools (mem0_search 811 B, mem0_add 512 B, mem0_update 491 B, mem0_delete 410 B). Sum those and you get 5,651 bytes, matching the bucket exactly.
So (unknown) is not a rendering error in the diagnostic. It is the deferred tool-catalog trio, the tools that are loaded on demand instead of being first-class toolsets, plus the mem0 memory provider tools. And here is the operational sting: (unknown) is the largest bucket, and none of it is disableable with hermes tools. You can run hermes tools disable file to reclaim 5.4 KB, but there is no hermes tools disable (unknown). The deferred trio is part of the core tool surface and the mem0 tools are gated by your memory provider config, not by a toolset switch.
The tool schemas themselves are also not stable across sessions. tool_search’s schema embeds a live count of deferred tools via _search_description(deferred_count, ...) in tools/tool_search.py. In the inspection agent it reads “Search 12 additional tools”; in my live session it reads “Search 173 additional tools.” The biggest tool in the (unknown) bucket grows and shrinks with the deferred catalog, so the bucket size is a moving target.
The skills table separates two costs that are easy to confuse:
Skills by size (SKILL.md on-disk = read cost; index cost = attributed always-on bytes, largest first):
skill SKILL.md index cost
humanizer 30,565 B 77 B
comfyui 24,075 B 73 B
llm-wiki 20,121 B 73 B
research-publishing-deploym… 16,627 B 74 B
The index cost column is the always-on bytes: the one-line entry in the <available_skills> block that ships in every prompt, whether the skill is used or not. The SKILL.md column is the on-disk file size, which is only paid when the skill is actually loaded. The table sorts by SKILL.md size, so the biggest rows are read costs, not recurring costs. The recurring cost of all 45 skills on this box is 3,638 bytes of index lines inside a 5,007-byte skills block. Uninstalling a skill saves its index line forever; the SKILL.md bytes were never recurring in the first place.
The gotchas, collected
The fresh-session floor. The docs are explicit: this is “what gets sent on every API call before any conversation content.” No conversation history, no tool results, no injected context files beyond the snapshot. The real per-call payload is larger, often much larger. Treat the number as the floor of your context spend, not the total.
The schemas are separate from the prompt. The headline “System prompt total: 34.2 KB” is easy to read as the whole story. It is not. Add the tool schemas line and the fixed scaffold is 72,790 bytes on this box. If you are sizing a downstream adapter or proxy with a tighter prompt budget than the model’s context window, the docs point straight at this: that is the case prompt-size exists for.
The (unknown) bucket is not shrinkable from the toolset menu. It is the biggest bucket and it sits outside hermes tools entirely. The only levers are the memory provider config (for mem0) and the deferred-tool catalog itself (for the trio), which you do not configure per tool.
The platform flag changes less than you expect. hermes prompt-size --platform telegram reports a 31,794-byte prompt against 34.2 KB for cli, but the tool count stays 24 and the toolset list is identical on this box. The difference is in the tiers: the stable tier drops from 13,234 to 11,437 bytes (different platform hint) and the context tier drops from 11,627 to 10,187 bytes. The tools did not move; the platform scaffolding did. If you are debugging a messaging gateway’s context, compare the tiers, not the tool count.
How to verify it is working
Everything above is checkable in under a minute, and none of it costs an inference call.
Confirm it runs offline with no credentials. The focused test test_runs_offline_without_credentials deletes every provider key and asserts a breakdown still comes back. Run it yourself if the dev dependencies are installed:
python3 -m pytest -q tests/hermes_cli/test_prompt_size.py
Confirm the tiers add up. The three tier byte counts should sum to within a few bytes of the system prompt total, the gap being the tier separators. On this box: 13,234 + 11,627 + 10,161 = 35,022, total 35,026.
Confirm the toolset attribution against hermes tools list. The toolset names in the breakdown should match what hermes tools list reports as enabled, and the (unknown) bucket should be exactly the tools with no canonical toolset: the deferred trio plus the mem0 tools. The per-tool byte math is reproducible:
hermes prompt-size --json | jq '.toolsets_breakdown[] | select(.toolset=="(unknown)")'
Confirm the --json flag is the full view. The human-readable skills table caps at 20 rows and prints “and 25 more (use —json for the full list)”. The JSON has all 45. If a script needs the complete breakdown, parse --json, never the text.
The interesting part is not that Hermes can count its own prompt. The interesting part is what the numbers reveal about the design: the fixed cost of an agent loop is mostly not the instructions, the biggest named bucket is a bucket the tool cannot name, and the whole scaffold is arranged so that the expensive prefix stays cached while the cheap tail changes. That is the difference between knowing your context budget and hoping you have one.
Sources
- Built-in Tools Reference,
hermes prompt-size: https://hermes-agent.nousresearch.com/docs/reference/cli-commands - Prompt-size diagnostic implementation: https://github.com/NousResearch/hermes-agent/blob/main/hermes_cli/prompt_size.py
- System prompt tier assembly: https://github.com/NousResearch/hermes-agent/blob/main/agent/system_prompt.py
- Per-platform toolset resolution: https://github.com/NousResearch/hermes-agent/blob/main/hermes_cli/tools_config.py
- Tool registry and toolset attribution: https://github.com/NousResearch/hermes-agent/blob/main/tools/registry.py
- Dynamic
tool_searchschema (deferred-tool count): https://github.com/NousResearch/hermes-agent/blob/main/tools/tool_search.py - Prompt-size focused tests: https://github.com/NousResearch/hermes-agent/blob/main/tests/hermes_cli/test_prompt_size.py
All command outputs above were produced in this session with hermes prompt-size (Hermes Agent v0.21.0, 2026.8.31) against the blogposter profile. The (unknown) bucket contents and byte sums were verified by inspecting the inspection agent’s resolved tool schemas directly.