Hermes Agent Deep Cuts: The Insights Report Counts Messages You Already Compressed Away
Part of the Hermes Agent: Deep Cuts series

Hermes Agent Deep Cuts: The Insights Report Counts Messages You Already Compressed Away

$ hermes insights --days 30

  📋 Overview
  ────────────────────────────────────────────────────────
  Sessions:          118           Messages:        9,212
  Tool calls:        5,012         User messages:   195
  Input tokens:      19,299,549    Output tokens:   2,867,689
  Total tokens:      395,748,930
  Active time:       ~22h          Avg session:     ~12m
  Avg msgs/session:  78.1

  💰 Cost
  ────────────────────────────────────────────────────────
  Estimated:          ~$1.39
  Unknown:            90 session(s) (no pricing data)

  🔧 Top Tools
  ────────────────────────────────────────────────────────
  Tool                            Calls        %
  terminal                        6,493    78.4%
  read_file                         433     5.2%
  search_files                      209     2.5%
  web_extract                       168     2.0%

Add up the two token lines and you get 22,167,238. The next line says 395,748,930. Add up the tool table and you get 8,285; the Overview says 5,012. Two headline totals disagreeing with their own detail sections by 17x and 1.65x, on one screen, from one command.

Query the same 30-day window straight out of the store and both gaps close:

$ python3 -c "..."
(19299549, 2867689, 373581692, 0)
$ python3 -c "..."
counter: 5012
tool rows: [(0, 3179), (1, 5009)]

The first row is input, output, cache read, cache write. Input and output match the printed lines exactly, cache read is the 373 million missing from the sum, and the fourth value is zero on every row in this store. The second shows 5,009 tool rows still active and 3,179 that compaction archived but the report still counted.

hermes insights runs a set of SQL queries over your own transcript rows in state.db and formats the result. It makes three decisions worth knowing before you quote a number from it: what it sums, which rows it reads, and what it refuses to claim.

The 94% lane the Overview never prints

The Overview computes its total as four lanes added together. From _compute_overview in agent/insights.py:

total_input, total_output, total_cache_read, total_cache_write = (sum(int(r.get(k) or 0) for r in rows) for k in _TOKEN_KEYS)
total_tokens = total_input + total_output + total_cache_read + total_cache_write

_TOKEN_KEYS is ("input_tokens", "output_tokens", "cache_read_tokens", "cache_write_tokens"). The terminal formatter prints three of those: input, output, and the total. Cache read and cache write never appear.

That is the whole 17x. On this box, 373,581,692 of the 395,748,930 total were prompt cache reads, 94.4% of everything the report counted, and the printed block shows you the 5.6% that was not. The arithmetic is exact, not approximate: 19,299,549 + 2,867,689 + 373,581,692 + 0 = 395,748,930.

Why that matters beyond tidiness: cache reads bill at a small fraction of input on most providers, so the aggregate measures throughput rather than spend. Treating it as a cost driver leads you to optimize the wrong thing, though optimizing it away is not the goal either, since a high cache-read ratio is what a well-behaved long session produces. What you can read from it is the shape of your workload: this profile’s 30 days ran 118 sessions that pushed 373 million tokens of cached prompt back through the model, because cron jobs keep re-sending a large stable system prompt with a small tail.

The per-lane numbers do exist elsewhere. A one-shot run writes them to a JSON sidecar with hermes -z "<prompt>" --usage-file report.json, and hermes chat -q ... --format stream-json emits a terminal result record whose tokens object carries input, output, total, cache_read, and cache_write. Those two surfaces are what a cost pipeline should parse, since scraped terminal output loses the cache lanes entirely.

The session that compaction rewrote

The tool-call gap is more interesting because the two totals are not the same measurement at all.

The Overview line is a sum of one column:

total_tool_calls = sum(s.get("tool_call_count") or 0 for s in sessions)

tool_call_count is a maintained counter on the sessions row. The recompute path in hermes_state_messages.py, _active_transcript_counts, counts only active = 1 rows:

rows = conn.execute("SELECT tool_calls FROM messages WHERE session_id = ? AND active = 1", (session_id,)).fetchall()
return len(rows), sum(_tool_calls_len(row[0], scalar=1) for row in rows)

The Top Tools table comes from a different query with a different rule. _get_tool_usage counts two representations of the same calls, tool result rows carrying tool_name (what the gateway writes) and assistant rows carrying a tool_calls JSON blob (what the CLI leaves behind), reconciles them per (session, tool) pair by taking the maximum, and sums across all pairs. Its SQL has no active filter at all:

SELECT m.session_id, m.tool_name, COUNT(*) as count
  FROM messages m JOIN sessions s ON s.id = m.session_id
 WHERE s.started_at >= ? AND m.role = 'tool' AND m.tool_name IS NOT NULL
 GROUP BY m.session_id, m.tool_name

So the Overview measures the transcript the agent still carries, and the Top Tools table counts every row that still exists in the window, including the ones a /compress archived.

In this window that is 3,179 rows, and every one of them belongs to a single session: the 4 hour 40 minute run the report itself flags in Notable Sessions as the longest and the busiest. For that session the numbers reconcile exactly. The report lists it at 813 tool calls. A row scan finds 3,992. The difference, 3,179, is precisely the archived count. The counter kept the active transcript; the table counted the archive too.

Nothing here is corrupt. The archived rows are real, they are still full-text searchable, and a tool table that ignored them would understate what the agent did. But the practical consequence is that two numbers on the same screen answer different questions, and the lower one is not “the truth” you should be quoting. If you want loop iterations for a cost or capacity model, take the Overview. If you want a fingerprint of what the agent touched, take the table. Comparing them to each other will only ever produce a fake regression.

The reconciliation also explains why the table can sum above either raw count: the union of (session, tool) pairs includes pairs that appear in only one representation, so 8,285 came out of 8,187 tool rows plus 8,191 assistant call blobs. And one caveat on precision: this box’s store is live, cron jobs write to it while you read, so a figure taken minutes apart from another can drift by a few calls. The snapshot above was taken in a single pass.

Cost arrives in three buckets

hermes insights will not collapse spend into one number, which is the most honest thing about it. The full cost block on a 30-day window:

  💰 Cost
  ────────────────────────────────────────────────────────
  Estimated:          ~$1.39
  Unknown:            90 session(s) (no pricing data)

Three buckets live in the code, and the formatter only prints the non-zero ones: estimated, included (subscription, no provider invoice), and unknown. The engine classifies each session through _estimate_cost, which routes to estimate_usage_cost and returns a status alongside the amount, and has_known_pricing decides whether the model plus billing provider plus base URL has a price at all. Sub-cent aggregates render at four decimal places through format_cost_label rather than collapsing to ~$0.00.

Group the same window by status and 90 sessions show up as unpriced:

[('unknown', 86), ('estimated', 28), ('(null)', 4)]

86 plus 4 is the 90 the report prints. That means the ~$1.39 is a floor over 28 of 118 sessions, not a bill. On this profile the unpriced sessions are exactly the ones you would expect: a local weights file addressed by path, and custom OpenAI-compatible endpoints that no pricing table knows about. A report that turned those into $0.00 would be lying by omission, and the reason the bucket exists is the bug class its own comments call out.

For real provider-side spend, hermes usage prints the /usage block without starting a session, and it needs the provider to answer. Verified here: on a custom provider it printed No account usage available for provider 'custom:api.buzzgw.com': no credential is configured for it, the provider has no usage endpoint, or the fetch failed. and exited 1, while hermes usage --provider openrouter returned a credits balance and exited 0. Local accounting always works. Remote accounting depends on the vendor.

What --source filters, and what the docs leave out

--source filters the sessions.source tag. The tags present in this profile’s store, with counts across all time: cron 246, a2a 28, cli 12, telegram 3, tui 2, desktop 1. That is a useful operational view on a box that mostly runs unattended, and it is how you separate what your cron jobs did from what your own typing did.

Two edges worth knowing before you wire this into a monitor:

An unknown tag is not an error. hermes insights --source teams --days 30 prints No sessions found in the last 30 days (source: teams). and exits 0. A typo is indistinguishable from an idle platform.

--days 0 is legal and always empty, because the cutoff is now. It prints No sessions found in the last 0 days.

The CLI reference documents exactly two options for this command, --days and --source, and nothing else. The slash command reference lists /insights [days] for messaging. The gateway handler for that command accepts more than the docs show: a bare number, --days N, and --source S, so /insights 7 --source telegram works on Telegram even though the option is undocumented there. The same handler normalises Unicode dashes before parsing, because Telegram clients on some platforms convert -- into an en dash in transit. If you have ever typed --days into a chat window and watched it do nothing, that is the failure mode being patched around.

The profile question bites harder. The gateway resolves which state.db to read from a contextvar at call time, and the handler deliberately hops to its executor thread with the context preserved. The comment in gateway/slash_commands_status.py is blunt about why: a plain executor hop starts with an empty context and would read the default profile’s database. The practical upshot is that /insights in a Telegram chat reports the store of the profile that gateway serves, not whichever profile you were thinking about. From a shell you pick explicitly: hermes -p <profile> insights --days 7.

One more property that makes this command safe to schedule: it opens the store read-only. cmd_insights constructs SessionDB(read_only=True), so the report never writes to your session history. It also calls flush_token_counts() before computing, which drains an asynchronous token-accounting queue whose deltas are applied by a background writer thread. That means hermes insights waits for queued usage to land, and a hand-written query against state.db during a live turn can legitimately read a smaller number than the CLI reports. I did not reproduce that drift here, since every read above was taken while the box was idle between cron runs, so treat it as a source-level mechanism rather than a measured effect.

How to check the numbers yourself

Three probes, all read-only, all run against the store the command itself reads. This box has no sqlite3 binary installed, so the interpreter is the Hermes venv Python; substitute sqlite3 -readonly "$HERMES_HOME/state.db" if your box has the CLI.

Token lanes, the split behind the missing 373 million:

$ /home/dazeb/.hermes/hermes-agent/venv/bin/python3 -c "
import os,sqlite3,time
c=sqlite3.connect('file:'+os.path.join(os.environ['HERMES_HOME'],'state.db')+'?mode=ro',uri=True)
print(c.execute('SELECT SUM(u.input_tokens), SUM(u.output_tokens), SUM(u.cache_read_tokens), SUM(u.cache_write_tokens) FROM session_model_usage u JOIN sessions s ON s.id=u.session_id WHERE s.started_at>=?',(time.time()-30*86400,)).fetchone())
"
(19299549, 2867689, 373581692, 0)

The counter against the row scan, the split behind the tool-call gap:

$ /home/dazeb/.hermes/hermes-agent/venv/bin/python3 -c "
import os,sqlite3,time
c=sqlite3.connect('file:'+os.path.join(os.environ['HERMES_HOME'],'state.db')+'?mode=ro',uri=True)
t=time.time()-30*86400
q='SELECT m.active, COUNT(*) FROM messages m JOIN sessions s ON s.id=m.session_id WHERE s.started_at>=? AND m.role=? AND m.tool_name IS NOT NULL GROUP BY 1'
print('counter:', c.execute('SELECT SUM(tool_call_count) FROM sessions WHERE started_at>=?',(t,)).fetchone()[0])
print('tool rows:', c.execute(q,(t,'tool')).fetchall())
"
counter: 5011
tool rows: [(0, 3179), (1, 5007)]

And the cost buckets:

$ /home/dazeb/.hermes/hermes-agent/venv/bin/python3 -c "
import os,sqlite3,time
c=sqlite3.connect('file:'+os.path.join(os.environ['HERMES_HOME'],'state.db')+'?mode=ro',uri=True)
print(c.execute('SELECT COALESCE(cost_status,\'(null)\'), COUNT(*) FROM sessions WHERE started_at>=? GROUP BY 1 ORDER BY 2 DESC',(time.time()-30*86400,)).fetchall())
"
[('unknown', 86), ('estimated', 28), ('(null)', 4)]

Those runs were taken a few minutes apart, which is why the counter reads 5,011 here against 5,012 in the snapshot above and 5,009 active rows. The store is live and cron jobs write to it while you read, so the last digit moves. Read the ratio instead.

The dashboard exposes the same ground through HTTP: GET /api/analytics/usage and GET /api/analytics/models in hermes_cli/web_routers/analytics.py. The models route differs from the CLI in one way worth knowing: it derives its per-model list from the session counters and then folds in the auxiliary usage rows, because auxiliary calls write to session_model_usage and never touch the session counters. Models that only ever appear through an auxiliary task get their own entry rather than being dropped.

Auxiliary work shows up as models you never picked

That reconciliation matters because the Models Used table is not a list of what you chose. In this window, 21 rows in session_model_usage carry a task label: compression at 1,371,427 tokens across 36 calls, background_review at 619,808, plus title_generation and vision rows. All-time task labels in this store are approval (54 rows), title_generation (29), background_review (14), compression (8), and vision (7).

So a model you never selected can appear in the table legitimately, and the aggregate in the Models Used column is wider than the sessions counters feeding the Overview. In an earlier snapshot the sessions table held 20,719,635 input plus output tokens while the model table held 22,114,269, and the difference is auxiliary spend the session counters are not supposed to know about. When a compression model and a title model are cheap flash models, that gap is small and harmless. When one of them is your main model, it is the line that shows you why.

What to do with the numbers

Everything above comes from one design choice: hermes insights reports what it measured and where, instead of presenting one canonical number and hoping you never compare it to anything. It prints the same value twice under two different definitions and lets you notice. It refuses to price the models it cannot price, and it tells you how many sessions that was. It sums cache reads into the aggregate it builds from and hides them from the headline, which is exactly the kind of gap you find by reading the source rather than trusting the total.

Use it as a query surface. Know that the token total is a throughput figure, that the tool table counts rows compaction already removed from your live transcript, and that the estimated cost is a floor. Then the 17x and the 65% stop looking like bugs and settle into the two measurements they always were.

Sources

  • agent/insights.py: _TOKEN_KEYS, _compute_overview (the four-lane total_tokens formula and total_tool_calls), _get_tool_usage (the tool_name and tool_calls reconciliation, the per-session max, the SQL with no active filter), _get_model_usage and _compute_model_breakdown (per-model rows from session_model_usage, residual reconciliation, reasoning_tokens excluded from total_tokens), _cost_lines and the three cost buckets, _compute_activity_patterns, format_terminal and format_gateway
  • hermes_cli/subcommands/insights.py: the entire option surface, --days (default 30) and --source
  • hermes_cli/main_agent_cmds.py: cmd_insights, the SessionDB(read_only=True) open and the swallowed-exception exit path
  • hermes_state_usage.py: flush_token_counts and the _token_writer_loop that applies queued token deltas
  • hermes_state_messages.py: _active_transcript_counts (the active = 1 recompute), _bump_session_counters, _tool_calls_len
  • gateway/slash_commands_status.py: _handle_insights_command, the Unicode-dash normalisation, the _run_in_executor_with_context profile-context comment
  • hermes_cli/web_routers/analytics.py and hermes_cli/web_server_profiles.py: GET /api/analytics/usage, GET /api/analytics/models, _aux_usage_rows, _merge_aux_into_by_model
  • Hermes Agent documentation, CLI Commands Reference (hermes insights [--days N] [--source platform], and hermes usage for account rate-limit windows) and Slash Commands Reference (/insights [days] in messaging, --days N and --source absent from the documented form)
  • PR #552, the /insights feature: the original report scope (sessions, messages, tokens, costs, models, platforms, activity patterns) in the CLI, slash, and gateway surfaces
  • Live measurements on the authoring box, 2026-09-27, Hermes Agent v0.21.4 (upstream a53b42dd), profile blogposter, against state.db (196 MB, 118 sessions in the 30-day window): hermes insights --days 30, --days 1, --days 0, and --source teams; hermes usage and hermes usage --provider openrouter; read-only Python probes for the token lanes, the active/archived tool-row split, the session counter reconciliation, and the cost-status buckets. The session_model_usage task labels and auxiliary token totals come from the same store on the same date.

Keep reading