Hermes Agent Deep Cuts: Turn a Session into a Trace Without Uploading It
Part of the Hermes Agent: Deep Cuts series

Hermes Agent Deep Cuts: Turn a Session into a Trace Without Uploading It

$ mkdir -p "$HOME/hermes-trace-demo"
$ hermes sessions export --format trace --session-id "$SESSION_ID" "$HOME/hermes-trace-demo/"
Exported 1 session trace to /home/dazeb/hermes-trace-demo/cron_5032b7d71ac3_20260930_024543.trace.jsonl

$ python3 - "$HOME/hermes-trace-demo/$SESSION_ID.trace.jsonl" <<'PY'
import json, sys
rows = [json.loads(line) for line in open(sys.argv[1])]
print("events:", len(rows))
print("assistant messages with tool use:", sum(
    any(block.get("type") == "tool_use" for block in row.get("message", {}).get("content", []))
    for row in rows
))
PY
events: 86
assistant messages with tool use: 40

That export produced 86 trace events, including 40 assistant events with tool calls. It stayed on this machine. The same command with --upload sends the trace to a Hugging Face dataset. Those two words change the data boundary, not the format.

Hermes already stores the conversation as OpenAI-shaped messages in SQLite. Trace export converts those records into the Claude Code JSONL shape the Hugging Face Agent Trace Viewer understands. You can inspect a tool-heavy session in the viewer without rebuilding the agent run or making another model call.

Conversion, not replay

The implementation lives in agent/trace_upload.py. The trace builder walks the stored conversation in order and emits one JSONL record for each non-system message. Assistant text becomes text blocks. Assistant tool calls become tool_use blocks, with the function arguments parsed back into JSON. Tool results are represented as user records containing tool_result blocks keyed to the tool call ID.

Each record gets a fresh UUID and points to its predecessor with parentUuid. Hermes omits system messages. It also replaces image payloads with [image omitted], since this viewer represents text turns and not base64 image content. That is a useful conversion boundary to understand: this is a readable trace of text and tool flow, not a byte-for-byte session backup.

The output is created locally by default:

hermes sessions export \
  --format trace \
  --session-id "$SESSION_ID" \
  "$HOME/trace-review/"

For a single session, the directory form writes a default-named file ending in .trace.jsonl. If no session ID or filters are supplied, the trace exporter selects the most recently active session. I prefer specifying the ID. “Most recent” can be a cron or a one-shot run, not the interactive conversation you meant to inspect.

Upload is a separate decision

The upload path calls Hugging Face’s API and creates or reuses a dataset named hermes-traces under the authenticated account. It writes sessions/<session-id>.jsonl. By default the dataset is private, and --public changes that explicitly.

hermes sessions export \
  --format trace \
  --session-id "$SESSION_ID" \
  --upload

The CLI checks for HF_TOKEN, HUGGINGFACE_HUB_TOKEN, HUGGING_FACE_HUB_TOKEN, or HUGGINGFACE_TOKEN. The account token needs write access. A local export does not need that credential or the huggingface_hub package; upload does.

Redaction is enabled by default for trace exports, both local and uploaded. The code forces the sensitive-text redactor on each text body and serialized tool argument. If redaction throws an error or produces invalid JSON for tool arguments, the export refuses to continue instead of falling back to raw content. --no-redact disables that protection. Treat it as an explicit escape hatch for a trace you have reviewed, not as a fix for a redaction error.

The redactor is not a privacy review. Paths, private source code, personal data, and non-secret prompt content can still be sensitive. Hugging Face’s own trace documentation warns that traces may include prompts, tool inputs, command output, screenshots, and secrets. Keep the dataset private unless you have reviewed the actual file and intend to publish it.

Make trace review repeatable

Export locally first, inspect the structure without dumping message bodies into a shared terminal log, then upload only the selected session. For a quick structural check:

python3 - "$HOME/trace-review/$SESSION_ID.trace.jsonl" <<'PY'
import collections
import json
import sys

rows = [json.loads(line) for line in open(sys.argv[1], encoding="utf-8")]
print("events:", len(rows))
print("record types:", dict(collections.Counter(row["type"] for row in rows)))
print("tool uses:", sum(
    block.get("type") == "tool_use"
    for row in rows
    for block in row.get("message", {}).get("content", [])
))
print("unlinked roots:", sum(row["parentUuid"] is None for row in rows))
PY

The normal result has one root with a null parentUuid; later rows chain back through preceding UUIDs. Tool results should point to their corresponding tool_use IDs. The check tells you whether the conversion produced a coherent event stream, not whether every secret was removed or whether the transcript is safe to publish. Open the file and review its content before uploading it publicly.

For repeat audits, use --session-id with a saved ID and keep the output path outside the repository. Avoid putting the trace itself in shell history, build logs, or a public artifact store. It is session content, even when the renderer calls it a trace.

The gotcha is the format name

trace means Claude Code JSONL, not Hermes’s own generic JSONL export and not the ShareGPT trajectory format. The exporter gives the viewer familiar turn and tool-use blocks, but it drops system messages and substitutes a marker for image content. If you expected a complete archive to restore later, use a session backup format instead.

The second trap is the --no-redact switch. It is easy to confuse it with an option that “adds redaction”; the polarity is opposite. Default behavior is redacted. --no-redact opts out. Leave it off unless you have a concrete reason and have read the resulting trace.

There is a useful failure signal too. On an upload, a missing token or missing Hub package returns a setup error. A redaction failure says the upload was blocked. Do not work around that by retrying with --no-redact on a live session. Export a copy for manual review, resolve the issue, then decide whether the content belongs in a remote dataset.

Verify the trace without leaking it

I ran the local export against a stored session and parsed every output line as JSON. The captured file had 86 records: 46 user records and 40 assistant records, with a tool-use block in each assistant record. I suppressed the message text during verification. The counts above are from that run.

The useful verification checks are mechanical:

  • The command reports one exported trace, or the expected count for a filtered batch.
  • Every line parses as JSON and carries type, message, uuid, and parentUuid fields.
  • The event chain starts at one root and tool results refer to tool-use IDs.
  • Image inputs show as omitted text markers, rather than pretending the trace contains the original media.
  • Before --upload, inspect the output file and confirm the destination dataset is private unless public release is deliberate.

The important distinction is not whether the JSONL parses. It is whether you know which parts of the session left the machine. Export locally to learn the shape. Upload only after you have looked at the content.

Sources

Keep reading