Hermes Agent Deep Cuts: The Context-Saving Power of execute_code
Most agents fail because they suffer from “Context Bloat”—the LLM gets overwhelmed by every intermediate step of a task, trying to remember every raw tool output before it can reach a conclusion.
The execute_code tool fixes this by moving the “mechanical” logic out of the LLM’s context window and into a child process. Instead of the agent seeing every single search result, every intermediate list, and every raw string, it only sees the final, polished output of a Python script. It’s the difference between a chef describing every single chop of a vegetable or just presenting the finished soup.
The Mechanism: RPC over Unix Sockets
Unlike the terminal() tool, which simply pipes a string of text into a shell, execute_code is a programmatic bridge.
When you invoke execute_code, Hermes does the following:
- Stub Generation: It creates a
hermes_tools.pymodule, providing a Pythonic interface for every built-in tool (e.g.,web_search,read_file). - Socket Connection: It opens a Unix domain socket (or a TCP loopback on Windows) to communicate with the child process.
- Script Execution: The Python script runs in a separate process. When the script calls a tool, the call is serialized to JSON and sent over the socket.
- The “Print” Payoff: Crucially, only the
print()statements in your script are returned to the LLM; intermediate tool results never enter the context window.
Advanced Usage: Logic vs. Reasoning
The real power of execute_code is revealed when you need to perform 3+ tool calls with logic between them.
If you want to find the top 3 news stories about “Quantum Computing” and then extract the content of just the first one, a standard agent might take 3-4 turns:
web_search("Quantum Computing")$\rightarrow$ (LLM sees 10 results)web_extract(result[0])$\rightarrow$ (LLM sees a huge block of text)summarize(extracted_text)$\rightarrow$ (LLM produces final result)
With execute_code, this happens in one turn:
from hermes_tools import web_search, web_extract
results = web_search("Quantum Computing", limit=10)
# The LLM doesn't see the 'results' list yet, only the final output
top_content = web_extract([results["data"]["web"][0]["url"]])
print(f"Summary of the top story: {top_content['results'][0]['content'][:200]}")
This “Context-Saving” approach is the most efficient way to handle:
- Loops: Iterating over a list of URLs to find the “best” one.
- Filtering: Reducing a massive search result down to 3 key data points before the LLM even looks at it.
- Data Transformation: Converting a raw list of tool outputs into a JSON object or a formatted string.
The Gotcha: project vs. strict
A common point of failure is the execution mode defined in your ~/.hermes/config.yaml.
projectmode (Default): The script runs in the same directory as your project. It uses your activeVIRTUAL_ENVorCONDA_PREFIX. Use this when your script needs to “know” your project (e.g.,from my_module import tool).strictmode: The script runs in a clean, isolated temporary directory. It uses Hermes’s own Python interpreter. Use this when you want to ensure that a script doesn’t accidentally “pollute” the context by reading a file in your project tree that you didn’t explicitly mention.
The Fix: If your script is throwing ModuleNotFoundError or finding files in your project you didn’t expect, check your mode. Most everyday tasks are better in project mode, but for reusable utility scripts, strict provides the cleanest boundaries.
How to Verify
Run this exact block to see the “Context-Saving” power in action. Instead of the agent listing 5 search results and then choosing one, it filters for the longest title and extracts the content in a single leap.
from hermes_tools import web_search, web_extract
# Search for 5 results
results = web_search("the impact of large language models on software engineering", limit=5)
# Find the result with the longest title programmatically
longest_title_idx = 0
max_len = 0
for i, r in enumerate(results["data"]["web"]):
if len(r["title"]) > max_len:
max_len = len(r["title"])
longest_title_idx = i
# Extract content from only that one
content = web_extract([results["data"]["web"][longest_title_idx]["url"]])
print(f"Longest Title: {results['data']['web'][longest_title_idx]['title']}")
print(f"Content Preview: {content['results'][0]['content'][:300]}")
Sources: