Hermes Agent Deep Cuts: Your Screenshot Is Routed, Not Attached
$ cd /home/dazeb/.hermes/hermes-agent && venv/bin/python3 image-routing-probe.py
HERMES_HOME -> /home/dazeb/.hermes/profiles/blogposter
main provider/model -> custom:api.buzzgw.com / deepseek-v4-flash
agent.image_input_mode -> auto
auxiliary.vision -> {'provider': 'auto', 'model': ''}
decide_image_input_mode -> text
_accepts_tool_result_images -> False
_should_use_native_vision_fast_path -> False
Three lines of that output are the whole story. The model answering this box cannot be handed an image, the tool that would attach one to a tool result is switched off for it, and the routing layer agrees on both counts before any turn is built. Every screenshot and every pasted path becomes a text description written by a different model than the one you are talking to.
That is the safe answer, and it is also a decision you have never seen made. The same function that picks text here decides whether an image path typed in a sentence becomes an attachment, whether the pixels go out at full size or halved four times, and whether a format the provider hates gets transcoded or dropped. None of it appears in the transcript.
The path from your message to the model
Two lanes share one decision.
The native lane is short. extract_image_refs() in agent/image_routing.py scans the user’s free-form text for image references. Local paths must be anchored at ~/ or /, carry one of the allowed image extensions, and exist as a regular file. URLs must be http(s). Then build_native_content_parts() turns each local path into a base64 data: URL, passes remote URLs through untouched, and prepends a text part that carries a string handle per image so a later tool call can find the same file.
Here is the scanner on a message shaped like the ones people actually send, including the parts a careful reader expects to be ignored:
$ python3 image-routing-probe.py # section 1
local paths -> [
".../img-routing-probe/shots/error-42.png"
]
urls -> [
"https://example.com/assets/diagram.webp?rev=3",
"https://example.com/b.png"
]
Four things did not make the list. The path that only exists inside fenced and inline code spans was skipped, because extract_image_refs() strips code spans before matching so a pasted command does not become a live attachment. The path to a file that no longer exists was skipped, because local candidates are validated with os.path.isfile. The file:// URL was skipped, because the URL pattern is strictly http(s). And /spec.pdf was skipped by extension: documents and archives are deliberately excluded because a PDF must never become a vision part.
The text lane is where this box lives. The gateway runs the configured vision model over each attached image, prepends its description to the user’s message, and hands the model the local cache path so it can look again if it wants to. The wrapper is literal, and it is worth reading once because it is what your model actually receives:
[The user sent an image~ Here's what I can see:
<description>]
[If you need a closer look, use vision_analyze with image_url: /path/to/cache.jpg ~]
When the description call fails, the note degrades instead of vanishing, and the failure text is chosen so the model cannot mistake it for a description:
[The user sent an image but I couldn't quite see it this time (>_<)
You can try looking at it yourself with vision_analyze using image_url: /path/to/cache.jpg]
Both lanes exist because the model may or may not be able to see. That is the only question the router asks.
The decision, in precedence order
decide_image_input_mode(provider, model, cfg) returns one of two strings, native or text, and the order of its checks is where the surprises live. Every line below is real output from the probe, run with explicit config dicts so no config file could interfere:
config override says the model can see -> native
override is the string 'false' (quoted YAML) -> text
per-provider override, main override absent -> native
explicit auxiliary.vision backend, main model CAN see -> text
agent.image_input_mode = native, capability unknown -> native
agent.image_input_mode = text, capability true -> text
no provider/model, no override (unknown capability) -> text
Read the fourth line twice. A main model that demonstrably supports vision still gets routed to text if auxiliary.vision names a backend, because the code treats an explicitly configured vision backend as the de-facto image route even in auto mode. The first three lines are the config overrides winning. The fifth and sixth show agent.image_input_mode as the absolute override in both directions. The last line is the fail-safe: unknown capability resolves to text, never to a guess.
Underneath the override sits the capability ladder, in priority order: the managed local runtime’s own capability table, then the models.dev catalog, then an Ollama /api/show probe, then a registered provider profile’s declaration. The catalog lookup is allowed to hit the network because a cold cache returning “unknown” would push a screenshot onto a cloud auxiliary model. Before any of that, _supports_vision_override() checks model.supports_vision, then providers.<name>.models.<model>, then legacy custom_providers entries. The coercion is strict, and the comment says why:
def _coerce_capability_bool(raw):
if isinstance(raw, bool): return raw
if isinstance(raw, int): return bool(raw) if raw in (0, 1) else None
return _BOOL_TOKENS.get(raw.strip().lower()) if isinstance(raw, str) else None
bool("false") is True in Python, so a quoted supports_vision: "false" in YAML would flip a text-only model onto the native lane and break every image turn. The strict parse means only real booleans, 0/1, and the literal tokens true/false/yes/no/on/off/1/0 count. Anything else falls through the ladder.
The three keys that change the answer
Ask the box what it currently thinks, then change it deliberately:
$ hermes config get agent.image_input_mode
auto
$ hermes config get auxiliary.vision
provider: auto
model: ''
$ hermes config set agent.image_input_mode native
$ hermes config get agent.image_input_mode
native
auto is the default and the right setting for most operators, because it only consults the capability ladder. native is the absolute override and the escape hatch for local or custom providers the catalog has never heard of. text is the deliberate downgrade, useful when you would rather pay for a small vision model’s description than send a multi-megabyte image to an expensive main model. Each key lands in the active profile’s config.yaml, so a second profile on the same box can sit on a different lane without touching the first one.
The second key is the one that bites, and the Vision & Image Paste page does document the rule: a provider other than auto, or any model or base_url, selects the description path even for a vision-capable main model. What the doc cannot tell you is whether your box has crossed that line, because nothing announces the switch at runtime. The precise condition is narrower than “you set something there”. auto as the provider with an empty model and no base_url does not count as explicit, so the probe output showing {'provider': 'auto', 'model': ''} is the not-explicit case. The text answer on this box therefore comes from the capability ladder finding nothing it can attest, not from an override. If your agent suddenly starts paraphrasing screenshots, that key is the first place to look and the last place anyone looks.
The same distinction explains a result that looks wrong on this host. The docs list DeepSeek Flash among the vision-capable models, and this box runs deepseek-v4-flash on a custom endpoint. It still routes to text, because capability is resolved per provider and a custom endpoint is not in the catalog, so nothing attests the model. The one-line fix is the absolute override:
$ hermes config set agent.image_input_mode native
What actually arrives on the wire
The native lane is where the byte-level care shows up. From the probe:
_guess_mime() on a JPEG named .webp -> image/jpeg
skipped -> [ ".../gone.png", ".../vector.svg" ]
text part -> "what is wrong with this render?
[Image attached at: .../shot.jpg]
[Image attached: https://example.com/far.png]"
image part -> data:image/jpeg;base64,/9j/4AAQSkZJRgABAQAAAQABA... (199,015 chars)
image part -> https://example.com/far.png... (27 chars)
A 121 KB JPEG becomes a 199,015-character data URL, attached at native size. The MIME comes from magic bytes before it comes from the name. That order matters because platforms lie about content types: Discord serves proxied stickers as PNG with image/webp in the header, and Anthropic rejects a mismatched media type with an HTTP 400 that kills the turn. Formats outside png/jpeg/gif/webp (BMP, TIFF, HEIC, AVIF, ICO) are transcoded to PNG through Pillow, with pillow-heif and pillow-avif-plugin registered on demand when present. A managed local runtime gets a narrower accepted set on purpose, because it decodes fewer formats and fails in the worst possible way: a WebP part it cannot read produces no error at all, and the model confabulates a description.
The send layer is more forgiving than the attachment layer. A Responses-style backend rejects the whole request over one bad inline part, and because that part stays in history every later turn fails the same way. So an inline SVG is rasterized to PNG when a rasterizer such as cairosvg or rsvg-convert is installed. Without one, the SVG becomes a short [image omitted: image/svg+xml is not a supported image format] placeholder while the valid images in the same message still go out.
Sizing is reactive. The image goes out at full size, and the agent waits to be told it was too big. On an image_too_large failure, the recovery pass runs one shrink attempt, halving dimensions up to four times with a 64px floor and walking a JPEG quality ladder, then retries. If the provider rejects the bytes anyway, the image parts are stripped from the tool messages rather than failing the turn. There is no proactive table of provider limits, and the comment in the source gives the reason: ceilings are partial, they change, and a stale table would quietly degrade quality for everyone whose provider got more generous last month.
vision_analyze has its own fork, and it is the third line of the opening probe. When routing says native and the provider carries images inside tool results, the tool returns a multimodal envelope: one short text part telling the model it can see the image now, and the base64 image alongside it. The docs list the stacks that do carry them, Anthropic, OpenAI, Azure-OpenAI and Gemini 3.x. Everywhere else the tool falls back to the older lane, where an auxiliary model describes the image and only its text comes back. The envelope also carries a text_summary field for providers that take the text part but not the image, which is the quiet half of the mechanism: a fast path can degrade to a summary line without raising anything. Two separate gates, one for routing and one for the transport, which is why _should_use_native_vision_fast_path asks both questions before it answers.
A native embed is not free after the fact either. It is baked into the tool result, so it rides every later API call in the session, which is what vision.embed_target_bytes (256 KB by default, clamped between 64 KiB and 4 MiB) and vision.max_calls_per_image exist to bound. Hit the call cap and the tool stops embedding and answers with vision_analyze refused: this image has already been loaded into context N time(s), which reads like a bug the first time you meet it.
The six ways this fails quietly
None of these throw. All of them are visible only if you look at the log or at the config.
A path inside backticks or a fenced block never attaches. That is the documented intent, and it is also the most common way to describe a screenshot without actually sending it. If your prompt quotes a command that happens to contain an image path, nothing happens.
A path that does not exist is dropped without a message. Local candidates are checked with isfile and unreadable paths collect in skipped. Remote URLs are never validated at all, at this layer or the next one, because the provider fetches them.
file:// URLs are invisible to the scanner. The pattern is strictly http(s), so a local file URI pasted as a link attaches nothing.
On a text-lane box with no working vision backend, vision_analyze fails with unknown variant image_url. This site’s publishing workflow records that failure on this profile. The routing code shows why that is the shape it takes: the text lane sends the image to an auxiliary model, this box has not configured one (the probe shows provider: auto with an empty model), and a provider that does not accept the image part answers with a 400. Treat that error as the skip signal it is, not as a retry.
The _accepts_tool_result_images gate is separate from capability. A provider profile can declare supports_vision=True and supports_vision_tool_messages=False, meaning it takes images in user messages but rejects list-type tool content with a 400. The shipped xiaomi/MiMo profile sets exactly that veto, because its endpoint answers text is not set. When the veto is on, the native fast path for vision_analyze stays off even though the model can see, which is why the gate exists at all.
Managed and local runtimes can disagree with the cloud catalogue in both directions. A local GGUF the catalogue has never heard of reads as text-only and detours every screenshot to a cloud model, unless the managed runtime reports its own modalities. That probe is the reason the ladder starts with the server that is actually receiving the bytes.
How to verify any of this on your own box
The probe behind every number above is read-only and prints no secrets:
# /home/dazeb/.hermes/profiles/blogposter/scripts/image-routing-probe.py
from agent.image_routing import (
_guess_mime, build_native_content_parts, decide_image_input_mode, extract_image_refs,
)
msg = open("your-message.txt").read()
print(extract_image_refs(msg)) # what will attach, and what will not
print(decide_image_input_mode(provider, model, cfg)) # native or text for this turn
print(build_native_content_parts(msg, paths, urls)) # the parts, and the skipped list
Run it from the Hermes checkout with the same interpreter the tooling uses, cd /home/dazeb/.hermes/hermes-agent && venv/bin/python3 your-probe.py, so the imports resolve against the installed source rather than your system Python. Then check the live values with hermes config get agent.image_input_mode, hermes config get auxiliary.vision, and hermes config get model. If the mode says auto and the auxiliary vision model is empty, you are on the text lane until the capability ladder says otherwise, and the ladder is one config line away from being bypassed.
The reason to care is not that vision is hard. It is that this pipeline is silent by construction. Every branch fails toward the same place, a text description or a dropped attachment, and the agent will happily reason about an image it never received. If you want to know which model saw your screenshot, measure the route, not the answer.
Sources
- Hermes Agent documentation, Vision & Image Paste (the routing table for vision-capable and text-only models, the
agent.image_input_modevalues, the explicitauxiliary.visionrule, the tool-result fast path on Anthropic/OpenAI/Azure-OpenAI/Gemini 3.x,vision.embed_target_bytes,vision.max_calls_per_image, the send-layer SVG placeholder): https://hermes-agent.nousresearch.com/docs/user-guide/features/vision - Hermes Agent documentation, Tools Reference (
vision_analyze, thevisiontoolset,video_analyze): https://hermes-agent.nousresearch.com/docs/reference/tools-reference - Hermes Agent documentation, Toolsets Reference (
visiontoolset membership, capability-gated tools): https://hermes-agent.nousresearch.com/docs/reference/toolsets-reference - Hermes Agent documentation, Tools Runtime (
check_fnavailability gating, async handler bridging): https://hermes-agent.nousresearch.com/docs/developer-guide/tools-runtime - Source in
NousResearch/hermes-agent@a53b42ddea76b250a3f5fa3f6dffa057f29e6374:agent/image_routing.py(extract_image_refs,build_native_content_parts,decide_image_input_mode,_explicit_aux_vision_override,_VISION_PROBES,_coerce_capability_bool,_MAGIC,_file_to_data_url),tools/vision_tools.py(_should_use_native_vision_fast_path,_accepts_tool_result_images,_build_native_vision_tool_result,_resize_image_for_vision,_EMBED_MAX_DIMENSION),gateway/run_inbound.py(_decide_image_input_mode,_enrich_message_with_vision),agent/turn_recovery.py(image_too_largeshrink-then-strip),agent/vision_message_prep.py,plugins/model-providers/xiaomi/__init__.py(supports_vision_tool_messages=False): https://github.com/NousResearch/hermes-agent/tree/a53b42ddea76b250a3f5fa3f6dffa057f29e6374 - Live output on the authoring box, 2026-09-25, Hermes Agent v0.21.4 (upstream
a53b42dd):agent/image_routing.pyimported and executed directly for the scanner, decision matrix, MIME sniff, and native-part assembly (/home/dazeb/.hermes/profiles/blogposter/scripts/image-routing-probe.py), plushermes config get agent.image_input_mode,hermes config get auxiliary.vision, andhermes config get model. The test files in the scanner probe were created for the run; the config values and routing verdicts were not.