The Router Does Not Need to Reason

The Router Does Not Need to Reason

A 340M parameter encoder now beats a 4B general-purpose classifier, and its own 1B sibling, at the job most agent stacks quietly spend their money on: deciding what to do next.

Fastino released GLiNER2.5-Decide on September 23, 2026, under Apache 2.0. It is a 340M encoder trained to answer typed questions against a label set you pass at call time, in one forward pass. Its own model card draws the boundary for you: “This release is not a general-purpose model. It does not reason, explain, or answer open questions. It is a specialist for operational decisions.”

That sentence is the whole story. The decision layer of an agent stack is being carved out as its own model class, with its own leaderboard, its own metric and its own failure mode, and the usual instinct about which model to reach for is wrong in a specific and useful way.

The work that never needed a generator

Denny Sentinel has argued since July that routing is the structural defect in most LLM products: tickets, email triage, sentiment, spam and priority calls run through the same frontier endpoint as multi-step reasoning, at 10 to 30 times the cost, producing one label where a 4B model would have done. The missing piece was never the idea. It was that “route with a small instruct model” still means generating tokens to emit one word, still means prompt templates that drift, still means sampling variance, and still means you cannot easily say how confident the router was.

GLiNER2.5-Decide is a direct attack on each of those. The labels are arguments, not prompt text. The output is a structured decision with a probability distribution and a confidence score, per the release announcement. There is no prompt template and no generated token, so the same input and label set produce the same answer every time. The library is CPU-first with no GPU required, and the base install of the GLiNER2 package does not even pull PyTorch:

pip install gliner2            # schema, API client, training utilities
pip install gliner2[local]     # local inference and LoRA support
from gliner2 import AutoExtractor

model = AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")

One call can score several heads at once, so the same forward pass answers intent, priority and policy at the same time. Single-label heads return a string. Multi-label heads return every label above a threshold. If you have written the same classification prompt four times with different label lists, that is the shape of a schema, and it was already the shape of the problem.

The table is the argument

The benchmark is the vendor’s own Fast Decisions suite: 17 domains, 300 held-out examples each, with the same text and candidate labels given to every model, scored as exact match.

ModelParametersAvg exact match
GLiNER2.5-Decide340M60.2%
GLiNER2.5-Decide-1B1B59.6%
JevK5not stated57.6%
GLiNER2.5-multi-Decide287M56.7%
SemIf (Qwen3.5-4B)4B56.4%
GLiFormer large-v1not stated49.0%
Laya Routernot stated46.6%

Read the top two rows together. The 340M model beats the 1B model of the same family. Scaling the artefact up made the decisions slightly worse, which is not how capability curves are supposed to look, and it makes sense only if you accept what the task actually is. A router with a closed answer space is not reasoning. It is mapping a passage onto one of several labels you already wrote down, and the quality lever is representation and calibration, not depth.

The same reading explains the SemIf row. A 4B general classifier, fine-tuned for classification, lands 3.8 points below the 340M specialist. Ask a big model to hold the world in its head and it will do it, slowly and expensively. Give a small model the label set and the passage and nothing else, and the question becomes tractable.

There is one small inconsistency worth noting, because it is the kind of thing that shows up in vendor numbers: the announcement rounds the 340M score to 60.1% and the card states 60.2%, and the same gap appears on JevK5 (57.5% in the post, 57.6% on the card). The card is the artefact that gets corrected later, so it is the one to quote.

What 60% means, and why the confidence score is the product

Four in ten decisions wrong is not a deployment. It is a threshold.

This is the part that separates a demo from a control loop, and it is where the release is more interesting than its benchmark table. A classifier that returns only a label forces a binary choice: trust it, or do not run it. A classifier that returns a probability distribution lets you build a gate. Route the confident cases to the cheap path, escalate the uncertain ones to the expensive model, and log the escalation rate. That rate is the number that determines whether routing saves you anything at all, and it is invisible if your router is a prompt that emits one token.

The operating consequence is worth stating plainly, because it inverts the usual wiring. If your router is an LLM, you cannot easily put a mathematically meaningful threshold on it, so you end up putting a frontier model behind every decision. If your router is an encoder that reports its own uncertainty, you can put a frontier model behind only the decisions that need one. The verifier is still the bottleneck, as this site argued in July. It is just that the verifier is now a threshold you can measure and tune instead of a vibe you argue about.

The category is forming around this

The narrow-model-for-decisions idea is not Fastino’s alone, and the surrounding evidence is thinner than the model card, so treat it as claims rather than findings. A user claiming to build the same shape posted on September 18: Kev-0.5B is described by its author as “a tiny open source Jev-like decision model with a TypeSafe-compatible API based on Qwen2.5-0.5B that you can train and run on a MacBook Pro.” The same week, another account claimed to have classified 1,697 conference events with Gemma 4 26B-A4B and no fine-tuning. Neither is verified here. What is verifiable is that a named decision model, JevK5, appears as a baseline in a competitor’s benchmark table, which is what a category with a scoreboard looks like.

The pattern is also already inside agent harnesses. Hermes Agent ships a Jev plugin whose own tool descriptions are blunt about the design: it “picks, ranks and gates; it never writes.” The same tool cut appears in three places in that stack, choosing which search results to read, filtering retrieved passages, and selecting the next browser or desktop action from a table of prevalidated steps. None of them need a writer model. All of them need a fast, repeatable, auditable decision with a confidence attached.

The adoption signal on the Hub is the part I would watch. Within roughly 48 hours of release, the ONNX port appeared, a third-party LoRA appeared, and a cultural heritage lab posted a domain-specific fine-tune for linked-artwork reconciliation. At capture on September 25, the base model had 1,048 downloads and 132 likes on the Hub. Domain fine-tunes arriving in a day and a half is the tell that people have label sets they were already maintaining by hand.

The gotcha that will make it look broken

Two things will waste your afternoon if nobody tells you.

First, the API split. The GLiNER2 repository ships two extraction architectures behind one loader, and the older constructor refuses the newer checkpoints:

GLiNER2.from_pretrained(...) remains span-only and will not load GLiNER2.5 boundary checkpoints.

So GLiNER2.from_pretrained("fastino/GLiNER2.5-Decide") is not a slightly wrong call, it is the wrong class. Use AutoExtractor.from_pretrained(...) and let the saved architecture field dispatch.

Second, the variants. The 287M multilingual sibling GLiNER2.5-multi-Decide scores lower on the English suite (56.7%) for the obvious reason, and the 1B model is not the upgrade. If you want to reason about which one to pull, remember the table above rather than the parameter counts.

What an operator should change

  1. Sort your calls into decisions and reasoning, then stop paying for the first kind. If the answer space is closed and known before the request arrives, you have a classification problem wearing a chat interface.
  2. Make confidence a first-class field. A label with no score cannot be gated. Route on the score, escalate the tail, and instrument the escalation rate as your primary routing metric.
  3. Fine-tune before you prompt-engineer. The label lists you keep rewriting are a training set. LoRA on a few hundred of your own examples is cheap on this class of model, which is exactly why third parties were shipping fine-tunes the next day.
  4. Keep sensitive classification local. Customer messages, tickets and clinical or financial text are the inputs these models are aimed at. One forward pass on a CPU with no external call and no prompt template removes an entire prompt-injection surface, because there is no instruction channel to hijack.
  5. Do not read a vendor suite as your accuracy. Seventeen internally generated English domains measured by exact match is a directional claim about a model class, not a prediction about your traffic. The published “unseen” suite is unseen to the models, and its author is the vendor.

The uncomfortable truth about the routing layer is that it was never a reasoning problem, and the industry spent two years solving it with a generator that reasons. A 340M encoder beating a 1B sibling is the moment that becomes hard to argue with. The router does not need to think. It needs to be right often, cheap always, and honest about the times it is not.

Sources:

Keep reading