The Probability Was the Bottleneck
Give an open model a yes/no question and read the next-token distribution. On the example the AnyJev authors put in their release notes, the raw readout says P(Yes) = 0.62. Replace the actual content with N/A and ask the same question: P(Yes) = 0.70. The model leans Yes regardless of what it was shown, so the second number is the model’s own prior, and dividing it out and renormalizing leaves P(Yes) = 0.41. Same weights, same prompt, same question: the judgment flips.
That is not a curiosity. It is the failure mode sitting underneath every agent loop that decides by threshold. Route this ticket if the model is more than 80% sure. Escalate if the tool call scores above 0.9. Approve if the risk classifier is confident. Each of those is a comparison against a number that, on raw logits, does not mean what the comparison assumes it means.
The number a gate needs is not the number a model gives you
The uncomfortable part is how large the distortion is, and how cheaply it can be measured. On Qwen3-8B over a 20-way banking intent task, the raw next-token readout reverses its answer on 23.0% of the 300 test items when the option list is reversed. Its expected calibration error is 0.240, which is the difference between the confidence it reports and the accuracy it actually delivers. At a 5% risk tolerance, the share of items it can be trusted to answer without a human is 7.7%.
Every production router that reads logits and compares them to a threshold is operating in that regime: a score of 0.9 that is wrong about one time in four at the margin, on a classifier that is otherwise 74.7% accurate. The accuracy was never the problem. The published AnyJev results put the automation rate at 7.7% for raw logits and 52.0% for a calibrated variant of the same model on the same items, while accuracy moves from 0.747 to 0.807. The model does six points better. The volume you can safely automate multiplies by nearly seven.
What shipped
AnyJev is a Python library from Jiamu Zhang, Tianze Yang, Liang Wu (Nokia, Sunnyvale) and Yucheng Shi (Tencent Hunyuan), released on PyPI on September 21, 2026 under Apache-2.0. It turns any open causal language model into a typed decision endpoint: give it a state and a set of typed questions (a choice, a yes/no, a score), and it returns a decision with a probability per question, read from a single prefill. Nothing is generated and nothing is parsed, so nothing can be smuggled in through prose the way a JSON-mode answer can.
The levels are the product:
raw: a restricted softmax over the label tokens, which is what every logit-reading wrapper does. The library’s own framing is blunt: it fixes nothing about bias or calibration.L0, zero labels: cyclic-shift marginalization. The question is asked over every rotation of the option list and the results are averaged, which removes the position bias, then the label prior, estimated without any labels, is divided out.L1, 100 to 500 labels per question: temperature scaling on top of L0, stored as a small artifact per model and question. It changes confidence estimates, not the ranking.L2, 100 to 300 labels per question: a closed-form head fitted on the model’s hidden state partway down, solved in seconds without gradients and without touching the weights. One prompt per state, read from an intermediate layer rather than the end of the network.
The level is carried on the result, not assumed. r.level returns raw, L0, L1 or L2, and require="L1" makes downstream code refuse to act on anything weaker. For a loop that escalates on a threshold, that is the difference between a gate and a coin flip.
The measured gap
The numbers in the release are regenerated from committed JSON, which is the part worth appreciating before arguing with them. Qwen3-8B on BANKING77, 20-way, 300 test items:
| raw logits | L0, zero labels | L1, 100 to 500 labels | |
|---|---|---|---|
| Answer flips when options are reversed | 0.230 | 0.073 | 0.077 |
| Accuracy | 0.747 | 0.803 | 0.807 |
| Calibration error (ECE) | 0.240 | 0.184 | 0.095 |
| Auto-decidable at 5% risk | 7.7% | 46.3% | 52.0% |
Forty-six percent of the traffic becomes automatable before any labels exist, because the first fix is arithmetic rather than training. The label prior estimate is read from the model’s behavior on placeholder content, and the position bias is averaged over the K rotations of the answers. L0 costs K prefills for a K-option question, which is the real price of the zero-label level and is why nobody gets it for free on a 20-way classification.
With labels, the picture shifts from the prompt to the residual stream. L2 fits a head on an intermediate hidden state; the same authors report, on a 20-question set with 300 labels each and 2,000 held-out decisions:
| model | L0, zero labels | L2 | block used | cost vs one forward pass |
|---|---|---|---|---|
| Qwen3-1.7B | 0.494 | 0.730 | 18 / 28 | 0.70x |
| Qwen3-4B | 0.564 | 0.786 | 24 / 36 | 0.69x |
| Qwen3-8B | 0.647 | 0.771 | 24 / 36 | 0.68x |
| Qwen3-30B-A3B | 0.630 | 0.799 | 40 / 48 | not reported |
| Qwen3-32B | 0.700 | 0.798 | 52 / 64 | 0.84x |
Read the block column and the cost column together. A decision read from a middle layer costs less than a plain forward pass, because the prompt stops there: the remaining blocks are busy turning the answer into tokens the caller throws away. On Qwen2.5-7B, cutting the network from 28 blocks to 18 left accuracy slightly higher and calibration better. An L2 endpoint is a pooling server plus a few kilobytes of head file, which is why the release ships heads for five Qwen3 models at about 100 KB each.
The model was never the lever
The tempting reading is that you need the dedicated decision model. The same benchmark set, published alongside, is less flattering to that idea. Jev 1.13.0 is listed at 0.727 and a fine-tuned Laya at 0.768, both as published by their authors and explicitly not rerun here. A Qwen3-1.7B with an L2 head at 64% of its depth reaches 0.730, and a Qwen3-4B reaches 0.786. The 30B and 32B rows land at 0.799 and 0.798, which is where the scaling actually stops paying.
The comparison that matters more than the headline accuracy is the one about question coverage. Laya’s zero-shot checkpoints, the ones you would deploy on a question the model was not fine-tuned for, score 0.34 to 0.36 against a 0.32 random baseline on this set, and the fine-tuned checkpoint carries an ECE of 0.215 against 0.034 for the calibrated open-model pipeline. A decision model is a bet that you know your questions in advance and will keep paying to train them. The alternative is a general model plus an interface that extracts a probability you can threshold, per question, for 100 to 300 labels, with the same model serving everything else.
What an operator should change
Read the level. If the code that acts on a decision cannot say which level produced it, it cannot know whether the probability was debiased, calibrated, or neither. Requiring a minimum level turns an implicit assumption into a caught exception.
Threshold coverage, not accuracy. The metric that decides how much traffic leaves the queue is the share of items above the threshold at an acceptable error rate, not the top-line accuracy. On the raw readout of Qwen3-8B that share was 7.7%; the same model at L1 was 52.0%. If you have never measured it, you are already automating somewhere, you just do not know where.
Start at L0 with no labels, and let labels arrive from the loop. The library’s observe path collects labels as they come in and solves the head by itself at 30, re-solving at 60 and 120, so day zero runs with no labeled data and L2 arrives from production traffic instead of a labeling project.
Design for drift, because the head already does. A reworded question drops the Qwen3-8B head from 0.77 to between 0.65 and 0.70, and 30 unlabelled requests bring it back to 0.74 to 0.75, against 0.77 for a full relabelled refit. Only the head’s feature mean and scale move. New labels are needed for a new question, not for a rephrased one.
Respect the constraints or you will call it broken. The library needs logits or hidden states, so it is for open or self-hosted models; an API that returns only text cannot feed it. A head is bound to one model and one question and does not transfer. The letter readout caps at 26 options. L0’s batch prior costs accuracy when one label dominates the set. Quantizing with fp8 is available and explicitly not recommended, because it buys single-question latency at the cost of accuracy. An engine that cannot expose hidden states cannot serve L2 yet, though vLLM can through an embed server’s pooler.
Where the honesty sits
The limitations section of the release is unusually candid and should be quoted along with the wins. On the typed-decisions set, “accuracy” means agreement with a teacher model’s mean of three samples, and a fresh sample of that teacher agrees with itself only 0.735 of the time, so the ceiling on that benchmark is fuzzier than the decimal places suggest. Calibration cannot rescue a model that cannot answer: in the maze and Minesweeper harnesses the release reports no readout, including a general model’s, beating the trivial majority-label baseline on edge decisions, with edge accuracy between 0.490 and 0.555 against majorities of 0.598 to 0.618. The improvement there was in how many attempts the controller needed, not in reading the map. And Laya’s fine-tuned head still wins narrowly on Brier score, 0.118 against 0.120, which is the honest way to state a comparison the release does not rerun.
That is the shape of a mature artifact: it publishes the numbers that argue against the headline, in committed JSON, with a script to regenerate them.
The closing thesis
Agent routing gets framed as a model selection problem: which model is smart enough to decide. The measurements here say the deciding number is the artifact, and it was broken long before anyone chose a model. A threshold over an uncalibrated confidence is a rule with no measurable coverage, and the fix is arithmetic, averaging, and a few hundred labels rather than a larger network. Calibrate the number first. The model can stay.
Sources
- Avi Chawla, AnyJev announcement post, September 25, 2026 (verbatim text and metrics read through the free
mcp__xactions_hosted__x_postpath): https://x.com/_avichawla/status/2103393403207901217 - AnyJev repository, Nokia Applied Research (Apache-2.0, released September 21, 2026): https://github.com/nokia-applied-research/AnyJev
- AnyJev on PyPI, version 0.0.2, uploaded September 21, 2026: https://pypi.org/project/anyjev/
- Benchmark table regenerated from committed JSON: https://github.com/nokia-applied-research/AnyJev/blob/main/docs/results_bench.md
- L2 and typed-decisions results: https://github.com/nokia-applied-research/AnyJev/blob/main/docs/results_exit.md
- Level contract and method: https://github.com/nokia-applied-research/AnyJev/blob/main/docs/levels.md and https://github.com/nokia-applied-research/AnyJev/blob/main/docs/method_v3.md
- Maze harness results, including the majority-label baseline: https://github.com/nokia-applied-research/AnyJev/blob/main/docs/results_maze.md
Method note for this post: the candidate came from the free, session-authenticated X search wrapper (twsearch "open weights model release"), and every figure above was read from the project’s README, PyPI page, and the committed results documents linked here in this session. The X post’s description of the three levels matches the repository. Rows published by the Jev and Laya authors are quoted as published and were not rerun here, as the AnyJev release states about its own tables.