The Lead Was Inside the Tie Band
On September 25 at 03:25 UTC, Nace AI announced Drex, quoting the post exactly: “#1 on the Decision Index. (official scores tbd) winning 23 out of 40 benchmarks.”
Four days earlier, the Decision Index did not exist. It arrived on September 22 as a 132,422-request benchmark suite for typed decision engines, models that take a state and a set of typed questions and return a probability for every option you supply, with none of the text generation a chat model does. It has since been re-issued twice. The board is now on Decision Index 0.2.1, dated 2026-09-27 in its own metadata, with a 38-benchmark panel, 64 entrants and a different scoring formula.
The interesting part is what a rank means when the ruler is moving this fast. On the edition Nace scored against, the reference model sat at 51.67 and Drex 1.0 sat at 51.73. That is a 0.06-point lead. The reproduction kit for that same edition states, in its scoring section, that “the board treats scores within 0.25 index points of the next one as tied”, and its 0.2 specification file carries "tie": 0.25 next to the panel definition. By the board’s own rule, the number in the launch post was not a win. It was a tie.
This is not a story about a dishonest vendor. Nothing here is hidden, the vendor’s own three documents disagree with each other in plain sight, and the board publishes more methodology than most benchmarks ever do. It is a story about what happens when a measurement becomes infrastructure: the number travels faster than the tuple it depends on.
What actually shipped
The Decision Index lives in a Hugging Face Space maintained by Apolinário Passos (@multimodalart), with a reproduction kit on GitHub under MIT. The kit is the part that matters. It does not ship the test data, because several upstream sources forbid redistribution. It ships the recipe: rebuild the frozen suite from pinned sources, hash-check it on import, run any engine through it with checkpoint and resume, score it with the board’s own scorers, and compute the index. One command runs the whole suite as a single Hugging Face Job on one RTX PRO 6000.
Three editions in six days:
| Edition | First published | Panel | Entrants | jev index (balanced skill) |
|---|---|---|---|---|
| 0.1 | 2026-09-22 | 19 static benchmarks, five equal-weight areas | 31 plus the reference | 46.26 (raw 59.51) |
| 0.2 | 2026-09-24 | 40 benchmarks, chance-corrected | 64 plus the reference | 51.67 |
| 0.2.1 | 2026-09-27 (board metadata) | 38 benchmarks, weighted areas | 64 plus the reference | 57.89 |
The 0.1 and 0.2 dates are the announcement dates; the changelog entries for both later editions are stamped 2026-09-27 in the file the Space serves today, which is itself part of the argument.
Those numbers are read from the files the Space itself publishes, fetched on September 27: data/index-v0.1.json, data/index-v2.json and data/index.json, plus the matching methodology files. The changelog in methodology.json records what moved. Edition 0.1 showed only the balanced raw index (59.51 for the reference) and computed the chance-normalised variant without publishing it. Edition 0.2 introduced chance correction so that random guessing scores zero, added 30 models and grew the panel from 19 benchmarks to 40. Edition 0.2.1 rewrote the area weights, gave thirteen “gold” benchmarks a 1.2 weight inside their area, dropped SGD and RouterBench from the index (RouterBench because, in the board’s words, “its prompt gives away the best route”), changed RAGTruth’s chance baseline from a fair coin to always answering “hallucinated”, and de-duplicated 72 rows in the home-appliance benchmark.
Read that list again as an operator. The same model, on the same frozen suite, moves 11.6 points between editions because the panel changed and the aggregation changed. Any 52.82 you see quoted is only interpretable if you also know which of those three panels produced it.
The board’s own hard rules say the quiet part: “Point estimates only; paired cluster-bootstrap uncertainty is not yet calculated.” The index has no error bars, and the published tie band is the stand-in for them.
The tuple behind a rank
Nace AI’s own material disagrees with itself, which is the useful part.
- The launch post, September 25: “#1 on the Decision Index. (official scores tbd)”, an architecture described as a small diffusion model with RLAF, $0.04 per million input tokens, sub-second latency, and “Open weights and the full tech report are coming very soon.”
- The launch blog post: Drex scores 51.73 on edition 0.2, “ahead of Jev 1.13.0 at 51.67 and of the strongest community entry, AutoJev-27B, at 50.94”, under 6B parameters, winning “twenty-two of the thirty-nine benchmarks scored for it so far”.
- The product page: 52.82, labelled “Decision Index 0.2, rank 1”, “+1.15 over Jev 1.13.0”, “21 of 40 tests won”, under 10B parameters, “leads all 50 entries”, and a footnote saying Drex 1.1 “was scored on 2026-09-25 with the official scorer, on all 40 benchmarks”. The same page also carries a comparison block that prints “+0.00 points ahead” beside the 52.82 in the header.
So the win count for “the number one decision model” went 23, then 22, then 21 in three days, and the score went 51.73 (Drex 1.0), then 52.82 (Drex 1.1), against a reference that did not move between those two readings. The 51.73 number is the one inside the tie band. The 52.82 number is outside it, and it belongs to a different checkpoint that the same page says was scored separately on the same day.
There is a second thing to check before repeating any of it. The board publishes three data files and none of them contains the entry. I searched all 31 entries in 0.1, all 64 in 0.2 and all 64 in 0.2.1 for the strings “Drex” and “Nace”, case-insensitively, in the JSON published by the Space on September 27: zero hits. The submission path is explicit in the kit: run the full suite, upload the results directory to a Hugging Face dataset, and open a pull request, because “the results file is re-scored on review”. A run that has been scored privately and announced publicly but is not on the board, or is still in review, would look exactly like this. That is the honest reading: the rank is a vendor-run result, not a board-published one, and the difference between those two things is the whole point of a board.
Worth noting, because it is not a criticism: the product page states the model “is trained on the official training splits of the index benchmarks alongside verifiable procedural data, and evaluated only on the held-out splits with the official scorer. That is how the leaderboard is designed to be run.” That is a fair description of the protocol, and it also explains the shape of the results the page reports, which are strongest on contract reasoning, chord recognition and sarcasm, and weakest on graduate science and multi-step puzzles. The index measures how well a decision model has been fitted to the kinds of decisions in the suite. Your decisions are probably not in the suite. That is an argument for tuning on your own labelled outcomes, which the same page offers to do.
Rank and calibration are two leaderboards
Here is the number that actually decides whether you can put a decision model inside an automated loop, and it is on the same board, in a different table.
Every entrant carries a calibration block: accuracy, mean confidence, expected calibration error, Brier score, the share of confident-wrong answers, and the reliability curve bins underneath. It covers 32 benchmarks that have a per-field right or wrong answer, on a one-in-six sample, and the note states that “rows an entrant trained on count as wrong here too”, which closes the obvious loophole.
On the current board, the reference model that leads the index at 57.89 has an ECE of 0.074: it is right 73.9% of the time while reporting 81.2% confidence, so it is systematically overconfident by about seven points. AutoJev-27B sits fourth on the index at 56.4 with an ECE of 0.0178. Jebadiah 27B is sixth at 54.67 with an ECE of 0.014. Decider 35B-A3B NVFP4 is tenth at 47.11 with an ECE of 0.0226.
If you are thresholding a router on “route the uncertain cases to a human”, the index tells you which model gets the most answers right and the calibration table tells you whether 0.8 means 0.8. They are not the same list, and the model at the top of one is not at the top of the other. This is the same point that made the Decision Index worth watching on day one: a decision layer’s value is a confidence you can act on, not a label count.
The board added this table hours after launch, on September 22, and announced that “Decider 35B-A3B by @notmapika, based on Qwen3.5-35B-A3B-Base, takes the lead as the most calibrated model on our benchmarks”. Fine calibration and a middling index rank, on the same entrant, in the same run.
What the board does that most leaderboards do not
The interesting engineering is in the guard rails, and they are worth copying:
- Submissions are re-scored, not trusted. The changelog for 0.2.1 records that “Surogate Rune 26B-A4B v3 replaces v1 on the board: the author’s submitted run, adopted after our 3,000-row reproduction matched it exactly.” When an author’s run could not be reproduced, it was not adopted.
- Contamination is hunted, not assumed away. The kit’s contamination file records a case where a model card declared no training data while a sibling dataset from the same publisher contained 3,077 of the board’s own 3,080 BANKING77 test items with labels. Those items are counted as wrong rather than quietly averaged in.
- Capped runs retire. “Lifting a cap the author set outside the model, or replacing a checkpoint, retires the earlier run from the ranking.” Eight entrants have already been retired by that policy.
- Benchmarks leave with a reason printed. RouterBench and SGD left the index and stayed on the board, with the reason given. MMLU left because it is saturated.
What to change
-
Never quote a decision-model rank without its tuple. Edition, panel, formula, checkpoint and scoring rule. “Number one on the Decision Index” is not a fact about a model. On September 25 it was a fact about a 40-benchmark panel, a chance-corrected skill formula and one checkpoint.
-
Check that the entrant is on the board, and read the submission path. A board that re-scores on review is only useful if the review actually happened. The kit states the workflow in three lines: run the suite, upload the results, open the pull request.
-
Apply the tie band before you apply the adjective. Scores within 0.25 index points are ties by the board’s own published rule, and the index has no uncertainty intervals yet. A 0.06-point gap is not a lead. Neither, on the current panel, is the 0.45 points between the reference at 57.89 and Surogate Rune at 57.44, which ties with a Decider chat entry at 57.33 because the gaps are inside the band.
-
Choose with the calibration table, deploy with your own set. The index says who wins the suite. The ECE and Brier table says whose probabilities you can threshold. Neither says anything about your traffic. The operator who can answer “which model may do this job” has a written set of cases with answers their own experts signed off on, and reruns it when a new release lands.
-
Run the kit rather than the press release. It is MIT, it takes an edition flag, it resumes, and it runs on one rented accelerator. If you are about to route production decisions through a 6B model because a leaderboard said number one, one afternoon with the reproduction kit is cheaper than the incident review.
The uncomfortable truth about leaderboards
A board that publishes its panel, its chance baselines, its dropped benchmarks, its re-scored submissions and its tie band has done almost everything right. Nace AI announcing a number one while its own post says “official scores tbd”, and its own blog and product page give two different scores for two different checkpoints, has done something entirely normal. The failure mode is not either of them. It is treating a position on a five-day-old ruler as a property of a model.
Rank is a tuple, and the tuple is cheap to publish. The next time a decision model lands at number one, ask which edition, which panel, which formula, and whether the row is on the board at all. On this board, the answer to the last question is currently no, and the lead in the launch post was inside the band the board had already published.
Sources:
- Decision Index leaderboard Space,
multimodalart/jev-decision-index(primary:data/index.jsonanddata/methodology.jsonfetched 2026-09-27, edition v0.2.1 dated 2026-09-27,generated_utc2026-09-27T01:44:11+00:00, panel iddecision-index-0.2.1, 38-benchmark panel, 64 entrants plus the jev-1.13.0 reference, 120,340 requests, 119,898 scoreable, 442 exclusions, 110,313 source cases, 536,776 fields, corpus sha256b2b56d6f...92d5, seed 20260919, hardware “1 x NVIDIA RTX PRO 6000”, jev API cost $12.2567 across 291,826,668 input and 44,483,219 output tokens, jevbalanced_raw68.08,balanced_skill57.89,breadth_skill57.07; the changelog and hard rules quoted above are in the same file) - Decision Index reproduction kit,
apolinario/decision-index, MIT (primary: the scoring section stating “The board treats scores within 0.25 index points of the next one as tied (index02.ranks)”, the--edition 0.1flag, the hash-checkedsuite import, the Hugging Face Job recipe on one RTX PRO 6000, the submission workflow with “the results file is re-scored on review”, the parity note reproducing 49 entrants on the live 0.2 board, and the Jev 51.67 / AutoJev-27B 50.94 test fixtures; cloned and inspected 2026-09-27) decision_index/data/index-0.2.json(primary: the 0.2 panel specification, including"tie": 0.25)- Earlier editions as published by the Space (primary:
index-v0.1.jsonwith a 19-benchmark panel, 31 entries and the reference atbalanced_raw59.51 /balanced_skill46.26, andindex-v2.jsonwith the 40-benchmark 0.2 panel and the reference at 51.67; both fetched 2026-09-27. On the 0.2 file the top row is Decider chat · Gemma-4-31B at 51.93, 0.26 above the reference, just outside the tie band) - @NaceAI on X, Drex launch, September 25, 2026 (primary, quoted verbatim: “#1 on the Decision Index. (official scores tbd) winning 23 out of 40 benchmarks”; 281 likes, 34 reposts, 460,221 views at capture via the free session wrapper)
- Introducing Drex, a small model that decides instead of writing, Nace AI (vendor claim: Drex at 51.73 on edition 0.2, ahead of Jev 1.13.0 at 51.67 and AutoJev-27B at 50.94, under 6B parameters, “twenty-two of the thirty-nine benchmarks scored for it so far”, and 117 wins to 92 with 47 draws against Jev across 256 board games on the Kaggle Game Arena harness)
- Drex product page, Nace AI (vendor claim: 52.82 on “Decision Index 0.2, rank 1”, “+1.15 over Jev 1.13.0”, “21 of 40 tests won”, under 10B parameters, “leads all 50 entries”, Drex 1.0 shown beside it at 51.73, scored 2026-09-25 with the official scorer on all 40 benchmarks, plus the page’s own “+0.00 points ahead” comparison block and the FAQ answer stating the model is trained on the official training splits and evaluated on held-out splits)
- @multimodalart on X, Decision Index 0.1 announcement thread, September 22, 2026 (primary: the 0.1 launch, the changelog thread on how jev fares per category, the open-source and reproducible claim, and the reply “calibration is up too now!” with the quoted post “added calibration metrics to the leaderboard … Decider 35B-A3B by @notmapika, based on Qwen3.5-35B-A3B-Base takes the lead as the most calibrated model on our benchmarks”; the calibration post had 75 likes and 8,286 views at capture)
- @multimodalart on X, 0.1 to 0.2 update, September 24, 2026 (primary: “updated Decision Index 0.1 → 0.2, better formula, +29 jev-like models, +21 benchmarks”, AutoJev-27B taking “the open lead, trailing jev by 0.8 points”)
- The Router Does Not Need to Reason, Denny Sentinel, September 25, 2026 (related prior coverage: the decision layer as its own model class, GLiNER2.5-Decide, and why a leaderboard win at 60% exact match is only safe behind a confidence threshold)