Near-Lossless Is a Claim About a Calibration Set You Cannot Read
“The point is not 3-bit.”
“The point is what survives at 3-bit.”
Those two lines are the close of the headline block on OrcaRouter’s model card for OrcaSAQ-2-27B, published September 24. What survives, per the card, is this: a 54 GB BF16 checkpoint compressed to 12.3 GB, with 93.2% token-level Top-1 agreement against the BF16 reference, 0.031 mean KLD, and perplexity of 5.6468 moving to 5.6482.
Two days later, the same account announced a cyber variant: “We compressed our most popular local cyber model down to 15.7 GB,” with the metric block reading “54.7 → 15.7 GB / 262K context / 94.4% Top-1 agreement.” The post stands at 977 likes, 95 reposts and roughly 188K views at capture. Its model card is gated. Fetching orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF unauthenticated returns “Access to model orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF is restricted,” and the Hugging Face API reports gated: auto.
So there are now two numbers, 1.2 points apart, produced by one team four days apart, both labelled Top-1 agreement, and only one of them ships with the corpus, the token count, the baseline and the evaluation path that make an agreement rate mean anything.
The uncomfortable truth is that the compression ratio is the part of this release you can verify, and the fidelity claim is the part you cannot.
What the number actually measures
The dense card is unusually specific, and that specificity is worth reading before it gets lost. The fidelity table is measured on WikiText-2 over 16,376 predicted tokens, through the same evaluation path, comparing the quantized model’s top-1 prediction at each position against Qwen3.8-27B BF16.
Run the arithmetic that the card leaves to the reader. An agreement rate of 93.2% over 16,376 positions leaves roughly 1,114 positions where the top prediction differs from the reference. Three published metrics, three different units: a rate (93.2%), a mean distributional divergence (KLD 0.031), and a ratio of two perplexities (5.6468 to 5.6482, which is the “+0.02%” the card quotes). None of the three says which tokens differ, or whether the differing ones sit anywhere near a decision.
The cyber variant’s 94.4% is vendor-reported and, as of this writing, not reproducible even in principle from the public artifact: the card is behind an access gate, the announcement carries no token count and no sentence identifying the evaluation corpus, and the BF16 reference for that build is a different checkpoint (the abliterated orcarouter/Qwen3.8-27B-Uncensored, not the base Qwen3.8-27B the dense card measures against). The claim may well be accurate. It is not checkable, and the 1.2 point gap between it and the sibling’s number is not a comparison between two measurements of the same thing.
Precision is allocated, not spread
The card lists a decoder average of 3.21 bits per weight across a 64-layer model. That average is the interesting artifact. A uniform 3.2-bit build would spread quantization error roughly evenly across the network; this one describes itself as “sensitivity-aware mixed-precision,” which means an allocator decided which matrices got more bits and which got fewer, under a sensitivity estimate fitted against some calibration distribution.
The card then states, in its Method section: “Detailed quantization methodology, calibration strategy, precision allocation and packing techniques are not currently disclosed.”
That one sentence is the whole trust story for a compressed security model. If precision is allocated by sensitivity to a calibration set, then the model is cheapest to compress exactly where that calibration set is thinnest. The resulting errors are correlated with an unpublished editorial choice rather than sprinkled evenly at a 6.8% rate. And the calibration set plays the role of the verifier for the compression step itself: whatever it under-represents got rounded hardest and was never checked during allocation. For an agent that emits tool-call JSON, unified diffs, file paths, hex addresses and error taxonomy, the token classes whose exact form decides control flow are not the classes a Wikipedia prose corpus weights heavily. That is inference, and it is the inference that matters when the model is wired to a shell.
The card’s fidelity number and the card’s serving settings describe different regimes
This one is worth separating carefully, because it is a mismatch between two sections of the same document rather than a vendor overclaim.
The card’s fidelity table compares top-1 predictions: the argmax token at each position against the BF16 argmax. The card does not state the decoding configuration used for that comparison.
Separately, the card’s recommended sampling block for deployment reads:
temperature = 1.0
top_p = 0.95
top_k = 20
At temperature 1.0 with top_p 0.95 and top_k 20, the model is sampling from a truncated distribution and will, by construction, often emit a token that is not its own argmax. So the number carrying the “near-lossless” label measures a regime in which the model always takes its single best guess, while the settings the card tells you to serve with deliberately abandon that behaviour. Both statements are correct. Read together, they mean the agreement figure is a property of the next-token distribution, not a promise that two runs, or two builds, will produce the same trajectory. Treat that as reasoned inference from two primary sections, not as a documented contradiction.
Why a per-token rate does not compose into agent reliability
If you take 93.2% at face value and assume divergences are independent, the probability that a 200-token tool call is token-identical to the reference is 0.932^200, about 7.6e-7. At 500 tokens it is 5.1e-16. That calculation is not a prediction of anything, since token divergences are neither independent nor uniformly harmful. It is a demonstration that token agreement is measured in the wrong unit for an agent, because a single divergence is either inert or it changes an action.
The card makes the compounding argument itself, in prose, in its long-horizon section: “A small model error can change a tool call. That changes the environment state. The changed state affects every decision that follows.” The same card then reports its fidelity table on WikiText-2, which contains no tool calls and no environment state. Both things are in the document. The honest reading is not that the vendor is hiding a failure mode; it is that the document contains an argument that its own headline metric cannot see, and states so: “93.2% Top-1 agreement means some token decisions differ from BF16,” and “+0.02% PPL is a model-fidelity measurement and does not guarantee identical downstream performance.”
The card’s own remedy is the right one, and it is the same advice this site gives about every routing decision: “For agent deployments, benchmark against the actual tool schema, context distribution and reasoning budget used in production.”
The baseline has three values depending on units
Hugging Face metadata for Qwen/Qwen3.8-27B reports 27,781,427,952 parameters in BF16. At two bytes each, that is 55,562,855,904 bytes: 55.6 GB in decimal units, or 51.7 GiB in binary units. The dense card calls the reference 54 GB. The tweet calls the cyber baseline 54.7 GB.
Nothing here changes the conclusion, and the uncensored base for the cyber build is a different checkpoint that may genuinely add parameters, so 54.7 GB could be a real measurement of a real file. The point is narrower and still worth making: the denominator of every compression ratio on these cards depends on which unit convention was used, and the published figures do not resolve to any single convention. When a page sets out to prove that a ratio is impressive, the denominator deserves the same scrutiny as the metric.
What an operator should take from this
- Keep the ratio, drop the adjective. “54 GB to 12.3 GB” is a fact about a file, verifiable with
ls -lafter download. “Near-lossless” is a claim about an evaluation you did not run, on a corpus that may not be published at all. - Build the eval before you route. A set of cases pulled from your own workload, with answers your own people signed off on, is the only fidelity measurement that transfers. The vendor says this in its own card. It is also the only way to decide between a 12.3 GB build and a 15.7 GB build on evidence rather than on the size of the checkpoint.
- Treat a fidelity number with no corpus and no token count as a marketing scalar. The dense card published its basis, its token count and its reference model. The gated card published a percentage in a tweet. Both are labelled Top-1 agreement.
- Pin the exact build. The quant family moved twice in 72 hours: the dense checkpoint was created September 24, the cyber GGUF on September 26 at 03:37 UTC and last modified that day at 19:22 UTC. A family that re-cuts weekly is not a series you can average across.
- For security work, buy locality, not compression. The genuinely verifiable property of a local cyber model is that source code, logs and findings do not leave the host. That is worth real money and it holds regardless of whether agreement is 93.2% or 94.4%.
- Note where the refusal boundary went. The cyber build’s base is an abliterated checkpoint (the Hugging Face tags on orcarouter/Qwen3.8-27B-Uncensored read “abliterated,” “uncensored,” “red-teaming”). This site covered the runtime version of that trade on September 19. An uncensored local model does not remove the safety decision; it relocates it into your harness, your approval gates and your logging. Test what the model will actually do before wiring it to a terminal.
The closing thesis
A compression ratio is a deployment fact. Near-lossless is a claim about a calibration distribution that the vendor states is not disclosed, evaluated on a corpus whose token count you may not be able to read, under a decoding regime that the recommended serving settings then abandon. Those are two different kinds of statement, and only one of them is yours to verify. When a headline number travels without its corpus, it stops being a measurement and becomes a scalar, and no agent should be routed on a scalar.
Sources
- OrcaRouter model card:
orcarouter/OrcaSAQ-2-27B(primary: 54 GB to 12.3 GB, 3.21 bpw, PPL 5.6468 to 5.6482, 93.2% top-1 agreement, 0.031 mean KLD, WikiText-2 over 16,376 predicted tokens, the undisclosed-method statement, the temperature 1.0 / top_p 0.95 / top_k 20 recommendation, the long-horizon compounding argument, the limitations list) - Hugging Face API:
orcarouter/OrcaSAQ-2-27B(primary: created 2026-09-24, 262K context tag, Apache-2.0, vLLM, exl3, 1,330 downloads and 163 likes at capture) - Hugging Face API:
orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF(primary: gated statusauto, created 2026-09-26T03:37Z, last modified 2026-09-26T19:22Z, base modelorcarouter/Qwen3.8-27B-Uncensored, 313 downloads and 63 likes at capture) - OrcaRouter model card (gated):
orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF(primary: unauthenticated fetch returns “Access to model … is restricted”) - X: @OrcaRouter announcement of OrcaSAQ-2 Cyber 27B Uncensored GGUF (primary: verbatim text “We compressed our most popular local cyber model down to 15.7 GB”, “54.7 → 15.7 GB / 262K context / 94.4% Top-1 agreement”; September 26, 2026, 19:33 UTC; 977 likes, 95 reposts, 27 replies, approximately 188K views at capture on 2026-09-27)
- Hugging Face API:
Qwen/Qwen3.8-27B(primary: 27,781,427,952 BF16 parameters, created 2026-08-05) - Hugging Face API:
orcarouter/Qwen3.8-27B-Uncensored(primary: abliterated and uncensored base checkpoint, created 2026-08-18, gated) - Denny Sentinel: The Weights Stopped Being the Model (prior coverage of OrcaRouter’s runtime ablation release and why compression moved inside the training procedure)
Web view counts captured through a syndication feed are approximate and are reported as such; likes, reposts and replies are the figures the capture returned at 2026-09-27 15:00 UTC. The 94.4% agreement figure is reported by the vendor and is not independently verifiable while the model card remains gated. Nothing in this post claims the compressed models perform badly, and no benchmark was run for this article.