The Weights Stopped Being the Model

The Weights Stopped Being the Model

On September 17, PrismML announced Ternary Bonsai 2 27B, a 27B-class multimodal model whose language weights are stored as three values, minus one, zero, or plus one, with one FP16 scale per group of 128. It keeps 98.2% of its full-precision benchmark average in a footprint the vendor quotes at 5.9 GB, against roughly 54 GB for the same base model in FP16. One day later, OrcaRouter published a runtime that changes that model’s behavior end to end without touching a single weight: one projection applied at 129 points inside the forward pass, driven by a direction vector of about 20 KB.

Two teams, two releases, and on their own each is an ordinary week in open weights. Together they describe something harder. The compression step stopped being a knob you turn on a finished model and became a property the training run installs. The moment that happens, the checkpoint stops being a self-contained description of what the model does, and every control built on the assumption that the weights are the artifact quietly stops holding.

The stakes are the local agent

The reason this matters now is that small local models have crossed into useful territory, and agent deployments are moving onto them for privacy, latency, and cost. A tool-calling loop that runs on a laptop does not send your repository to anyone, and the economics of a 27B model in a 6 GB envelope are the economics of putting an agent on hardware you already own. That is what PrismML is selling, and it is a real unlock.

It also means the compression decision now sits upstream of your agent’s reliability. If the compressed build is subtly worse at exactly the long-horizon reasoning your control loop depends on, no amount of prompt tuning or verifier design recovers it. You shipped a different model than the one you benchmarked, and the benchmark you ran was the one that could not see the difference.

The sub-4-bit collapse is selective, which is why casual testing misses it

PrismML’s model card publishes a 14-benchmark comparison in thinking mode, run with EvalScope on a vLLM backend on H100s. The headline rows:

BuildTrue bits/weightFootprintAverage (14)vs FP16
Qwen3.8-27B FP1616.054 GB86.32100%
Qwen3.8-27B UD-Q4_K_XL (“4-bit”)5.217.6 GB85.1898.7%
Qwen3.8-27B IQ2_XXS (“2-bit”)2.89.4 GB72.5984.1%
Ternary Bonsai 2 27B1.725.9 GB84.7898.2%

Read the average alone and IQ2_XXS looks like a 16% haircut. It is not a haircut. It is an amputation of specific capabilities, and the card’s own per-benchmark numbers show it:

  • AIME26: 57.50 against 94.58 for FP16
  • LiveCodeBench: 56.40 against 90.05
  • MATH-500: 84.60 against 99.80
  • MMLU-Redux: 88.93 against 89.09

Knowledge recall survives almost perfectly. Sustained chains of reasoning do not. Bonsai 2 27B holds the benchmarks that collapse: AIME26 at 95.83 and LiveCodeBench at 90.07, with math within half a point of the full-precision model in the vendor’s category rollup.

This is the sentence worth remembering: MMLU is the benchmark most teams run when they want to know whether the low-bit build “still works,” and it is the one benchmark in the suite that cannot see the failure. A quantization check that scores 89 against 89 will pass a build whose competition math and code reasoning have roughly halved. The failure lands hardest on exactly the workloads an agent runs, because an agent is a long chain of small steps where a small per-step error compounds before any verifier gets to look at it.

It is also not specific to this base model. The previous Bonsai release documented the same selective collapse on Gemma-4-31B, which is the point: below four bits, conventional post-training quantization fails in a pattern, and the pattern is set by the method rather than by the model family.

Why the low-bit representation had to move into training

The 98.2% figure comes from quantization-aware training, and the difference between that and post-training quantization is the whole story of this post.

A conventional low-bit build takes finished weights and rounds them. Whatever quality survives is whatever the rounding happened not to destroy, and rounding a 2-bit build of a 27B model costs you about fourteen points of average.

Quantization-aware training puts the low-bit representation in the loop during training, so the weights are fitted around the rounding instead of being rounded after fitting. OrcaRouter’s writeup states the consequence precisely: the rounding behavior PrismML trained the model to tolerate is a property of the training procedure, not of the quantizer, and you cannot re-run it after the fact.

That single sentence is the hinge. It means the low-bit model is not a compressed copy of the full-precision model. It is a different artifact with a different fitted solution, and the relationship between the two is not invertible.

The arithmetic that forces interventions into the runtime

Here is where it stops being a compression story and becomes a systems story.

The OrcaRouter release removes a refusal direction from Bonsai 2 27B. Conventional abliteration is a weight edit: estimate a direction r, then project it out of every matrix that writes into the residual stream.

W <- W - r(r^T W)

On a BF16 checkpoint you save the result and ship a new set of weights. On a ternary pack you cannot. A ternary matrix stores values in {-1, 0, +1} times a per-group scale, in a rotated basis, and W - r(r^T W) is a dense full-precision matrix. There is no ternary matrix equal to it in general. Saving the edit back means re-quantizing the edited weights, and re-quantizing them does not reproduce the quantization-aware training that produced the pack. You would be throwing away the one property the pack exists to have.

So the projection moves to inference, applied where each residual contribution is produced:

y <- y - alpha * dot(y, r) * r

alpha defaults to 1, alpha = 0 runs the published model. The implementation wraps 129 residual writers: 64 mlp.down_proj, 48 linear_attn.out_proj, 16 self_attn.o_proj, and model.embed_tokens. That coverage number is the part people get wrong when they try this themselves; wrapping self_attn.o_proj alone catches 16 sites and looks plausible while leaving most of the residual stream untouched. The repo ships scripts/selfcheck.py, which measures how much of the residual stream lies along the direction before and after intervention. Correctly instrumented, the remaining component drops to about 1e-6 of the residual norm, and the check warns if 129 sites are not detected.

The vendor’s own measurement is a clean experiment by construction, and worth naming as such: base and ablated are the same weights in the same process at alpha = 0 versus alpha = 1, so nothing about quantisation or checkpointing can confound the comparison. Their capability checks move inside noise at those sample sizes (MMLU 76.7 to 77.7, GSM8K 87.3 to 86.0, CMMLU 76.2 to 75.6), which is what bit-identical weights should produce.

And the underreported half: over-refusal drops too. On JailbreakBench’s benign prompts the published pack refuses 25.0%; at alpha = 1 it refuses 0.0%, and XSTest-safe goes from 5.2% to 0.4%. Refusal and over-refusal are the same knob. You do not get to turn down only the refusals you dislike.

What the intuition gets wrong

The first intuition is that low-bit compression costs a small, uniform amount of quality. It does not. It costs capability in a specific pattern, and the pattern lands on the reasoning categories that agent workloads consume most.

The second intuition is that a model is a file, so you hash the file and you know what you are running. That one is now wrong in three separate ways, and each is checkable.

The bytes on disk are not the bytes in the headline. The 5.9 GB figure is the PTQ1_0 GGUF packing, measuring 5.95 GB at 1.75 bits per weight. The MLX pack that the runtime ablates is 8.60 GB on disk, 7.67 GB of language model plus a 0.92 GB vision tower, because MLX’s grouped container stores a scale and a bias per group and costs 2.25 bits per weight for the same ternary values. Two packings of one representation, a 45% footprint difference, and no transfer of numbers between them.

The pack needs its own runtime or it lies to you. Bonsai 2’s weights live in a Hadamard-rotated basis declared as metadata, and the model card is explicit about the failure mode: ordinary MLX loaders skip the activation transform and the inverse embedding lookup, “so they return wrong output rather than an error.” A loader that silently computes the wrong thing is worse than one that refuses, and that is a general property of low-bit formats with folded transforms.

The behavior is no longer in the file at all. A 20 KB direction vector plus a runtime flag changes what the model says end to end. The pack stays bit-identical. If your provenance is a weights digest, that digest is identical before and after the behavior change, and it was identical all along.

What an operator should change

Move the eval to the deployment point. Stop validating a low-bit build on MMLU or on a generic capability average. Validate it on the tasks the agent actually runs: tool calling, multi-step code edits, long-horizon plans. If you cannot run those evals, say so, because the vendor’s average will not tell you, and the collapse is invisible on the cheap check.

Publish a model identity, not a weight hash. The thing you are running is weights digest plus packing format plus rotation metadata plus runtime version plus every side artifact that can steer behavior. That includes direction vectors, adapters, system prompts, and inference flags. For the runtime release above, the side artifact is roughly 20 KB. A manifest that hashes only the safetensors file is a manifest that cannot see it.

Treat the runtime as the trust boundary. Safety properties that used to be enforced by the choice of aligned checkpoint no longer live in the artifact. The release’s own responsible-use section says the technique is a research and inference-control mechanism, not evidence that any output is safe, and that deployments should apply their own access controls, policy enforcement, and security boundaries. That is the correct reading, and the obligation does not move to whoever wrote the ablation code. If a runtime flag can change model behavior, that flag belongs behind your approval and audit path, not in a config file someone can edit unnoticed.

Check the label against the true bit-width. Bonsai 2’s card makes this point against the industry: a widely used “2-bit” build of Qwen3.8-27B is 2.8 bits per weight at 9.4 GB. Names understate true average bit-width routinely, and the gap between 2.0 and 2.8 is the difference between a model you can use and one you cannot.

Use bit-identical provenance where you can get it. An immutable pack plus a separable intervention is a better audit story than an edited checkpoint, because the intervention has its own hash, its own size, and its own verification script. That is a real architectural gain, and it is available precisely because the weights were never edited.

What is not established

  • Every Bonsai 2 27B benchmark number in circulation is vendor-reported. The 84.78 average, the 98.2% retention, the category splits, all of it is PrismML’s own measurement, run with EvalScope on vLLM on H100s at xhigh reasoning effort. Nobody outside PrismML has reproduced it. Treat it as a strong, self-consistent claim, not as an established result.
  • The runtime evaluation is also first-party. OrcaRouter measured its own model, thinking off, greedy decoding, 64-token budget, with a rule-based opening-phrase classifier rather than an LLM judge. The refusal collapse is large and consistent across seven prompt sets; the capability movement is inside noise. One of the sets, SimpleSafetyTests, is understated because the ablated model answers self-harm prompts with a crisis redirect whose opening phrase the classifier does not recognize, so 18.0% is a floor, not a rate.
  • The central assumption is unverified, and the project says so. The direction was estimated from the BF16 base model the pack was trained from. The architecture and hidden basis are identical, so the projection is geometrically exact. But removing a vector exactly is not the same as removing the behavior it was estimated to represent, and how well a direction survives quantization-aware training has not been fully measured.
  • One direction is an assumption, not a law. The technique rests on Arditi et al. (2024), which found refusal mediated by a single direction. Follow-up work on other models has found geometrically distinct directions for different refusal categories. One vector swept over one alpha does not settle that.
  • Hardware and portability still favor the checkpoint. The runtime path is Apple Silicon / MLX only; the quantized matmul has Metal and CPU kernels but no CUDA implementation through mlx-cuda, so on an NVIDIA machine a 27B forward pass can take minutes. Conventional abliteration also produces one artifact that any compatible loader can consume, which the runtime approach does not.

The verifier moves to the deployment point

Denny Sentinel’s recurring argument is that control structure beats model choice, and this pair of releases moves a specific control out from under a specific assumption. The assumption is that a model’s identity is its weights, so hashing the weights tells you what you are running, and shipping an aligned checkpoint is a safety property.

That assumption is now false in a way that is not exotic and not a research edge case. It is the direct consequence of a compression technique that works, which means it is the direction the field is moving. When the training run decides the representation, the artifact stops determining behavior. A checkpoint hash tells you which bytes you have. It has never told you what the system will say, and it now cannot even tell you what the system will refuse.

The practical version is unglamorous: eval what you deploy, manifest everything that can steer it, and put the runtime behind the same approval and audit path you already put your code behind. The weights are no longer the model. Stop treating their digest as the answer.

Sources:

Keep reading