The Model Is Not Faster. The Decode Loop Is.

The Model Is Not Faster. The Decode Loop Is.

The fastest way to run a local language model may be to stop asking it for one token at a time.

That is the idea behind TensorFold, an open source inference engine from Ash Hart. It drafts several candidate positions, checks them in parallel lanes with the target model, then commits only the prefix that matches what serial decoding would have produced.

The target model still chooses the answer. TensorFold changes how much useful work happens in one verification pass.

That distinction matters more than the headline numbers. The interesting optimization is not a smaller model, a new quantization format or a more aggressive sampling trick. It is a control loop that spends the expensive model pass verifying several guesses at once, while keeping exactness as the contract.

Serial decoding leaves hardware waiting

Normal autoregressive decoding is sequential:

run the model -> emit one token -> append it -> run the model again

The dependency is real. The next token depends on the previous token. But the implementation also pays the cost of repeatedly reading model weights, launching kernels and moving through the same execution machinery for a single position.

TensorFold’s published method separates drafting from acceptance. A drafter proposes positions ahead of the target model. TensorFold places those positions into lanes, sends them through one target pass and compares the result with the target model’s own choices.

If the first two candidates match and the third does not, the first two are committed, the target’s token at the mismatch is used, and the remaining guesses are discarded. The next round starts from the committed sequence.

one stream
  -> draft A, B, C, D
  -> target verifies all candidate rows
  -> commit A, B
  -> take target token at C
  -> draft again from the new state

This is speculative decoding, but the engineering emphasis is unusually clear: the draft is disposable, and the verification path is the authority.

Exactness is the product feature

A speedup that changes the answer is a different model. That may be acceptable in some serving workloads, but it is not the same promise as accelerating one fixed decoding process.

TensorFold’s public site describes its rule as simple: accepted draft tokens must equal what serial decoding would have produced from the same cache state. Its repository documents row specific kernels intended to keep arithmetic consistent when several verification rows share a call. The project also exposes a comparison path using "draft": false, so a drafted reply can be checked against the serial path.

The release notes make the operational cost visible. TensorFold 0.3.4.1 reported 9 of 9 drafted replies matching the serial request on each supported model, plus resumed multi turn conversations matching fresh conversations. Those are project measurements, not an independent benchmark, but they are the right kind of evidence. The system is not merely reporting tokens per second. It is testing whether the optimization preserved the output contract.

That is the systems lesson most inference discussions skip. The verifier is not a cleanup step after the optimization. It defines which optimizations are allowed to exist.

The numbers only make sense with the machine attached

The TensorFold site reports separate measurements for separate hardware and workloads. Its current headline figures include 188 to 206 tokens per second for Nemotron 3.5 Lightning 30B-A3B on an M5 Max, 120 to 124 for Qwen3.8 27B on the same class of machine, and 88 to 92 for Qwen3.8 Flash Next on an M3 Ultra.

The project repository gives a more specific example. In a published M3 Ultra test, Qwen3.8 27B with a DFlash2 drafter reached 141.3 tokens per second for sampled code against a 38.2 token per second serial MLX baseline. That is a 3.70x result for one setup, with 64 token replies and thinking disabled.

The conditions are not decoration:

  • model weights and quantization affect memory traffic
  • prompt length changes prefill and cache behavior
  • code and predictable continuations draft better than fresh prose
  • thermals and MLX or CUDA versions affect the result
  • the baseline engine belongs beside the reported speed

TensorFold’s own site explicitly warns that the numbers are observations, not a cross model race. That is the honest framing. A local inference result without its machine, checkpoint, prompt and baseline is closer to marketing than measurement.

The gotcha is the prompt, not just the decode

A faster decode loop does not automatically mean a faster answer.

TensorFold’s 0.3.4 release made resumed prompts more exact by routing prefill through row exact kernels. That preserved a consistency property, but prompt processing became several times slower than MLX’s own path. Version 0.3.4.1 moved prefill back through MLX’s forward in fixed 2,048 token chunks, while keeping cache boundaries stable enough for resumed conversations to match fresh ones.

The release notes report 32,000 token cold prompt processing for Nemotron at 2,456 tokens per second on an M5 Max, compared with 2,448 for mlx_lm in the same session. For Qwen3.8 27B at 64,000 tokens, the figures were 465 and 467.

That repair is more important than a single decode chart. An agent does not experience tokens per second in isolation. It experiences time to first token, time between tokens and the delay introduced by context growth. An inference engine that accelerates generation while making every long prompt crawl has optimized the wrong part of the loop.

This is an agent infrastructure story

For an agent, the endpoint matters less than the behavior around it. TensorFold exposes an OpenAI compatible local API, so an existing coding assistant can point at http://127.0.0.1:8080/v1. The practical change is not that an agent suddenly has a smarter model. It is that the same model can spend less time waiting between tool decisions, code edits and verification steps.

That makes the verification contract useful beyond inference benchmarks. Agent systems already need to compare proposed state with accepted state:

  • a model proposes a tool call, policy checks it
  • a planner proposes a patch, tests verify it
  • a drafter proposes tokens, the target model verifies them

The scale is different, but the architecture is familiar. Separate proposal from authority. Make the acceptance rule explicit. Measure rejected work instead of hiding it behind a single throughput number.

The common intuition is that faster agents require a larger or newer model. Sometimes they do. But a large fraction of the wait comes from the control loop around the model: serial execution, repeated setup, cache handling, tool latency and unnecessary verification passes. Improve that loop and the same weights become more useful without pretending they became more capable.

What operators should measure

If you are evaluating a local inference engine for an agent workload, do not stop at the vendor’s top tokens per second number. Record:

  1. time to first token at the context lengths you actually use
  2. inter token latency during tool calls and code generation
  3. drafted tokens proposed, accepted and discarded
  4. output equality against a serial reference when exactness is promised
  5. cache behavior across resumed conversations
  6. memory headroom under the longest realistic prompt
  7. the workload and baseline behind every speed claim

TensorFold is interesting because its public artifacts expose much of that shape: the repository, the recipe book, the release measurements and the local serving path.

The uncomfortable truth is that inference speed is not a property of the model alone. It is a property of the model, the draft strategy, the kernels, the cache and the verifier acting as one loop.

The model did not get faster.

The loop got less wasteful.

Sources

Keep reading