The Harness Is Part of the Model

The Harness Is Part of the Model

A model can keep the same weights and score 62% in one agent harness, then 33% in another. The gap is not a footnote about prompting. It is evidence that the harness is part of the model’s operating environment.

That is the important result behind FineEnvs’ new multi-harness reinforcement-learning release. The team trained an open LFM2.5-2.6B checkpoint through four real agent harnesses and reports 54.2% pass@1 across 1,000 graded task and harness cells. The checkpoint and evaluation details are public on Hugging Face.

The claim is self-reported, single-run, and tied to a specific task set and pinned harness versions. It is not a universal leaderboard result. It is still a useful systems signal: the interface around a model changes what the model can do.

The proxy is the key idea

The obvious way to train an agent for several products is to edit each product’s integration. That creates a maintenance problem and risks training on one implementation’s quirks.

FineEnvs takes the opposite route. Its multi-harness guide describes a capture proxy that sits between the harness and the model. The proxy speaks the API formats used by the target agents, records the token IDs and log probabilities sampled by vLLM, and passes training sequences to TRL’s asynchronous GRPO implementation.

The harnesses remain unmodified. The training system sees the same kind of calls the agent would make in practice, including tool interactions and the resulting task outcome.

That boundary matters. It turns the harness from an unexamined wrapper into part of the training distribution.

The authors state on X that the proxy supports OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and Gemini formats, and that ten harnesses can run through it unmodified. The announcement thread is discovery context; the model card is the source for the released checkpoint and its reported evaluation.

What the released numbers actually say

The released main checkpoint is step 1,000. It was evaluated on 250 fixed SmolDataEnvs test tasks under four harnesses, producing 1,000 graded cells.

HarnessCorrectPass@1
OpenCode124 / 25049.6%
Claude Code122 / 25048.8%
Codex134 / 25053.6%
Mini-SWE-Agent162 / 25064.8%
Overall542 / 1,00054.2%

The base-to-checkpoint comparison also reports 31.1% fewer tool calls across 356 task and harness pairs solved by both models. That is a matched-success comparison, not a reduction across every task, and it does not isolate how much of the change came from the efficiency bonus.

The training reward was correctness multiplied by a tool-call efficiency term. Incorrect answers received zero. Missing or unverified action counts did not receive the efficiency bonus. That is a better contract than rewarding short traces unconditionally, because a short wrong answer is not efficient.

There is also an important evaluation caveat. The model card says the best observed step 700 checkpoint, at 54.6%, was selected using this test set rather than a separate validation set. That makes the number useful as a reported result, but weaker as an unbiased estimate of generalization.

Why one harness is not enough

Agent harnesses decide more than where a prompt is sent. They determine the tool schema, how errors are represented, how much history is retained, when a tool call is retried, how files are mounted, and what counts as task completion.

Two harnesses can expose the same model to different effective programming languages. One may return a structured tool error. Another may put a stack trace into ordinary text. One may preserve a long tool transcript. Another may summarize it. The model is not solving the same control problem in both environments.

This is why copying a successful rollout from a larger model is a weak shortcut. The FineEnvs guide reports that imitation training on 3,189 successful Qwen3.8-27B rollouts plateaued at 47.5%, below both reinforcement-learning runs. The important contrast is not that RL always wins. It is that behavior learned inside the environment can matter more than reproducing a transcript from a different policy.

The result is an argument for environment-aware training, not a license to treat one benchmark as proof that a small model has become a general coding agent.

The deployment consequence

Teams building agents often version the model and loosely version the harness. That is backwards when the harness changes the control loop.

A production evaluation should record at least:

  • The model revision and tokenizer.
  • The harness name and exact version.
  • Tool schemas, retry behavior, and sandbox configuration.
  • The task set and grading contract.
  • The number of tool calls and how those calls were counted.
  • Whether the score is pass@1, pass@k, or a retry-assisted result.

A model score without those fields is an incomplete measurement. A 54.2% result on 1,000 specific cells does not tell an operator what will happen after changing the harness from Claude Code to an internal runner. The per-harness table is not noise. It is the operational result.

The proxy approach also suggests a cleaner architecture for future training. Keep the harness stable, capture the boundary between harness and model, and train against several real callers. That lets teams improve the policy without turning every harness into a custom research fork.

The model is the whole loop

The open-weights story is usually framed around parameter count, context length, and a download link. Those properties matter, but agents are not evaluated in a vacuum.

They are policies embedded in a loop: model, harness, tools, sandbox, errors, retries, and verifier. Change the loop and you change the task.

The interesting part of this release is not that one 2.6B model reached one more benchmark percentage point. It is that the training artifact treats the harness as a first-class variable and publishes the mismatch instead of hiding it behind one average.

The model is not just the weights. For an agent, the model is the weights running inside a control loop. Train and measure the loop you intend to deploy.

Sources

Method note: this article was selected from the two required free, session-authenticated twsearch radar queries. The reported scores and limitations come from the public model card. The X post was retrieved read-only and is quoted only for the proxy description.

Keep reading