The Patch Passed. The Server Still Failed.

The Patch Passed. The Server Still Failed.

A coding agent can pass the tests in front of it and still fail at the thing the software was built to do.

That is the central result of SWE-Serve, a new benchmark from NVIDIA and the SGLang community. Across 19 tasks that included live-serving checks, the same patches passed 69.4 percent of the time when those checks were excluded, but only 45.9 percent when the full verifier started a real server and exercised it.

Put differently, about one in three patches that passed the other checks failed when the system had to serve a real model through its normal interface.

This is not a story about a weak model. It is a story about a weak definition of done.

The test that matters is the request path

SWE-Serve turns 83 merged SGLang pull requests into 53 executable tasks covering six parts of inference engineering:

AreaTasks
Speculative and advanced decoding14
Model and backend enablement12
Kernels, quantization, and performance8
Serving APIs and runtime correctness8
Caching and runtime state7
Distributed execution and scheduling4

The task is not simply to edit a repository until a visible unit test turns green. The agent receives a containerized checkout from before the target change, then its patch must satisfy a hidden verifier on the declared hardware.

Nineteen tasks start a real server. The verifier can check model loading, OpenAI-compatible requests, batched generation, image inputs, log probabilities, expert routing, and other behavior that only appears after the full request path is alive.

That is the important design choice. The benchmark does not ask whether the patch looks plausible in source code. It asks whether the deployed shape of the program produces the expected result.

Why local tests create false confidence

A serving stack is a chain of stateful boundaries:

request
  -> routing and scheduling
  -> model loading
  -> execution and kernels
  -> KV cache and runtime state
  -> response serialization

A patch can be locally coherent while breaking one of the handoffs.

A model registration change may parse correctly but fail to load real weights. A scheduler change may pass a narrow unit test but return requests out of order under batching. A cache change may work for one prompt and corrupt state when several requests share a server. An API change may satisfy a function test while failing through the public protocol that clients actually use.

Those are not exotic edge cases in an inference server. They are the product.

SWE-Serve reports that its 19 live-serving tasks contained 276 serving tests. Of those, 242 were sourced or adapted from SGLang, with the remaining tests covering behavior introduced by the target changes. The benchmark’s point is not that every verifier is perfect. It is that a verifier which never starts the service is measuring only a slice of the engineering problem.

The verifier is the real benchmark

The phrase “the agent solved the task” hides a policy decision. Solved according to which observation?

SWE-Serve makes that observation explicit. A patch counts only when it passes the complete verifier on the declared hardware. NVIDIA says the unmodified repository fails the new-behavior tests while passing regressions, and that a reference patch passes the complete verifier. The benchmark also blocks web access during evaluation, so the agent cannot retrieve an upstream solution while it works.

That structure matters more than the leaderboard headline. It separates three things that are often collapsed:

  1. The agent’s explanation of what it changed.
  2. The repository tests that happen to run.
  3. The externally observable behavior that the service must provide.

Only the third one tells an operator whether the system works as deployed.

This is the same control-loop problem that appears in security agents. A model proposes a vulnerability, but a separate validator must observe an exploit. A deployment agent reports healthy, but an independent check must query the endpoint. A coding agent reports success, but the running service must answer correctly.

The verifier is not a reporting accessory. It defines the boundary between a hypothesis and a result.

Cross-domain work is where the score falls

SWE-Serve also reports a second gap. Its 26 tasks confined to one runtime domain had a 69.0 percent pass rate across the best setting for each evaluated model. The 27 tasks spanning more than one domain had a 47.7 percent pass rate, a difference of 21.3 percentage points.

The pattern is familiar to operators. Bugs hide at boundaries, not inside the cleanest abstraction.

Request handling meets scheduling. Scheduling meets model execution. Execution meets cache management. Cache state meets distributed coordination. Each component can look correct in isolation while the composition fails.

This is why adding a larger model is not a complete response. A stronger model may navigate more of the codebase, but it still needs tools that expose the runtime and a verifier that can observe the right state. The agent cannot reason its way around a missing test oracle.

What the benchmark does not prove

SWE-Serve is useful, but its pass rate is not a universal measure of coding-agent intelligence.

The first release covers SGLang and 53 tasks. Twelve tasks run on CPU and 41 use a single NVIDIA H100. It does not evaluate other inference engines, multi-GPU execution, or multi-node serving. NVIDIA also says a pass means only that the patch satisfies the benchmark verifier. It does not establish that the patch is ready to merge, deployable in every environment, or endorsed by SGLang maintainers.

Those limits make the result more credible, not less. A benchmark with a declared scope is more useful than a universal claim built from an opaque test harness.

The reported model scores also show why a single leaderboard number is a poor routing policy. Across the best tested configurations, mean pass@1 ranges from 34.6 percent to 75.5 percent. Four models tied at 64 percent, yet their reported mean cost per task ranged from $0.95 to $7.24 and their mean wall time from 25.5 to 99.9 minutes.

The model is one variable. The task family, verifier, hardware, effort setting, latency budget, and cost are other variables.

What agent builders should change

If your coding agent edits infrastructure, make the runtime part of the task contract.

Start with the user-visible path. If the code change affects an API, start the service and make the request. If it affects model loading, load the real model or a representative artifact. If it affects batching, test concurrent requests and response ordering. If it affects caching, test repeated and interleaved state transitions. If it affects scheduling, measure behavior under the queue shape the system will actually see.

Then separate the lanes:

  • Exploration: the agent may generate hypotheses, partial patches, and likely explanations.
  • Verification: an independent harness checks the resulting state through the real interface.
  • Promotion: only a patch that passes the declared verifier can reach review, deployment, or automated merge.

Keep the verifier outside the agent’s narrative. Do not let the same process both declare success and define what success means.

Record the environment as part of the result. Hardware, model version, dependency lockfile, request shape, timeout, and test selection all change what a serving patch means. A green result without those details is not reproducible evidence.

Finally, test the boundary where the system becomes useful. For an inference server, that is not the import statement. It is the request that loads a model, consumes runtime state, and returns a correct response.

The deployment substrate is part of the intelligence

The interesting part of SWE-Serve is not the claim that agents can edit a complex codebase. We already know they can produce patches that look convincing.

The useful finding is that live serving changes the answer. A verifier that reaches the running system rejects a large class of patches that ordinary checks would accept. The benchmark turns deployment behavior into part of coding competence.

That is the direction agent evaluation needs to take. Measure the control loop, not just the text of the patch. Give the agent a real environment, then ask a separate mechanism whether the environment works.

A patch is not done when the tests pass. It is done when the system survives contact with the interface its users depend on.

Sources

Keep reading