The Patch Grader Was the Problem: Who Actually Verifies an AI Security Fix
In August, 1Password’s new security research lab, Off-by-1 Labs, published a benchmark with a headline most of the industry repeated verbatim: frontier AI models produced a “clean” security fix only 26% of the time. The number was picked up by ZDNet, The Register, and eSecurity Planet. Teams read it and concluded AI patching was not ready. Then, on September 15, Trail of Bits ran the same class of work through its own records and found the point was never the 26%. It is that the benchmark graded the models with the models themselves, gave its graders a buggy answer key, and credited a clean fix to patches that a broken grader could not catch. The verifier, not the patch, is what nobody has been able to measure.
This is exactly the boundary that keeps slipping in agent security. It is not that the agent is too smart or too dumb. It is that nobody has built an independent way to check whether the agent’s output is actually correct, so every headline about agent patching is really a headline about whichever grader happened to be in place.
The headline number, and the three choices that produced it
Off-by-1 generated 6,080 patches from ChatGPT 5.5 (default medium effort, with OpenAI’s Trusted Access for Cyber guardrails) and Claude Opus 4.8 (default high effort, with Anthropic’s Cyber Verification Program guardrails) against six recently disclosed, complex CVEs: four remote code executions, a use-after-free in Chrome, and a Linux local privilege escalation. Its average clean-fix rate, where a patch fully remediated the flaw without materially changing behavior, was 26.0%. That single figure, carried into the abstract, the conclusion, and the press release, is the product of the following choices:
- The sample was handpicked for difficulty. The six bugs were chosen because their fixes were complex. Clean-fix rates per bug ran from 3% to 60%. The mean has a standard error of roughly nine percentage points, which the report never discloses. A general claim about “AI patching” is being drawn from a cherry-picked hard case.
- Two of the nine prompt templates told the agent to apply the wrong fix. Those trials are 22% of the data. Combine “agents given bad advice” with “ordinary repair attempts” and the reported rate is partly a measure of how often the researchers chose to mislead the agents.
- More than a third of the trials were prohibited from testing. One evaluation mode stopped the agent from building or running code and accounts for 36% of the data. The headline pools “was not allowed to test its patch” with “was allowed to test and still failed.”
Trail of Bits did not just dispute all three. It re-ran the published test results, keeping only the trials where the agent could run code and was not told to apply the wrong fix, and got 2,634 of 3,067 patches (86%) defeating the supplied exploit. Blocking a provided proof of concept is not a complete repair, and ToB says so itself. But 86% versus 26% is the difference between “can’t patch” and “can patch when given the tools and honest instructions.”
The grader was grading itself, and its answer key had a bug
The deepest problem is not the sample. It is the instrument. Off-by-1 graded its patches with the very models under test: ChatGPT and Claude scored their own output and each other’s. Section 2.8 of their own paper reports the automated grades matched the authors’ human review only 65.9% of the time on the full five-category outcome. Every number in the abstract comes from those grades.
The answer key was wrong in a way the authors themselves documented. For the Linux “Copy Fail” CVE, the graders were told to treat the maintainers’ upstream fix as ground truth. That fix was a revert, a664bf3d, which carried an off-by-one bug corrected one commit later in 31d00156. The key given to the grader contained the bug and omitted the correction. A model that reproduced the kernel’s own bug matched the key and was scored as correct:
- 248 of the generated patches reintroduced the kernel off-by-one (129 ChatGPT, 119 Claude).
- The grader caught only 24 of those 248.
So the headline clean-fix rate was overcounting by an order of magnitude on one of six CVEs, and the paper’s authors recorded exactly how much on page 20 even as the abstract carried the 26.0%. For the Chromium use-after-free, between 38.5% and 41.9% of patches using the correct fix architecture merely moved the vulnerability into a callback instead of removing it, and the grader called many of those clean. Davi Ottenheimer (flyingpenguin) walked the same numbers and reached the same conclusion from the other side: the report’s own sections 4.4 and 4.9 prove the new-vulnerability rate was undercounted and the clean-fix rate overcounted, reproducibly from the released dataset.
The reason this matters is that a defender reading “26%” and trusting it will skip the tool entirely, leaving repairable vulnerabilities unpatched. A defender reading “agents produce clean fixes reliably” and trusting that without an independent check will merge a buggy patch. Both failures come from the same place: there is no trustworthy independent grader for security patches yet, and the people making the loudest claims built their grader out of the thing they were grading.
The two patches that prove plain humans are not the baseline
1Password’s paper offered no human baseline. Its only human comparison was an impression, not a measurement. But on the two codebases where a human fix exists to check, it measured humans without saying so:
- Linux. The kernel maintainers shipped the off-by-one into mainline in the “reference” fix, and corrected it one commit later. The clean-fix baseline for the platform’s own maintainers, on the bug the benchmark used as ground truth, was defective.
- freenginx. The Trail of Bits Patch the Planet agent-authored fix for a memory-safety bug in the embedded Perl module, PR #35, left one vulnerable code path open and introduced a new crash during request cleanup. freenginx maintainer Maxim Dounin closed the PR and committed his own fix,
cf26435. That fix covered all three vulnerable paths but introduced the same crash during cleanup.
Two authors, one agent and one human domain expert, working separately on the same bug at the same time, both missed the same early-request-cleanup trap. The original bug let Perl destroy a callback before freenginx used it; both fixes kept the callback alive but forgot what happened when a request ended early and released it.
That is the uncomfortable truth the “AI can’t patch” framing hides. On the only data point where a direct human comparison exists in the record, the human’s first fix was just as bad as the agent’s, for the same reason. Patching complex bugs is hard, and it is hard for people too. Trail of Bits owns a landed measurement of this: across 2,265 vulnerabilities in 236 of its own security assessments from 2024 to 2026, where the developers had detailed reports and knew their fixes would be reviewed, 283 first fixes (12.5%, one in eight) failed to fully resolve the issue, with a confidence interval of 10.5% to 14.5%.
What real-world acceptance looks like
Set the synthetic benchmark aside and look at what maintainers actually merged. Through Patch the Planet, Trail of Bits and OpenAI co-authored patches to open-source projects with engineers directing the work and completing the review. As of September 14, maintainers had merged or closed 186 of the public pull requests in the dataset, and the outcome distribution is nothing like a story of agent incompetence:
| Review outcome | PRs | % of merged |
|---|---|---|
| Accepted with no security-relevant revision observed | 91 | 72.2% |
| Accepted with a security-relevant revision observed | 33 | 26.2% |
| Indeterminate | 2 | 1.6% |
| Total merged | 126 | 67.7% of all 186 |
Of the 186, 60 were closed without merging. Only four were rejected on technical grounds. The rest were superseded by other work, declined for policy or maintenance reasons, or duplicates. ToB is explicit that maintainer acceptance does not establish every patch is correct, and its own post-merge review of ~33,500 subsequent commits found ten functional bugs, four automation bugs, and one performance bug introduced by its patches, and found no exploitable security vulnerability.
The systems lesson: patch generation ran ahead of patch verification
The sharpest line in the whole dispute belongs to Anthropic, quoted inside 1Password’s own report: “Progress on software security used to be limited by how quickly we could find new vulnerabilities. Now it’s limited by how quickly we can verify, disclose, and patch.”
That is the verifier problem, stated by the group that breeds these agents. Discovery is cheap now. Verification is the bottleneck, and the reason nobody has a trustworthy number is that the verifier is a model, or a human, or both, and neither has yet been structured to catch the class of bug that killed both freenginx fixes.
Trail of Bits is shipping the structural answer, not more slogans. Alongside the rebuttal it released two agent skills for exactly the gap:
post-patch-validationmakes the agent prove a patch before submitting it. It starts from the vulnerability report and the diff, then forces four moves: reproduce the original bug with a check that fails on vulnerable code and passes on the patched version; test a second path to the same failure (a different caller, a cleanup path); check for regressions and new vulnerabilities including sanitizer or bounded fuzzing; and, critically, treat a broken build as inconclusive, so a failed compile is never mistaken for evidence a vulnerability was reproduced.review-walkthroughturns a branch’s complete diff into an interactive, code-adjacent walkthrough for the engineer responsible for merging, and can prepare a GitHub review.
The interesting part is the discipline inside post-patch-validation: it refuses to let “I could not build it” masquerade as “it is fixed,” and it insists on a second path to the same failure. Both rules exist because the ones before them, the grading models in the 1Password benchmark, were structurally blind to exactly those two categories, the unverified build and the parallel vulnerable path.
What operators should change is not “trust agents” or “never trust agents.” It is that an AI patch is an unverified claim until an independent check has run, and the check has to be external to the model that wrote the patch.
- Never let the grader be the thing being graded. A model scoring its own patch, even with a different model’s help, propagates the same blind spot. Validation has to be execution-grounded, a test that demonstrates the bug is gone and that the build actually ran.
- Treat a reference fix as a claim, not a key. The kernel off-by-one that shipped in the “correct” answer is a standing reminder that upstream code can carry the bug you are trying to kill. Verify against the corrected history, not the first commit.
- Measure human first-fixes so you have a baseline. One in eight human first fixes fails under ideal conditions. Judge agent patches against that, not against a perfect ideal that no human meets.
- Require the second path and the cleanup path. Both freenginx authors missed the same case: what happens when a request ends early. A skill that makes the agent walk the cleanup path and a distinct caller catches the trap that killed both fixes.
The closing thesis is not that AI patching is ready. It is that we have been arguing about a number neither side could honestly measure, because the instruments were not up to the task. 26%, 86%, one in eight, 72%: all of them are really statements about whoever happened to be checking. The only durable move is to build an independent, execution-grounded verifier and put it between the agent and the merge. Until the grader is trustworthy, the patch numbers are just noise with opinions attached.
Sources:
- Trail of Bits: “1Password’s AI patching benchmark is misleading” (September 15, 2026; reanalysis of 1Password’s data, 186-PR outcome tables, 2,265-human-fix baseline, post-merge regression review, freenginx case)
- 1Password / Off-by-1 Labs: “Why AI-generated vulnerability patches still require expert human review” (Keith Hoodlet, August 6, 2026; 26.0% clean-fix headline, five scenarios, the six CVEs)
- Frontier Models’ Vulnerability Patches are Often F.L.A.W.E.D. (PDF) (sections 2.8, 4.4, 4.9; 65.9% grader agreement, 248 off-by-one reintroductions, Chromium callback blind spot)
- flyingpenguin (Daavi Ottenheimer): “Disinformation Pushed by 1Password: Their AI Patching Report is False” (August 6, 2026; the buggy answer key
a664bf3d, the 24-of-248 grader miss) - freenginx/nginx PR #35 (agent-authored fix, closed by maintainer)
- freenginx/nginx commit
cf26435(maintainer’s reference-counting fix) - Patch the Planet dashboard (Trail of Bits / OpenAI) (186-merged dataset as of September 14, 2026)
- Trail of Bits:
post-patch-validationskill - Trail of Bits:
review-walkthroughskill