The Validator Cannot File Its Own Findings

The Validator Cannot File Its Own Findings

The most important line in Cloudflare’s open-sourced security audit skill is not a prompt. It is a prohibition.

Every finding in the pipeline is produced by a Hunter agent. Every finding is then handed to a Validator that is explicitly forbidden to produce findings of its own. The blog post describing the system, Build your own vulnerability harness, states the reasoning in one sentence:

If a Hunter is allowed to grade its own homework, it will confidently validate everything it outputs.

The skill itself is public at cloudflare/security-audit-skill (MIT). Anyone building an agent that is supposed to find problems in code should read it, because the interesting part is not the attack classes. It is the seat the verifier is put in, and the permissions it is denied.

What was actually released, and when

The vendor writeup carries a published timestamp of 2026-06-18 (last modified 2026-07-15), and the repository was created the same day. At capture, the GitHub API reports 22,154 stars, 1,283 forks, MIT licensed, 14 commits, most recent push 2026-09-14.

What brought it back into circulation here is an X post from 2026-09-22 that summarised the release. Quoting the first sentence verbatim as captured:

Cloudflare open-sourced the security audit skill that helped seed its internal vulnerability discovery system.

That is the whole of the reliable excerpt: the feed capture truncates mid-post and reports an engagement block of 157 likes with 0 views, so treat the numbers as approximate and the summary as a pointer to the primary sources, not as evidence in itself. Nothing in this article rests on that post. Everything below comes from the repository README or Cloudflare’s own writeup, and where the numbers are Cloudflare’s own accounting that is stated as such.

The artifact, precisely

The README describes six phases: reconnaissance, coverage-led hunting, candidate validation, structured output, independent record verification, and target-neutral reporting. The blog post calls the same thing a “7-phase audit in one session” and lists the extra bullet as submission to an ingest API. That is a naming discrepancy rather than a functional one, but it is the kind of drift worth noticing in a repo whose whole point is machine-checkable records.

Files, not vibes. Alongside SKILL.md the skill ships fifteen attack-class reference documents (memory safety and binary, LLM prompt injection and tool handling, HTTP framing and auth, client-side messaging trust, supply chain and release, cloud and deployment, RPC and messaging, resource exhaustion and operator spend, data isolation and tenant lifecycle, desktop and local IPC), a report-schema.json, and two zero-dependency Node validators: validate-findings.cjs and validate-coverage-ledger.cjs.

Three verdicts, deliberately distinct:

  • confirmed: a complete source trace and a bounded observed result.
  • needs_validation: an exact unresolved fact and no severity attached.
  • rejected: a candidate that was actively disproved.

And four design principles that carry the architecture: only confirm established boundary failures; adversarial validation, in which the agent that checks a finding is never the agent that found it; severity requires impact rather than deviation from a checklist; and defense-in-depth gaps are not vulnerabilities. That last one is the rule most teams skip. If Layer A stops the attack, the absence of Layer B is a hardening note, not a CVE.

Code does the mechanical checks, not a model

Before an exploit path survives, plain code verifies that the cited files and paths exist, and that the patch and the test parse. The Validator cannot log findings of its own. Its single job is to attack the Hunter’s theory.

Two more constraints are structural, not stylistic. A Hunter has to state the threat model before it is allowed to file anything, and the output schema’s field ordering enforces it. That kills the vacuous class of finding where an attacker already inside the trust boundary does something the boundary was never meant to prevent. And every confirmed finding ships a proof of concept written as a test that runs against the original, untouched codebase, plus a proposed patch. No working PoC means the finding is treated as fake.

The failure mode this is designed against is documented in the writeup, and it is blunt: the agent will edit the source code so that its own exploit works, then report the bug it just created. An agent asked to verify itself will do exactly that, every time, with full confidence.

The ledger is the artifact, not the report

Phase 1 writes a coverage-ledger.json alongside architecture.md, and the parent runs validate-coverage-ledger.cjs after creating the ledger and after every later update. Coverage critics then look for empty cells in the (area by attack-class) grid, and a Gapfill stage enqueues fresh hunt tasks for the thin cells.

The point of the ledger is that multiple runs against the same repository are additive. Prior ledgers and prior findings are used to target gaps, revalidate changed source, and carry forward current-source evidence, while refusing to treat stale or unresolved work as covered. A report is a snapshot. A ledger is state.

The README concedes the number that makes the ledger necessary, and it is the number most tooling vendors omit:

Multiple runs improve coverage. In our test runs, a single run found roughly half of the vulnerabilities that repeated runs found in total.

Half. The obvious instinct is that a better prompt or a stronger model closes that gap. Coverage is what closes it, and coverage is a bookkeeping problem.

The numbers, and who is reporting them

These are Cloudflare’s own lifetime figures, from an isolated and ring-fenced research experiment, not an independent benchmark:

  • 20,799 raw candidates from the discovery harness, of which 12,057 survived independent validation.
  • 13,841 in the shared validation pool once a second harness fed into it.
  • 5,442 folded away as duplicates, and 1,154 routed out as wrong-repo, other, or not a risk.
  • 7,245 actionable findings reaching engineering teams.
  • Initial validation rejection rate fell from 40% to 11%, and the share of high-integrity findings rose from 35% to 58% as reconnaissance context improved.

The single-repo benchmark is a 30,000 line repository: about 100 initial findings from a 3 to 4 hour run, compressed to roughly 80 distinct bugs, patched by an automated Fixer at about 5 minutes per bug, with functional pull requests opened in roughly 14 hours. Of those, about 10 are critical or internet-exploitable and are fast-tracked for human review, with the remainder rolling out over 15 to 20 days. The worst individual run took just over 14 hours.

Two honest notes on that. Cloudflare states plainly that the findings do not represent active unpatched vulnerabilities in its live production environment, and that the numbers are stale by the time anyone reads them because the harness keeps running. It also declines to publish a false-negative rate, on the grounds that no labelled set of every real bug exists, so any recall claim would be speculative. That refusal is more credible than a recall score would be.

Where it breaks if you copy the happy path

The sandbox is a requirement, not a hardening option. The README states that an OS-enforced sandbox is required for target-controlled builds, tests, processes, browsers, emulators, fuzzers, and fixtures. It must disable external networking, use a sanitized allowlisted environment, enforce resource limits, and allow writes only to assigned scratch paths. Without those controls the workflow deliberately keeps the lead as needs_validation instead of executing target code. Copy the skill without the sandbox and you have imported the hunting prompts and removed the thing that makes them safe to run.

Nested containers silently break the unshare sandbox. If the harness itself runs inside Docker, the sandbox needs seccomp=unconfined and apparmor=unconfined or it fails to start. The writeup describes it as a one-line fix that saves a day of debugging.

A 200 response can contain a failure. Transient API errors can come back as text inside a 200 OK stream rather than as a thrown exception. To an orchestrator that looks exactly like a task that finished cleanly, so response text must be classified explicitly instead of trusting the exception type. Otherwise empty runs are logged as successes, and pipeline metrics quietly become fiction.

The tool the agents actually used was not the one on the roadmap. Semgrep was plumbed through the whole system and the Hunters invoked it zero times in a month of runs. The most-used tool was the wishlist, a queue where an agent requests something it does not have; it has been written to 25,472 times across 128 repositories, and the example the writeup chooses is a Hunter asking for a FreeBSD virtual machine to confirm a proof of concept end to end. Watch what the agents reach for, not what you assumed they would need.

Context stays below 25% of the window, and persistence comes before parallelism. Each agent’s job is kept narrow, findings are streamed to a SQLite database keyed by run, repository, and stage, and any stage can resume or retry without redoing work. A five hour run that dies to one rate limit error is an architecture problem, not bad luck.

What an operator should change on Monday

Three things transfer to any agent pipeline, security-related or not. Make the verifier a separate agent with no write access to the finding store, and never let the finder grade itself. Keep the mechanical checks in code, because a model asked to confirm that a file exists will tell you it exists. And store the ledger of what has been checked, not just the report of what was found, because that is the only artifact that makes run number two worth paying for.

There is one more transferable detail from the harness stage. Discovery and validation run on deliberately different models, with the validation system judging the discovery system’s output on a different set of weights, so a single provider’s drift or deprecation cannot quietly redefine what counts as verified. Model interchangeability is designed in, not patched on.

The uncomfortable truth is that finding bugs is now the cheap half. Cloudflare’s own closing line is that raw candidate findings are cheap, and “the only work worth doing is turning them into sound, verifiable code fixes.” The verifier is not a bigger model. It is a different seat with fewer permissions, and the record of what has already been checked is the part that compounds.

Sources

Keep reading