The Bug Bounty Became a Spam Filter

The Bug Bounty Became a Spam Filter

A bug bounty program is supposed to turn outside research into a second security team. It cannot do that if the intake queue becomes a landfill of plausible-looking guesses.

Tom’s Hardware reported that Google paused product vulnerability submissions in part of its open-source bug bounty program until 2027, amid an influx of invalid AI-generated reports. A widely circulated X post from Tom’s Hardware made the same claim on October 3, 2026.

The claim is reported here as an alleged program change based on those sources. The source material does not establish the exact volume of invalid reports, the precise cost to Google’s triage team, or whether AI-generated submissions were the sole cause.

Those details matter, but they are not the main story. The systems lesson is clearer: generating a vulnerability report is easy. Proving that it deserves a human’s time is the missing layer.

The queue is the security boundary

A modern vulnerability report is a structured request for scarce attention. It includes a target, a suspected flaw, reproduction steps, impact, and often a proposed fix. Language models are good at producing that shape even when the underlying claim is false.

That creates a dangerous asymmetry:

  • A researcher or agent can create a polished report quickly.
  • A maintainer still has to reproduce the behavior, identify the affected version, assess impact, and decide whether it is in scope.
  • A false report consumes nearly the same first-pass attention as a real one.

When the cost of generating claims falls faster than the cost of checking them, the intake system becomes the bottleneck. The program has not become less secure because models can write prose. It has become less useful because the queue no longer communicates signal.

This is the same control-loop problem that appears inside autonomous security agents. Discovery is not validation. A stack trace is not a vulnerability. A suspicious data flow is not an exploit. A report is not a finding until a verifier closes the loop.

Why language quality makes the problem worse

Bad reports used to be easy to discard. They were vague, incomplete, or obviously copied. AI-assisted reports can be different. They may contain the right file names, a realistic severity score, and a reproduction narrative that sounds like a security engineer wrote it.

That polish is operationally expensive. It pushes triage away from cheap syntactic filtering and toward expensive semantic checks.

The alleged pause is therefore not evidence that AI has made vulnerability research useless. It is evidence that the submission channel was designed around a world where producing a plausible claim had a higher cost.

The uncomfortable truth is that a better report template does not solve this. A model can fill in a template. It cannot, by formatting alone, establish that the vulnerable path is reachable, that the behavior is new, or that the impact is real.

The missing verifier

A useful security-research pipeline needs at least one stage that the report generator cannot substitute for itself. Depending on the program, that stage might include:

  1. Version anchoring. Does the claimed behavior exist in a supported version and disappear in a fixed version?
  2. Reachability. Can an untrusted input reach the claimed sink under the stated deployment conditions?
  3. Reproduction. Can a clean environment reproduce the behavior without hidden state from the reporter’s machine?
  4. Impact confirmation. Does the result cross the program’s security boundary, rather than merely produce an error or a local crash?
  5. Scope and novelty. Is the report in scope, and is it distinct from an existing issue?

An agent can help with every one of these stages. It can build a minimal test case, bisect versions, trace data flow, compare a patch, and assemble evidence. But the pipeline should record those artifacts separately from the prose summary.

A report that says “remote code execution” should carry a test result. A report that says “information disclosure” should identify the principal that can read the data and the boundary it crosses. A report that says “fixed by this patch” should show the before and after behavior.

Without those artifacts, the report is a hypothesis with a professional tone.

Do not make the model the submit button

The tempting response is to ban AI-generated reports. That may reduce volume, but it misses the architecture problem. A human can submit an unverified claim too. A model simply makes it cheaper and more scalable.

The better design is to separate the stages:

model proposes claim
        ↓
agent gathers evidence
        ↓
independent verifier runs checks
        ↓
human or policy gate decides submission

The final gate should not ask whether the prose sounds credible. It should ask whether the evidence contract is complete.

For a security agent, that means storing command output, affected revisions, test fixtures, traces, and patch comparisons as first-class artifacts. The narrative should be rendered from those artifacts, not used as their substitute.

For a bounty platform, it means making high-value evidence easy to submit and low-value prose hard to mistake for proof. Structured fields for affected versions, deterministic reproduction, and environment details can help. So can rate limits, duplicate detection, and staged disclosure workflows. None of those replaces technical validation, but they keep the verifier focused on the claims most likely to matter.

The operator consequence

If you run an AI-assisted vulnerability workflow, measure more than report count. Track:

  • the percentage of claims reproduced independently,
  • time from intake to first technical decision,
  • duplicate and out-of-scope rates,
  • evidence attached per report,
  • and how often a human reviewer had to reconstruct the entire experiment from prose.

A rising submission count can be a regression. If every additional report adds more triage than evidence, the system is scaling its noise channel, not its research capability.

The same rule applies to internal agent evaluations. Do not reward an agent for opening more issues unless the issue carries a verified behavior, a bounded impact claim, and a reproducible artifact. Otherwise the reward function selects for confident narration.

The queue needs a proof layer

The alleged Google pause is a warning about infrastructure, not just etiquette. Security programs are control loops. They accept observations, spend human attention, and produce decisions.

AI reduces the cost of observations. That makes the verifier more important, not less.

The next generation of security agents will not win by writing the most convincing vulnerability report. They will win by producing the smallest package of evidence that lets another system or person say yes, no, or not yet.

A report without proof is not a finding. It is queue pressure with markdown.

Sources

Keep reading