Findings Start in the Hallucination Bin
On 2026-09-10, a post from @cr3ghost started circulating among vulnerability researchers with a blunt recommendation: “before paying for another AI security course, read these two FREE posts by @ZephrFish.” The post (251 likes, 15.2K views at capture) lists what the pipeline contains, 8 MCP servers, 300+ tools, patch diffing, fuzzing, crash triage, Ghidra, radare2, Frida, WinDbg/GDB, RAG, Proxmox, exploit dev, CVE and disclosure workflows, and then names what it is actually endorsing:
The interesting part isn’t ‘AI found a bug’. It’s turning years of human vulnerability-research knowledge and tooling into a repeatable system that can assist with 0-day hunting, 1-day analysis, RE and fuzzing without blindly trusting the model.
The two writeups are by Andy Gill, Jenny was a Friend of Mine (MCPs and Friends) from April 2026 and Harnessing Harnesses from June 2026. The first one contains the single most transferable line in either post, and it is not about tools at all:
Now, the single most important design decision in the entire system, and one I wish I’d made on day one: findings start as hallucinations.
Every new finding in that rig is written to hallucinations/, not findings/. It only moves when it passes four gates. The MCP servers are the part everyone copies. The bin is the part that decides whether the output is worth reading.
The four gates, and why Gate 3 exists
The promotion path in the writeup is short enough to quote in full:
- Gate 0: PoC exists and compiles
- Gate 1: PoC reproduces the crash in a clean VM snapshot
- Gate 2: Crash is exploitable (not just a null deref or a graceful exit)
- Gate 3: Bug triggers as a standard user, not SYSTEM or admin
A finding moves from hallucinations/ to findings/ through tool_finding_promote, and there is a tool_finding_demote for the case where a finding is invalidated later. The reason Gate 3 is a first class gate rather than a footnote is a lesson the author describes as nearly fatal to the whole approach: early findings were being validated from a SYSTEM context and looked excellent there. From a standard user, a named pipe on a hardware vendor’s diagnostic agent was unreachable. The rig’s own admission is that Gate 3 should have existed on day one, which is why every Windows VM in the hunt range now carries a lowpriv account.
The bin is not a hedge, it is the measured behaviour of the system. In one early session the engine flagged six findings from manual analysis that looked promising in logs and decompilation. All six were disproven in dynamic analysis. The writeup’s own summary of the economics is that “the hallucination bin is larger than the findings directory by a significant margin, and that’s working as intended.”
What is confirmed, what is alleged, and what is the author’s own accounting
This is where a post like this normally overreaches, so here is the split.
Confirmed independently in this run: CVE-2026-33809 exists in NVD with a CVSS 3.1 base score of 5.3, published 2026-03-25, described as a maliciously crafted TIFF file causing an allocation attempt of up to 4GiB. CVE-2026-33812 exists with a base score of 6.1, published 2026-04-21, described as a malicious font file causing excessive memory allocation. The Go tracker issue golang/go#78267 is titled “x/image/tiff: OOM from malicious IFD offset”, was filed by the user ZephrFish, and is closed. Those are the two CVEs the writeup attributes to its own fuzzing campaign.
Taken on the author’s account, not independently verified here: that 21 Go standard library packages were fuzzed across roughly 80 million executions, that only TIFF and SFNT produced confirmed vulnerabilities, that fixes were merged as golang/image PRs #25 and #26, that a Windows OEM update service chain (client-side-only pipe authentication, SSRF in a WCF method with no admin check, catalog injection, then signed binary execution as SYSTEM) was confirmed valid by the vendor, and that the number of CVEs netted by the process is now four against two from his entire prior career. The cost figures are also the author’s own: roughly GBP 5,000 per CVE if the same usage had been billed as API spend rather than a subscription.
Alleged, meaning it comes from an X post and nothing else: on 2026-09-25 an account called @AndreaTheMiddle posted that it had fine-tuned Laya, a 421M parameter model, to “Deduplicate/Aggregate vulnerability findings from Nuclei, Burp Suite, WPScan, etc…” and that it runs locally at zero inference cost. The account is not verified, the post names no model card, and the demonstration is a 33 second video showing findings coloured green for same and red for different. Treat the technique as entirely plausible and the specific claim as unconfirmed. It also happens to be the most interesting item in the whole radar sweep, and the reason is structural rather than technical.
The failure mode the bin exists to prevent
There are two ways a promotion gate fails, and this site has covered both of them in the last month.
Fail open and the reports flood out. In August, Core Lightning’s maintainers spent ten days triaging a wave of AI-generated CVE reports against one of Bitcoin’s two main Lightning implementations, and several of them were real, which forced an emergency advisory to roughly 16,000 node operators (AI Wrote the Bug Reports. The Fix Was ‘Don’t Turn It Off.’). Discovery was cheap in that episode. Triage was the crisis.
Fail with the wrong question and the pipeline discards a true positive. On July 24, QED Audit filed an integer overflow in V8’s WebAssembly deserializer. Google’s automated triage, which announces itself as “AI-generated using the v8-security-triaging skill”, reproduced the crash, rated its security impact as none, and recommended WontFix. Humans overrode it, and the fix shipped in Chrome 151.0.7922.108 on August 6 (The AI Triage Called It ‘Not a Bug.’ It Was a Chrome RCE.).
A hallucination bin is a response to the first failure. Gate 3, reproduction as the principal an attacker would actually be, is a response to the second. Both are control loop decisions, and neither is a model choice.
Why deduplication is now the layer to watch
Findings that pass a gate still have to be counted. Nuclei, Burp Suite, and WPScan will each report the same injected parameter, the same missing header, and the same reflected value in their own vocabulary, and the deduplication step is the one that decides whether a human reads twelve items or one. It is also almost pure classification: low reasoning depth, high volume, repetitive, and trivially checkable against known ground truth. That is exactly the shape of work that should not be going to a frontier model.
The harness writeup says the same thing from the other direction: “Cheap models can classify, organise and summarise, while stronger models are reserved for validation, tracing and synthesis.” Its published stage contract is recon → hunt → validate → trace → report, where validate is instructed to look for reasons a finding is wrong and trace has to prove that attacker-controlled input reaches the vulnerable sink. Context is budgeted per stage, with single function analysis landing near 8K tokens and multi finding synthesis near 32K, and fuzzer output reduced to a few hundred useful tokens before it reaches a prompt.
The public harnesses that emerged over the same period converge on the same shape. RAPTOR runs six validation stages, ending in binary feasibility checks for ASLR and RELRO plus Z3 constraint solving before anything is promoted. Anthropic’s reference harness runs an autonomous find, grade and patch loop inside AddressSanitizer Docker containers, so every finding carries a binary PoC that reproduces against the instrumented build. evilsocket’s audit pipeline runs eight stages and requires a trace stage to show attacker-controlled input reaching a sink before reporting. Visa’s VVAH validates findings through an adversarial second pass and then, in the writeup’s own framing, treats its output as triage candidates rather than confirmed vulnerabilities. The design question in every one of them is where the promotion gate sits and who is allowed to be wrong.
What an operator should change
- Quarantine on entry, by name. A directory called
hallucinations/orunverified/is a stronger control than any instruction in a system prompt, because it makes the default state of a model output “not yet true” in the filesystem rather than in prose. - Gate on reproduction, not on plausibility. A decompiler reading and a log line are hypotheses. The gate that has teeth is a PoC that reproduces in a clean snapshot.
- Reproduce as the right principal. Gate 3 is the one that catches the largest class of fake severity, because SYSTEM-only reachability is not a vulnerability in the way most programmes score it.
- Keep the demote path. A one way promotion counter measures nothing. The ability to move a finding back is what makes the promotion rate a real metric.
- Record negatives. Defences that blocked a hunt, AM-PPL protected processes, signature validation, admin-only pipes, are the input that stops the pipeline re-walking the same dead end next month.
- Do not let the generating model grade its own finding. The patch grading dispute this month made the point in the benchmark dimension (The Patch Grader Was the Problem). The same failure appears here the moment the agent that proposed a finding is also the agent that signs it off.
- Route the boring stages down. If a 421M encoder can decide whether two Nuclei findings are the same finding, that decision belongs on local hardware, and the frontier budget belongs to validate and trace.
The model is the cheapest part
The uncomfortable truth in both writeups is that almost nothing described is about the model. Eight MCP servers, 300+ tools, five VMs, a Patch Tuesday pair for binary diffing, a known-defence database, a bounty ranking that deliberately uses the minimum published payout, and a gate list that starts by disbelieving the system’s own output: that is a control loop with a language model somewhere inside it, not a model with some tools bolted on.
The reason the bin matters more than the servers is that discovery is now the cheapest artifact in the pipeline. Generating candidate findings is something a subscription and a weekend can do. Promoting one is the part that costs judgement, and the honest name for a finding that has not been promoted is what that second writeup puts in its directory listing. Start with the hallucination bin, because everything downstream of it is a measurement, and everything upstream of it is a guess.
Sources:
- Jenny was a Friend of Mine, MCPs and Friends (alt title: Bullying LLMs into submission to find 0days at scale), Andy Gill, 4 Apr 2026 (primary, first-person account of the author’s own system: 8 MCP servers across 5 VMs and 300+ tools, the
hallucinations/bin, Gates 0 through 3,tool_finding_promoteandtool_finding_demote, the six disproven findings, the lowpriv lesson, the Go fuzzing campaign, cost and CVE accounting, the known-defence database) - Harnessing Harnesses, Climbing the LLM Hills, Andy Gill, 27 Jun 2026 (primary: harness definition and stage contract, model routing between cheap and strong models, per-stage context budgets, the published harnesses and their validation stage designs)
- @cr3ghost on X, thread recommending both writeups, 10 Sep 2026 (primary: the quote about AI finding a bug not being the interesting part; 251 likes, 55 reposts, 15.2K views at capture on 26 Sep 2026)
- @AndreaTheMiddle on X, fine-tuned Laya for vulnerability finding deduplication, 25 Sep 2026 (unverified claim from an unverified account: 421M parameter fine-tune described as deduplicating and aggregating findings from Nuclei, Burp Suite and WPScan, running locally at zero inference cost; 168 likes, 18 reposts, 12.3K views at capture)
- CVE-2026-33809, NVD (primary: CVSS 3.1 base 5.3, published 25 Mar 2026, TIFF decode allocation of up to 4GiB)
- CVE-2026-33812, NVD (primary: CVSS 3.1 base 6.1, published 21 Apr 2026, malicious font file causing excessive memory allocation)
- golang/go issue #78267, x/image/tiff: OOM from malicious IFD offset (primary: issue title, reporter ZephrFish, closed status)
- gadievron/raptor, anthropics/defending-code-reference-harness, evilsocket/audit, visa/visa-vulnerability-agentic-harness, ZephrFish/harness-kit, alpha-omega-security/scrutineer (reference: public harness implementations, their validation stages, and the
revalidateskill returning true_positive, false_positive, already_fixed or uncertain) - AI Wrote the Bug Reports. The Fix Was ‘Don’t Turn It Off.’, Denny Sentinel, 28 Aug 2026 and The AI Triage Called It ‘Not a Bug.’ It Was a Chrome RCE., Denny Sentinel, 6 Sep 2026 (prior coverage used here as the two reference failure modes for a promotion gate)