The Good News Is in the Verifier: What AI Actually Delivers by 2036
The most consequential AI result of the last twelve months is not a benchmark score. It landed on 30 January 2026 in The Lancet: the full results of the MASAI trial, the first randomised controlled trial of AI-supported mammography inside a national screening programme.
Over 100,000 Swedish women, four sites, one specialist AI system triaging low-risk mammograms to single reading and high-risk cases to double reading. The AI arm produced 12% fewer interval cancers (1.55 vs 1.76 per 1,000 women), 16% fewer invasive cancers, 21% fewer large tumours and 27% fewer aggressive subtypes. False positives were flat — 1.5% versus 1.4%. The interim analysis had already found a 44% reduction in screen-reading workload for radiologists.
That is not a demo. It is a randomised result inside a public health programme, and it is the shape of nearly all the AI good news worth trusting this decade: a model inserted into a loop that already had a way to check its work.
Both the doom discourse and the hype discourse miss this, because both are arguments about the model. Models are the cheap part. The verifier — the trial, the second reader, the hardware interlock, the synthesis step, the power plant — is where reliability is manufactured and where the gains are bounded.
Three verifier classes, and what each has actually delivered
Sort the credible claims by the kind of arbiter checking them, and the pattern stops being mysterious.
The detector inside an existing measurement loop
Here the AI’s output is compared against something already authoritative: a diagnosis, a confirmed case, a definite outcome. The verifier is unambiguous and cheap relative to the decision.
Mammography is the clearest case, and the UK’s EDITH trial extends it — around 700,000 women across 30 testing sites, £11 million of NIHR funding, launched February 2025, built on the fact that AI-assisted reading lets one specialist do what currently takes two. The MASAI authors are careful about the boundary: AI still requires at least one human radiologist, and their finding “does not support replacing healthcare professionals with AI.” The win is throughput and earlier detection, not autonomy.
Rare disease shows the same structure with a stranger arbiter. A study published in JAMA Network Open (VUMC, August 2025) gave two LLMs the standardised intake summaries of 90 Undiagnosed Diseases Network patients whose median diagnostic odyssey had already run 7.6 years. The models named the eventual final diagnosis 13.3% (ChatGPT-4o) and 10.0% (Llama 3.1 8B) of the time, against a 5.6% historical clinical review rate — at $0.03 and five seconds per case. The reason that number means anything is that the patients already had ground truth. The journal’s reviewers could falsify it.
The control loop, where the safety limit is deliberately outside the model
The fusion result from September 2026 is the most instructive piece of engineering in this whole inventory. PPPL and Princeton’s PACMAN framework (published in Nuclear Fusion) runs a repeating control loop at roughly 20 milliseconds, against a human operator’s floor of seconds. In one of five experiments on the DIII-D tokamak it predicted a tearing-mode instability about 200 milliseconds before it formed and changed the plasma to avoid it — something a conventional controller cannot do, because conventional controllers only see an instability once it has started.
The parts that matter for anyone building agent systems are not the latency. They are:
- Hardware safety limits are enforced regardless of what the model recommends. The model proposes; the interlock disposes. The safety property is not a property of the model.
- Humans set the objectives. The loop optimises; the physicist chooses the target and reviews after every shot.
- Modularity turned demonstrations into infrastructure. The first model took months to install. The second took a couple of days. As Egemen Kolemen put it, “that modularity is what turns AI plasma control from a series of one-off demonstrations into infrastructure the whole fusion community can build on.”
That is the loop-engineering argument stated by plasma physicists: the verifier and control structure determine reliability, not the model.
Search-and-confirm, where generation finally stopped being the bottleneck
Generative proposals with a hard physical arbiter have moved furthest, fastest.
Insilico Medicine went from initiating a target-discovery programme to dosing humans in an anti-fibrotic programme in under 30 months — a pathway that has historically consumed 10 to 15 years. The preclinical half took around 18 months at roughly $2.6 million. AlphaFold2 predicted structures for essentially all 200 million known proteins and is used by more than two million people in 190 countries; the 2024 Chemistry Nobel recognised both protein structure prediction and David Baker’s computational protein design, which is the half people forget.
Then there is the antibiotic result, which is the most honest story in the field. MIT’s team used generative AI to design compounds atom-by-atom against drug-resistant gonorrhoea and MRSA, interrogating 36 million compounds, and two candidates killed the bacteria in laboratory and animal tests — a result serious enough that reviewers called it a potential “second golden age” of antibiotic discovery. And the same paper contains the sobering detail: of the top 80 gonorrhoea designs the AI produced in theory, two were synthesised.
Two out of eighty. That ratio is the thesis of this post in a single number. Candidate generation is no longer the constraint. Synthesis, safety testing and clinical trials are.
What intuition gets wrong: the model was never the product
The standard framing treats the model as the deliverable and everything else as plumbing. The evidence says the reverse: the arbiter is the product, and the wins scale exactly as far as the arbiter can be cheaply applied.
Which explains the failures as cleanly as the successes. Where there is no bounded, repeatable way to check the output, nothing moves this decade — and the sources say so in their own words.
A 150-year lifespan is a ceiling, not a roadmap. The number comes from a 2021 Nature Communications study (Pyrkov et al.) that measured loss of physiological resilience across three cohorts using blood cell counts and daily step counts, and concluded that even with disease and accidents removed, resilience decays to a hard limit of 120 to 150 years. Jeanne Calment’s verified record is 122. Co-author Peter Fedichev’s own framing of the next step is precisely the verifier problem: “Measuring something is the first step before producing an intervention.” There is no cheap arbiter for an intervention that takes eighty years to read out. Expect healthspan gains and better measurement; do not expect the lid to move.
Humanoid robots arrive in factories, not living rooms. IDTechEx’s March 2026 forecast puts the market at roughly $29.5 billion by 2036, scaling first in automotive manufacturing and then logistics — because those are controlled environments with structured workflows, repeatable uptime metrics and available real-world data. Home deployment is expected only after 2030, with volume beyond 2035, held back by “long-tail scenario coverage” and the lack of datasets to validate embodied behaviour. Houses have no verification harness. Factories do.
Fusion reaches the grid in the 2030s, not this decade. PACMAN is genuinely remarkable, and it is a control result. The verifier for a power plant is a power plant, and each trial costs a fortune and takes years. The WEF’s own framing of AI’s fusion contribution is grid-scale in the 2030s.
Road deaths will not halve by 2030. The EU pledged a 50% cut against a pre-2020 baseline; at the halfway mark, deaths are down about 15%, and mandatory ADAS is modelled to remove roughly 6% of injury and fatal accidents by 2030. Verification here is expensive for a structural reason: crashes are rare relative to miles driven, so you need enormous exposure before you learn anything. Rare-event verification is the hardest kind, which is why autonomous trucks remain under 1% of the US commercial fleet on Goldman Sachs’ 2030 projection.
What operators should change
Four transfers from this inventory into systems work:
Instrument the verifier before you scale the generator. It is your throughput ceiling, and it is usually invisible in your metrics. The 80-to-2 ratio is what it looks like when a pipeline reports the size of its candidate pool instead of the yield of its confirmation step. Measure the confirmation step.
Put the safety limit outside the model. PACMAN enforces hardware limits regardless of what the model asks for. Translate that: approval gates, capability scopes and interlocks that the model cannot argue with. A model that can talk its way past the check is not checked.
Prefer loops to one-shot judgements. The MASAI trial’s benefit came from a triage loop, not a better radiologist replacement — and the finding that the AI still requires at least one human reader is a description of the architecture, not a limitation to be optimised away.
Expect the economics to be smaller and better documented than the discourse. The OECD’s 10-year estimate for AI’s contribution to annual aggregate labour productivity growth in high-exposure G7 economies is 0.4–1.3 percentage points. Wharton’s estimate is a 1.5% higher GDP level by 2035, peaking at 0.2 percentage points of annual growth in 2032. Those are unglamorous numbers — and unlike the headier ones, they come with a stated mechanism and a stated uncertainty range.
The closing thesis
The good news about AI is not that the machines got smart. It is that in a handful of domains — a screening programme, a tokamak, a protein database, an antibacterial assay — someone built a place to check the machine’s work, and the checks are passing.
That is why the wins cluster where they do. Detection, control and search have arbiters that are cheap, fast and unambiguous. Longevity, domestic robotics and grid-scale fusion do not, and they will keep slipping for reasons that have nothing to do with model capability.
The uncomfortable truth is that this reframes the decade’s opportunity. The scarce resource is not intelligence. It is the number of domains where we can afford to find out that we are wrong.
Sources
- The Lancet, AI-supported mammography screening: full results of the MASAI randomised controlled trial (2026): https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(25)02464-X/abstract
- ecancer, summary of the MASAI full results (30 January 2026): https://ecancer.org/en/news/27721-ai-supported-mammography-screening-results-in-fewer-aggressive-and-advanced-breast-cancers-finds-full-results-from-first-randomised-controlled-trial
- GOV.UK / DHSC, World-leading AI trial to tackle breast cancer launched (EDITH, February 2025): https://www.gov.uk/government/news/world-leading-ai-trial-to-tackle-breast-cancer-launched
- Shyr et al., LLM diagnostic performance in the Undiagnosed Diseases Network, JAMA Network Open (August 2025), via VUMC: https://news.vumc.org/2025/09/05/ai-tools-could-shorten-diagnostic-odyssey-for-patients-with-rare-diseases/
- Princeton / PPPL, PACMAN framework tested on DIII-D, Nuclear Fusion (September 2026): https://www.sciencedaily.com/releases/2026/09/260903064215.htm
- Insilico Medicine, From Start to Phase 1 in 30 Months (ISM001-055): https://insilico.com/phase1
- BBC News, AI designs new superbug-killing antibiotics for gonorrhoea and MRSA, MIT study in Cell (August 2025): https://www.bbc.co.uk/news/articles/cgr94xxye2lo
- Nobel Prize in Chemistry 2024 press release (Baker; Hassabis and Jumper): https://www.nobelprize.org/prizes/chemistry/2024/press-release/
- Kestin et al., AI tutoring outperforms in-class active learning: an RCT, Scientific Reports (June 2025): https://www.nature.com/articles/s41598-025-97652-6
- Pyrkov et al., longitudinal analysis of blood markers and step counts sets a limit to human lifespan, Nature Communications (2021), via Scientific American: https://www.scientificamerican.com/article/humans-could-live-up-to-150-years-new-research-suggests/
- IDTechEx, Humanoid Robots to Reach Nearly US$30 Billion by 2036 (March 2026): https://www.idtechex.com/en/research-article/humanoid-robots-to-reach-nearly-us-30-billion-by-2036/34443
- OECD, Macroeconomic productivity gains from Artificial Intelligence in G7 economies (June 2025): https://www.oecd.org/en/publications/macroeconomic-productivity-gains-from-artificial-intelligence-in-g7-economies_a5319ab5-en.html
- Penn Wharton Budget Model, The Projected Impact of Generative AI on Future Productivity Growth (September 2025): https://budgetmodel.wharton.upenn.edu/p/2025-09-08-the-projected-impact-of-generative-ai-on-future-productivity-growth/
- ETSC, A two-speed Europe on road safety (2025): https://etsc.eu/a-two-speed-europe-on-road-safety-a-handful-of-countries-on-course-to-halve-deaths-most-falling-behind/
- CEPI, artificial intelligence and the 100 Days Mission: https://cepi.net/artificial-intelligence
- Google Research, Flood Hub (river basins in over 80 countries as of June 2025): http://sites.research.google/gr/floodforecasting/