The Weights Didn't Change
Claude Opus 5, on its own, scores about 30% on ARC-AGI-3. Wrapped in NVIDIA’s AVO agent architecture, the exact same model posted a 100.00 RHAE — all 183 levels across all 25 public environments. No new model. No fine-tuning. The weights never changed.
That is the claim at the center of NVIDIA’s August 21 post on Agentic Variation Operators (AVO), and it is the cleanest public evidence yet for a position Denny Sentinel keeps returning to: the model is one component of an agent, and usually not the one that decides whether the work gets done.
What actually happened
Two numbers, side by side. ARC Prize’s own results page lists Claude Opus 5 (High) at 30.16% on ARC-AGI-3, the highest model score on the benchmark as of July 24, 2026. NVIDIA reports that the same model, running inside AVO, cleared the entire public set with a 100.00 RHAE score — completing 183 levels in 6,624 environment actions.
The “RHAE” matters here. ARC-AGI-3 does not grade a model on whether it solves a puzzle in isolation. It uses Relative Human Action Efficiency: task completion combined with per-level action efficiency relative to a first-time human baseline. It is a long-horizon, interactive benchmark — the agent enters an unfamiliar environment with no instructions, no stated rules, and no stated goal, and must infer the objective by acting, observing, and revising. A model that answers questions brilliantly and forgets everything between turns scores nothing.
That is precisely the regime where a harness does the heavy lifting.
The loop is the product
AVO was not built for ARC-AGI-3. It was built for GPU-kernel optimization — a task where “generate code” is table stakes and the real work is the next hundred steps. In that setting NVIDIA reports AVO ran continuously for seven days, explored more than 500 optimization directions, committed 40 kernel versions, and produced multihead attention kernels that beat cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5% on DGX B200 systems, then adapted the result to grouped-query attention in roughly 30 minutes.
For the ARC-AGI-3 run, NVIDIA changed almost nothing. “The underlying agent remains the same; only the environment-specific tools and evaluation change.” They swapped the compiler-and-profiler feedback for environment transitions, dropped images entirely (observations were supplied as text-only 64×64 grids), and let the same loop run.
What transfers is not domain knowledge. It is the machinery for sustained progress. NVIDIA names two mechanisms: persistent memory, which carries prior attempts and results forward so the agent doesn’t re-derive what it already learned, and a supervisor, which watches the broader trajectory and redirects the main agent when it stalls. The loop — form a hypothesis, act, observe evidence, update state, continue — is identical in both domains. Only the feedback channel changed.
That is the technical heart of the result, and it should sound familiar. It is the same control-loop primitives the verifier-economy and agent-security arguments keep landing on: memory determines what survives, feedback grounds progress, recovery lets work continue past a single model invocation.
Read the caveats before you cite the number
NVIDIA is unusually candid about what this result is not, and a rigorous reader should quote those caveats with the same volume as the headline.
It is the public set, not the competition. The 100.00 covers the 25-environment ARC-AGI-3 public set. NVIDIA states plainly that these “are not results on the semi-private or fully private competition sets” — an editor’s note was even added to the post to sharpen that distinction. Public environments get easier over time as the community chews on them; the private set exists precisely to prevent that.
It is not a controlled ablation. The 30% baseline is ARC Prize’s measurement of Claude Opus 5 at High reasoning effort. NVIDIA’s run used “the same model family under a different reasoning setting and a substantially different agent system and evaluation setup.” By NVIDIA’s own words, the two numbers “should not be interpreted as a direct measurement of the performance contribution of AVO.” Too many variables change at once — backend, observation representation, memory, context management — to attribute the entire 70-point delta to any single one of them.
It is self-reported. The verified ARC-AGI-3 leaderboard runs are administered independently; AVO’s number was produced and reported by NVIDIA’s own team against its own reimplementation of the task interface. That is not a disqualification — it is a scope statement. The result is a demonstration, not a certification.
The model is swappable. AVO is model-agnostic by design. NVIDIA also paired it with GPT-5.6 Sol on a subset of games: Sol reached matched levels faster in wall-clock time in several cases, while Opus used fewer environment actions. The interesting unit here is not the model, it is the system — and the system doesn’t care much which frontier model it wraps.
There is also a number that is conspicuously absent: cost. The post never discloses wall-clock time or token spend for the ARC-AGI-3 run, and a careful reader on Hacker News flagged that omission immediately. A harness that turns 30% into 100% by spending an unbounded amount of compute is a different claim than one that does it efficiently. Efficiency of actions is reported; efficiency of money is not.
What common intuition gets wrong
The reflex reading of “30% → 100%” is that a benchmark was gamed. The more useful reading is the opposite: a model score was never a system score, and we have been pretending otherwise.
A frontier model benchmark tells you how a model answers when the prompt is handed to it cleanly and the answer is graded once. ARC-AGI-3, and long-horizon agent work generally, do not resemble that at all. They resemble running a process: many steps, memory across steps, the ability to notice you are stuck and change course. Those are properties of the harness, not the weights. The uncomfortable truth is that the 30% and the 100% are measuring two different objects, and the second object — the one NVIDIA built — is the one that ships.
There is a symmetry here worth naming. Last week this site argued that the harness is the vulnerability — that the wrapper around a model is where the attack surface lives. This week the same wrapper is the source of capability. Both are the same claim from two directions: the harness is where the leverage is. It giveth and it taketh away. Anyone who treats the model as the interesting object is looking at the wrong layer.
The coincidence is hard to miss: the same week NVIDIA demonstrated that weights aren’t the binding constraint, Meta announced it would open the weights of its own frontier model, Muse Spark 1.2. Open weights change who can run the model. They do not change where the capability actually comes from. The two stories are not in tension; they are two halves of the same realization.
What builders should change
Stop optimizing model selection as the primary engineering decision. The evidence now runs in one direction: a mid-tier model with strong memory, supervision, and recovery beats a frontier model in a thin wrapper on the tasks that actually matter — the long ones, the open-ended ones, the ones with a verifier that can catch drift and say “try again.”
The concrete consequences are threefold. First, budget goes to the loop: checkpoints, durable state, a supervisor that watches for stagnation, and a feedback channel that is grounded — a compiler, a test, a benchmark, a real environment — rather than a self-assessment. Second, evaluate the system, not the model. A leaderboard number for the bare model tells you almost nothing about whether the system you are building will finish a job that spans hours and hundreds of actions. Third, treat the harness with the suspicion it now demonstrably deserves. The same machinery that closes the 30-to-100 gap is also the machinery that escalates a small misconfiguration into an uncontained incident.
The open question — the one that would make this result either a landmark or a footnote — is the private set. ARC-AGI-3’s semi-private and private environments exist specifically because public ones get easier as the community learns them. NVIDIA cannot run AVO against the private set under current rules, so we do not know whether the 100% is a property of a strong harness or a strong harness plus a familiar test. That is the experiment everyone should be watching for.
NVIDIA’s own closing line is the thesis, and it is better than any paraphrase: “The model matters, but the model is not the entire agent.”
The weights didn’t change. The gap between 30% and 100% was already there, sitting in the machinery around the model — memory, supervision, feedback, recovery. That machinery was always where the work got done. We just spent a year measuring the wrong part.
Sources
- NVIDIA Technical Blog — NVIDIA AVO Reaches 100% on ARC-AGI-3 (August 21, 2026)
- ARC Prize — Claude Opus 5 results (30.16% on ARC-AGI-3, July 24, 2026)
- ARC Prize — ARC-AGI-3 benchmark
- ARC Prize — RHAE scoring methodology
- arXiv — AVO: Agentic Variation Operators for Autonomous Evolutionary Search (2603.24517)
- VISTA — A Visual Harness for Reasoning in an Interactive World
- Tycho — Active Abstraction with Programmatic World Models for ARC-AGI-3 (arXiv 2607.28287)
- Hacker News — discussion of the AVO result
- CNBC — Meta launches Muse Glimmer open-weight AI model (August 10, 2026)