What's the most essential difference between ARC-AGI-3 and the earlier ARC-AGI-1 and ARC-AGI-2?
The most essential difference is when information gets revealed. Both ARC-AGI-1 and ARC-AGI-2 use a static format: the model can see the complete set of input-output examples right from the start, the rule is given upfront, and the model's job is to generalize that rule from the examples and apply it to a new problem. ARC-AGI-3 moves the point of rule disclosure to after the action — the AI has to take an action in the environment first, and only then sees the result that action produced. The rule is revealed piece by piece through action itself, not handed over as a pre-packaged set for the model to study.
This difference makes it nearly impossible for ARC-AGI-3 to be gamed through brute-force sampling, because before taking any action, there's no list of candidate answers to enumerate in the first place.
Why did StochasticGoose score 12.58% on the preview set but drop to only 0.25% on the full benchmark, and what does that gap itself reveal?
This gap reveals a long-standing problem that becomes especially visible on ARC-AGI-3: optimizing for a specific question set and possessing general-purpose ability to solve that entire category of problem are two different things. If part of StochasticGoose's 12.58% on the preview set came from learning the specific patterns of a handful of environments in that preview set — rather than learning the general Skill of "how to explore an unfamiliar environment" — then once it faced entirely new, unseen environments in the full benchmark, that targeted learning simply didn't transfer, and the score dropped sharply as a result.
This is also why benchmark teams typically hold back a portion of test questions that have never been made public, specifically to check whether a submitted approach genuinely has general capability or has merely memorized the specific patterns of the public question set.
What specific capability does ARC-AGI-3's "exploration efficiency" refer to, and how does it map onto a human playing a new video game?
"Exploration efficiency" refers to how many attempts and mistakes an agent needs, with zero instruction manual and no tutorial level, before it figures out what the environment's rules are and what counts as winning — and then achieves that goal efficiently. It's not about randomly trying things out of luck; it's about extracting useful information from the outcome of each attempt and using it to correct the next move.
This maps closely onto a human playing a video game they've never touched before. Most people don't read the entire manual before starting — they just pick up the controller, press buttons, watch how the character reacts, fail, try again, and within a few minutes have built a rough mental model of "roughly how this game works." ARC-AGI-3 is testing exactly whether AI has this capability, and current results show the gap to human performance on this specific Skill is far larger than the gap AI shows on tasks where large amounts of training data are available to lean on.
As a reader, when I see a headline like "new benchmark instantly collapses AI scores," how do I tell whether that means AI got worse, or the test just changed?
First check two things. One: did the score collapse because the model itself got worse (impossible — the model didn't change), or because a harder question set that's better at exposing weaknesses replaced the old one? Two: does the new question set measure the same capability as the old one? ARC-AGI-3's score collapse falls into the second category — it was deliberately designed to measure a capability the old benchmark couldn't test (actively exploring an unfamiliar environment), so a high score on the old benchmark and a low score on the new one aren't actually contradictory; they were never measuring the same thing to begin with.
For readers, spending a few minutes checking whether the old and new question sets measure the same capability gets you closer to the truth than accepting either "AI got better" or "AI got worse" as a single, simple conclusion.
Over the past two years, AI model scores on standardized tests climbed so fast they looked like they were approaching a ceiling — exactly the situation ARC-AGI-1 fell into in 2025, when one model surged from 0% to 92.5%, only to crash back to 0.37% the following year once a harder question set replaced the old one. The entire benchmark system seemed stuck in a loop where the test-makers couldn't outrun the test-takers. On March 25, 2026, the ARC Prize Foundation — co-founded by Keras creator François Chollet and Zapier co-founder Mike Knoop — released ARC-AGI-3, which abandons the old question format entirely. Instead of static puzzles where models guess a pattern from a picture, it drops AI directly into real-time, interactive game environments.
The result: human players cleared every single test environment, hitting 100%. Multiple frontier models, including Grok-4.20, scored under 1%.
Both ARC-AGI-1 and ARC-AGI-2 use a static format: the model is shown several "input pattern to output pattern" examples, then infers a rule from those examples and produces an answer. This format has a hidden weakness — a model can rely on heavy sampling, brute-force enumerating many possible answer combinations, and picking whichever one looks most plausible. A model can stumble into the right answer this way without ever truly "understanding" the underlying rule. ARC-AGI-3's design closes off that path entirely: the test environment is live and interactive, and the rules are only revealed after the AI takes action. A model can't preview a complete set of input-output examples and work the problem out at leisure the way it could with a static puzzle — it has to probe, make mistakes, and adjust its next move based on the outcome of those mistakes. The whole process resembles "learning how to play a new video game" far more than "solving a math problem."
Published results show most mainstream AI systems scoring under 1%. Grok-4.20 exceeded the action-count limit on every level it attempted, ending with a final score of 0%. StochasticGoose, an approach combining reinforcement learning with a convolutional neural network, scored only 0.25% on the full benchmark — a steep drop from the 12.58% it posted on the preview version, suggesting that earlier score may have come partly from overfitting to the preview question set, with the gap exposed once it faced the full set. Human testers, by contrast, achieved a 100% clear rate across every environment. That gap — not a gradual 60% versus 90% difference, but a near-cliff of roughly 0% versus 100% — is itself the most notable thing about ARC-AGI-3.
Many prior benchmarks effectively measured whether a model had seen similar questions before and how many patterns it had memorized — a capability that improves steadily by stacking more training data and compute, which is exactly why AI scores climbed so fast over the past few years. ARC-AGI-3 measures something different: "Skill-acquisition efficiency" — whether, with zero training data and no prior instructions, a model can actively explore an unfamiliar environment, figure out for itself what "winning" looks like, and then achieve it efficiently. This capability can't be patched by simply feeding in more data, because the benchmark is specifically designed so that each environment's rules differ — a model gains no advantage from "having seen a similar puzzle before."
This also explains why an approach like StochasticGoose, optimized against the preview question set, performed worse on the full set. If a system scores well by memorizing the specific patterns of environments in the preview set, rather than possessing genuine, general exploration ability, that kind of shortcut optimization collapses the moment it faces entirely new, unseen environments.
Next time you see a headline about some model achieving a breakthrough score on a benchmark, it's worth first checking what kind of capability that benchmark actually measures — is it testing how much the model remembers and has seen, or whether it can handle a genuinely novel situation it has never encountered? ARC-AGI-3's sub-1% score doesn't mean AI capability overall has stalled. It reflects a more precise fact: current AI systems improve rapidly on tasks where large amounts of training data are available to lean on, but on tasks that require exploring from zero with no example to imitate, the gap to human performance hasn't narrowed. Which of these two capabilities matters more for the application you actually care about determines which benchmark's score you should be using to judge where AI genuinely stands today.