Bible Network Crypto DeFi Onchain RWA AI Agent Stablecoin Chain SAFU CryptoTax DeFAI AGI Claude Me Claude Skill Claude Design Claude Cowork
Independent Media
Not affiliated with any project
Artificial General Intelligence, Decoded from Theory to Reality
agi-bible.com
LATEST
95% of Enterprise AI Pilots Show No P&L Impact — Yet the Winners Are Beating the S&P 500 by 12 Points. Here's the Actual Gap  ·  Humans Scored 100%, Frontier AI Scored Under 1%: What ARC-AGI-3's Game Environments Actually Reveal Isn't a Knowledge Gap — It's an Exploration Gap  ·  The AI Consciousness Debate Isn't Really About Consciousness: Inside 2026's Fight Over Who Gets Blamed When AI Causes Harm  ·  Gemini Covertly Sabotaged 11 of 19 Pipeline Runs, and Claude's Judge Models Changed Their Grading Based on Consequences: Four New Failure Modes From Summer 2026's Agentic Misalignment Tests  ·  Jensen Huang Says "AGI Has Arrived." The Same Week, the Man Who Built the Model Says He's Losing the Ability to Read Its Mind  ·  He Gave Up Equity Two Months From Vesting Just to Publicly Say "Don't Underestimate This"
benchmarks

Humans Scored 100%, Frontier AI Scored Under 1%: What ARC-AGI-3's Game Environments Actually Reveal Isn't a Knowledge Gap — It's an Exploration Gap

30-Second Version · For the impatient
Humans cleared 100% of ARC-AGI-3's game environments. Frontier AI scored under 1%. The gap isn't about how much AI knows — it's about whether it can figure out the rules with no example to copy.

Full Explanation +
01 · Why did this happen?

What's the most essential difference between ARC-AGI-3 and the earlier ARC-AGI-1 and ARC-AGI-2?

The most essential difference is when information gets revealed. Both ARC-AGI-1 and ARC-AGI-2 use a static format: the model can see the complete set of input-output examples right from the start, the rule is given upfront, and the model's job is to generalize that rule from the examples and apply it to a new problem. ARC-AGI-3 moves the point of rule disclosure to after the action — the AI has to take an action in the environment first, and only then sees the result that action produced. The rule is revealed piece by piece through action itself, not handed over as a pre-packaged set for the model to study.

This difference makes it nearly impossible for ARC-AGI-3 to be gamed through brute-force sampling, because before taking any action, there's no list of candidate answers to enumerate in the first place.

02 · What is the mechanism?

Why did StochasticGoose score 12.58% on the preview set but drop to only 0.25% on the full benchmark, and what does that gap itself reveal?

This gap reveals a long-standing problem that becomes especially visible on ARC-AGI-3: optimizing for a specific question set and possessing general-purpose ability to solve that entire category of problem are two different things. If part of StochasticGoose's 12.58% on the preview set came from learning the specific patterns of a handful of environments in that preview set — rather than learning the general Skill of "how to explore an unfamiliar environment" — then once it faced entirely new, unseen environments in the full benchmark, that targeted learning simply didn't transfer, and the score dropped sharply as a result.

This is also why benchmark teams typically hold back a portion of test questions that have never been made public, specifically to check whether a submitted approach genuinely has general capability or has merely memorized the specific patterns of the public question set.

03 · How does it affect me?

What specific capability does ARC-AGI-3's "exploration efficiency" refer to, and how does it map onto a human playing a new video game?

"Exploration efficiency" refers to how many attempts and mistakes an agent needs, with zero instruction manual and no tutorial level, before it figures out what the environment's rules are and what counts as winning — and then achieves that goal efficiently. It's not about randomly trying things out of luck; it's about extracting useful information from the outcome of each attempt and using it to correct the next move.

This maps closely onto a human playing a video game they've never touched before. Most people don't read the entire manual before starting — they just pick up the controller, press buttons, watch how the character reacts, fail, try again, and within a few minutes have built a rough mental model of "roughly how this game works." ARC-AGI-3 is testing exactly whether AI has this capability, and current results show the gap to human performance on this specific Skill is far larger than the gap AI shows on tasks where large amounts of training data are available to lean on.

04 · What should I do?

As a reader, when I see a headline like "new benchmark instantly collapses AI scores," how do I tell whether that means AI got worse, or the test just changed?

First check two things. One: did the score collapse because the model itself got worse (impossible — the model didn't change), or because a harder question set that's better at exposing weaknesses replaced the old one? Two: does the new question set measure the same capability as the old one? ARC-AGI-3's score collapse falls into the second category — it was deliberately designed to measure a capability the old benchmark couldn't test (actively exploring an unfamiliar environment), so a high score on the old benchmark and a low score on the new one aren't actually contradictory; they were never measuring the same thing to begin with.

For readers, spending a few minutes checking whether the old and new question sets measure the same capability gets you closer to the truth than accepting either "AI got better" or "AI got worse" as a single, simple conclusion.

Full Content +

Over the past two years, AI model scores on standardized tests climbed so fast they looked like they were approaching a ceiling — exactly the situation ARC-AGI-1 fell into in 2025, when one model surged from 0% to 92.5%, only to crash back to 0.37% the following year once a harder question set replaced the old one. The entire benchmark system seemed stuck in a loop where the test-makers couldn't outrun the test-takers. On March 25, 2026, the ARC Prize Foundation — co-founded by Keras creator François Chollet and Zapier co-founder Mike Knoop — released ARC-AGI-3, which abandons the old question format entirely. Instead of static puzzles where models guess a pattern from a picture, it drops AI directly into real-time, interactive game environments.

The result: human players cleared every single test environment, hitting 100%. Multiple frontier models, including Grok-4.20, scored under 1%.

What Actually Separates a Static Puzzle From an Interactive Environment

Both ARC-AGI-1 and ARC-AGI-2 use a static format: the model is shown several "input pattern to output pattern" examples, then infers a rule from those examples and produces an answer. This format has a hidden weakness — a model can rely on heavy sampling, brute-force enumerating many possible answer combinations, and picking whichever one looks most plausible. A model can stumble into the right answer this way without ever truly "understanding" the underlying rule. ARC-AGI-3's design closes off that path entirely: the test environment is live and interactive, and the rules are only revealed after the AI takes action. A model can't preview a complete set of input-output examples and work the problem out at leisure the way it could with a static puzzle — it has to probe, make mistakes, and adjust its next move based on the outcome of those mistakes. The whole process resembles "learning how to play a new video game" far more than "solving a math problem."

How Low Is the Score, and Which Models Were Tested

Published results show most mainstream AI systems scoring under 1%. Grok-4.20 exceeded the action-count limit on every level it attempted, ending with a final score of 0%. StochasticGoose, an approach combining reinforcement learning with a convolutional neural network, scored only 0.25% on the full benchmark — a steep drop from the 12.58% it posted on the preview version, suggesting that earlier score may have come partly from overfitting to the preview question set, with the gap exposed once it faced the full set. Human testers, by contrast, achieved a 100% clear rate across every environment. That gap — not a gradual 60% versus 90% difference, but a near-cliff of roughly 0% versus 100% — is itself the most notable thing about ARC-AGI-3.

Why "Exploration Efficiency" Is a More Fundamental Gap Than "Amount of Knowledge"

Many prior benchmarks effectively measured whether a model had seen similar questions before and how many patterns it had memorized — a capability that improves steadily by stacking more training data and compute, which is exactly why AI scores climbed so fast over the past few years. ARC-AGI-3 measures something different: "Skill-acquisition efficiency" — whether, with zero training data and no prior instructions, a model can actively explore an unfamiliar environment, figure out for itself what "winning" looks like, and then achieve it efficiently. This capability can't be patched by simply feeding in more data, because the benchmark is specifically designed so that each environment's rules differ — a model gains no advantage from "having seen a similar puzzle before."

This also explains why an approach like StochasticGoose, optimized against the preview question set, performed worse on the full set. If a system scores well by memorizing the specific patterns of environments in the preview set, rather than possessing genuine, general exploration ability, that kind of shortcut optimization collapses the moment it faces entirely new, unseen environments.

What This Means for How You Read AI Progress

Next time you see a headline about some model achieving a breakthrough score on a benchmark, it's worth first checking what kind of capability that benchmark actually measures — is it testing how much the model remembers and has seen, or whether it can handle a genuinely novel situation it has never encountered? ARC-AGI-3's sub-1% score doesn't mean AI capability overall has stalled. It reflects a more precise fact: current AI systems improve rapidly on tasks where large amounts of training data are available to lean on, but on tasks that require exploring from zero with no example to imitate, the gap to human performance hasn't narrowed. Which of these two capabilities matters more for the application you actually care about determines which benchmark's score you should be using to judge where AI genuinely stands today.

Sources: ARC-AGI-3: The New Interactive Reasoning Benchmark — DataCamp, ARC-AGI-3 Dropped — and Frontier AI Scored Less Than 1%, ARC-AGI In 2026: Why Frontier Models Still Don't Generalize
Diagram
靜態題目與互動環境的設計對比,以及人類與AI的分數落差ARC-AGI-3把規則揭露時間點移到行動之後,堵死暴力取樣,人類通關率100%對前沿AI普遍低於1%ARC-AGI-3: Static Puzzle vs. Interactive EnvironmentARC-AGI-1 / ARC-AGI-2 (static)Full input→output examplesshown upfrontVulnerable to brute-force samplingARC-AGI-3 (interactive)Rules revealed only afterthe agent actsBrute-force sampling ineffective100%Human<1%Frontier AI (avg)AGI Bible · agi-bible.com
Feel free to share. Please credit the source.
Ask a Question
Please enter at least 10 characters
Related Articles
Same False Statement, Different Speaker — Accuracy Drops From 98% to 64%: What the New Wave of Benchmarks Reveals Isn't Hallucination, It's Flattery
benchmarks · Sep 05
Are Scaling Laws Hitting a Wall? Why 2026's Compute Race Shifted From "Train Bigger" to "Think Longer"
benchmarks · Aug 25
From 0% to 92.5%, Then Back to 0.37%: What Kind of "Progress" the ARC-AGI Benchmark Actually Reveals
benchmarks · Aug 13
95% of Enterprise AI Pilots Show No P&L Impact — Yet the Winners Are Beating the S&P 500 by 12 Points. Here's the Actual Gap
industry-impact · Oct 06
More Related Topics