Bible Network Crypto DeFi Onchain RWA AI Agent Stablecoin Chain SAFU CryptoTax DeFAI AGI Claude Me Claude Skill Claude Design Claude Cowork
Independent Media
Not affiliated with any project
Artificial General Intelligence, Decoded from Theory to Reality
agi-bible.com
LATEST
Will AI "Fake Being Good"? What Apollo Research and OpenAI's Scheming Evaluations Actually Found  ·  How Long Can AI Work Autonomously? METR's Time Horizon Doubles in Months — But the Number Is Messier Than It Looks  ·  AI Regulation Splits Three Ways: The Paths the EU, US, and China Are Each Taking in 2026  ·  From 0% to 92.5%, Then Back to 0.37%: What Kind of "Progress" the ARC-AGI Benchmark Actually Reveals  ·  How Many Jobs Has AI Actually Taken? The 2026 Data Doesn't Quite Match the Headlines  ·  How Many Years Until AGI, Really? Lab CEOs and Academic Researchers Look at the Same Evidence and Reach Opposite Answers
benchmarks

From 0% to 92.5%, Then Back to 0.37%: What Kind of "Progress" the ARC-AGI Benchmark Actually Reveals

30-Second Version · For the impatient
ARC-AGI-2 went from zero across the board to 92.5% — looking like an intelligence explosion — until ARC-AGI-3 launched and the same models got knocked back to 0.37% on day one. What benchmarks teach us: progress only ever happens on the exact slice you decided to measure.

Full Explanation +
01 · Why did this happen?

If benchmark scores can be knocked back to square one so easily by the next-generation test, do these scores still mean anything for assessing AI progress?

Yes, but the meaning is local rather than comprehensive. ARC-AGI-2's climb from 0% to 92.5% genuinely reflects real progress on the specific slice of "static abstract reasoning ability" — this isn't inflation or gaming; capability on this clearly defined task type genuinely improved substantially. The issue isn't that the score itself is untrustworthy, but that the inference "this score represents overall intelligence" doesn't hold up.

A more accurate way to think about it: treat each benchmark as a ruler of a specific shape — it precisely measures one dimension, but can't tell you what happened along dimensions outside its scope. ARC-AGI-2 measures static abstract reasoning; ARC-AGI-3 measures interactive exploration. The gap between their scores isn't a contradiction — it reveals that "AI capability" itself is a multidimensional concept, with no single number able to fully capture it.

02 · What is the mechanism?

What does NVARC's low-cost model beating the high-cost brute-force approach actually reveal?

This directly challenges a common assumption: that a higher score directly equals higher capability, and that the only path to improved capability is throwing more compute at the problem. NVARC used a fine-tuned 4-billion-parameter small model to achieve competitive accuracy at a cost far below the brute-force sampling strategy — showing that, for specific task types, a small-scale system with targeted optimization can outperform a large-scale attempt built on simply stacking computational resources, at least on efficiency.

This finding offers an important complementary perspective to the compute-scaling narrative: the broad observation that Compute Scaling keeps delivering capability gains isn't wrong, but that doesn't mean the "efficiency" dimension can be ignored. The benchmark designers' deliberate decision to factor per-task cost into scoring is itself a reminder to the industry that an evaluation culture focused only on the top score, without accounting for how much resource it took to achieve that score, risks rewarding engineering-level resource stacking rather than genuine cognitive-level breakthroughs.

03 · How does it affect me?

How does the "interactive exploration" that ARC-AGI-3 tests concretely relate to agentic AI mentioned in earlier articles?

The connection is very direct. ARC-AGI-3 deliberately gives a model no instructions at all, requiring it to infer a game's rules through exploration, experimentation, and trial and error — this is precisely the core capability agentic AI needs in real-world workflows: facing a novel situation with no ready-made script, where the rules have to be figured out on the fly, can the system efficiently build an understanding of the situation, plan actions, and adjust strategy based on feedback? The ARC-AGI technical report explicitly states that this test evaluates "efficiency at acquiring new skills through exploration, model formation, goal inference, and planning," measuring performance through action efficiency relative to a human baseline, rather than just a binary success/fail outcome.

This is also exactly why it's especially worth taking note that Gemini 3.1 Pro approached human-level performance on ARC-AGI-2 (static reasoning) yet scored only 0.37% on ARC-AGI-3 (interactive exploration): if an enterprise assessing whether a given model is suitable for a highly autonomous Agentic Workflow references only its score on static benchmarks, it risks seriously overestimating that model's actual ability to handle unfamiliar situations requiring active exploration.

04 · What should I do?

Will this pattern of "a test gets cracked, so we switch to a harder one" keep continuing? What does this mean for readers tracking AGI progress?

Looking at the ARC-AGI series' evolutionary trajectory, this pattern is quite likely to continue. The ARC-AGI series itself had an evolution path planned back in 2022; ARC-AGI-2 was deliberately designed to resist known weaknesses of the previous generation (like brute-force search and data contamination), while ARC-AGI-3 deliberately shifted the test axis from static reasoning to interactive exploration — each new generation's design logic responds directly to limitations exposed once the previous generation got cracked.

For readers, this means that when tracking AGI progress, rather than fixating on whether any single benchmark's score has been "cracked," it's more worthwhile to watch the signal revealed by this whole cycle of test evolution: if a test for a given capability dimension keeps getting cracked relatively easily, that suggests current model architectures have matured on that dimension; if a next-generation test knocks the score back near zero, that reveals where the next capability frontier that needs breaking through actually lies. This ongoing process of continually finding "what the model still can't do" is itself a signal more informative for judging how far AGI actually is than any single score.

Full Content +

The ARC-AGI benchmark series is one of the few evaluation tools widely regarded as genuinely testing "novel reasoning ability" rather than "patterns already seen in training data." Its creator, François Chollet, calls this the "only thing that actually matters in intelligence": not retrieval, not pattern-matching across a vast training corpus, but solving a genuinely novel abstract puzzle from a minimal number of examples. This design logic has made the ARC-AGI series' score trajectory one of the most dramatic illustrations of AI capability progress available today.

ARC-AGI-2: From Zero Across the Board to 92.5% in About Seventeen Months

When ARC-AGI-2 launched in March 2025, frontier models — the very same systems posting strong scores elsewhere, passing bar exams, and generating usable production code — scored exactly 0% on this new test. Not close to zero — precisely zero. That number carried a message all by itself: whatever capability these models had previously demonstrated was a different thing entirely from the "solve a genuinely novel problem from minimal examples" capability ARC-AGI-2 was designed to measure.

What followed was a remarkably fast pace of progress. OpenAI's GPT-5.2 became the first to cross 50% in December 2025, landing around 54%; by April 2026, multiple teams had pushed scores to roughly 98%. As of August 2026, tracking service BenchLM's leaderboard shows GPT-5.6 Sol leading at 92.5%, with Claude Opus 5 (90.4%) and GPT-5.5 (85%) close behind — measured against a benchmark where human expert panels achieve full completion and average individual human performance sits around 66%, these frontier models have clearly surpassed typical human-level performance.

The Efficiency Threshold: The Score Isn't the Only Metric — How You Got the Score Matters Too

ARC-AGI-2's designers deliberately built in an angle that's easy to overlook: per-task cost. In the 2025 ARC Prize competition, the NVARC team used a fine-tuned 4-billion-parameter model to achieve 24% accuracy at roughly $0.20 per task — an efficiency showing that outperformed brute-force approaches spending up to $200 per task by exhaustively sampling thousands of candidate solutions and filtering them. The benchmark's designers explicitly consider this brute-force sampling strategy "a loophole, not a solution," because it reflects engineering-level resource stacking rather than cognitive-level improvement in reasoning ability. This is also why prize-eligible submissions are required to publish their compute budget and cost-per-task metrics — to prevent scores from being inflated purely by throwing more compute at the problem.

ARC-AGI-3: The Moment a Score Gets Cracked, the Difficulty Axis Shifts

Just as ARC-AGI-2 scores were rapidly approaching saturation, the design team had already, on March 25, 2026, unveiled ARC-AGI-3 at Y Combinator's headquarters — the first fully interactive benchmark: with no instructions given at all, a model has to rely on exploration, experimentation, and trial and error to infer the rules of a game. What it tests is "the efficiency of acquiring new skills through exploration," not static abstract reasoning. The results were sobering: the strongest frontier model at the time, Gemini 3.1 Pro, scored just 0.37% on ARC-AGI-3. A model that had just approached human-level performance on ARC-AGI-2 was, once switched into a scenario testing "active exploration" rather than "passive reasoning," nearly knocked right back to square one.

What This Means for Your Money

For readers assessing the pace of AI industry progress, the full trajectory of the ARC-AGI series offers an important calibration framework: any single benchmark's score only reflects the specific slice of capability it was deliberately designed to measure, not the full picture of "intelligence" itself. ARC-AGI-2's climb from 0% to 92.5% looks like an explosive capability leap, but the immediate follow-up with ARC-AGI-3 demonstrates that progress on static abstract reasoning and progress on interactive exploration ability are nearly independent curves. This matters particularly for assessing the feasibility of enterprises adopting agentic AI — autonomous systems that need to actively explore, plan, and adapt. Assuming a model already has the capacity to autonomously handle novel situations just because it scores well on static benchmarks risks overestimating current systems' actual reliability. The benchmark designers' deliberate move to keep an "unsaturated" test axis available is itself a reminder: behind any score that looks like it's already been cracked, there may still be an entire dimension of capability that hasn't been measured yet.

Diagram
ARC-AGI-2 分數攀升,ARC-AGI-3 立即歸零折線圖顯示 ARC-AGI-2 分數從 2025 年 3 月的 0% 快速攀升至 2026 年 4 月的約 98%,同時標示 ARC-AGI-3 於 2026 年 3 月推出時前沿模型僅拿到 0.37%ARC-AGI-2 Rises, ARC-AGI-3 Resets0%100%Mar 20250%Dec 202554%Apr 2026~98%ARC-AGI-2 (static reasoning)Mar 20260.37%ARC-AGI-3(interactive)AGI Bible · agi-bible.com
Feel free to share. Please credit the source.
Ask a Question
Please enter at least 10 characters
Related Articles
How Many Years Until AGI, Really? Lab CEOs and Academic Researchers Look at the Same Evidence and Reach Opposite Answers
perspectives · Aug 13
How Long Can AI Work Autonomously? METR's Time Horizon Doubles in Months — But the Number Is Messier Than It Looks
milestones · Aug 13
How Many Jobs Has AI Actually Taken? The 2026 Data Doesn't Quite Match the Headlines
industry-impact · Aug 13
Will AI "Fake Being Good"? What Apollo Research and OpenAI's Scheming Evaluations Actually Found
risk-alignment · Aug 13
More Related Topics