Lead · Milestones
What separates the top model from the rest isn't whose first attempt is smarter — it's who's still willing to try a 47th revision after 46 rejections.
Elena Voss
·
September 05, 2026
In June 2026, a cross-institutional research team released a new benchmark called AutoLab, specifically designed to test whether frontier models can work like real researchers — spending hours, sometimes over a dozen hours, cycling through "inspect the code, propose a change, run the experiment, read the result, refine again" — rather than being judged, as most existing benchmarks do, on...