If even OpenAI's own internal definition amounts to a financial threshold, does that mean the charter's definition of "outperforming humans in economic value" was just lip service?
Not entirely — the two are actually two expressions of the same underlying definitional logic, just at different levels of specificity. The charter's phrase "highly autonomous systems that outperform humans at most economically valuable work" is itself already a definition revolving around economic output, just without a concrete numerical threshold; the $100 billion profit clause with Microsoft was, in a sense, translating that abstract concept of "exceeding economic value" into a quantified metric both parties could concretely enforce in a contract, without room for interpretation.
This also reveals a deeper issue: any abstract AGI definition, once it needs to be operationalized into a concrete, enforceable contract or policy, has to be converted into some measurable concrete metric — and that conversion process itself introduces new assumptions not present in the original definition (for example, "using profit to represent economic value" is itself already a reduction, overlooking many ways economic impact can manifest beyond profit).
With Demis Hassabis and Marc Andreessen making completely opposite judgments about whether AGI has arrived, is there any way to tell who's more credible?
Rather than asking "who's more credible," a more practical question might be "what standard is each of them presupposing." Hassabis, as a scientist who's spent his career deeply invested in frontier AI research, likely has a standard behind his "still a long way off" claim that leans closer to the technical metrics academic definitions emphasize — versatility, generalization ability, data efficiency. Andreessen, as a venture capitalist who's long been bullish on AI industry investment potential, likely has a standard behind his "already here" claim that leans closer to the practical capability current AI systems already demonstrate in concrete commercial applications — capability already sufficient to create enormous economic value.
Both of their judgments could be internally consistent, reasonable conclusions under each of their own presupposed standards — the real disagreement isn't that one of them misjudged the evidence, but that both started out measuring the same phenomenon with a different yardstick. This is also why, when you see this kind of head-to-head public statement, it's worth figuring out what standard each side is presupposing before rushing to pick a side.
OpenAI's o3 scored 87% on a benchmark — why would anyone still question whether that counts as evidence of progress toward AGI?
The core of this skepticism lies in the fact that "scoring high" and "how you scored high" are two different things. According to commentary, o3's score required massive Test-Time Compute to achieve — meaning the model, when answering each question, invested computational resources far beyond typical usage to repeatedly reason and verify. What benchmark designers typically intend to measure is whether a model can generalize to new problems in a relatively efficient way, not the capability that "as long as you're willing to throw enough compute at it, almost any sufficiently capable model can score high on a specific test."
This is also why understanding the concept of test-time compute helps readers more accurately interpret news like "a model set a new record on some benchmark": the issue isn't just the score itself, but also what computational cost was paid to achieve it — if the answer is "it deployed a massive amount of compute that wouldn't be used under typical conditions," whether that score can be directly equated with "getting closer to AGI" is itself a question that needs further scrutiny.
If even experts can't reach consensus on AGI's definition, should the average reader just not bother worrying about whether something "counts" as AGI?
Rather than saying you shouldn't bother, it's more that where your attention goes needs adjusting. Rather than spending energy on the binary judgment of "does this system count as AGI" (given that even the definition itself lacks consensus, this question may never get an answer everyone agrees on), a more practical approach is to focus on concrete, observable capability changes themselves — like how much a given model's performance improved on a specific task type, whether that improvement came from algorithmic innovation or from stacking more compute, and whether that improvement reliably reproduces in real-world application scenarios.
This is also why this site's content tends to focus more on concrete, verifiable technical developments — like how ARC-AGI benchmark scores have changed, how Test-Time Compute affects performance, or whether Chain-of-Thought Monitoring can actually verify a model's reasoning process — rather than flatly asserting "AGI is just a few years away." Judgments of the latter kind ultimately rest on a definition that still lacks consensus today, and any answer that looks certain is worth pausing on with one question first: under which definition does this hold true?
If you're planning to start following AGI-related news, the first thing worth knowing might be a little counterintuitive: there's currently no single definition of "AGI" that everyone agrees on. This might sound like a harmless terminology quibble, but in reality, this definitional disagreement has genuinely shaped the direction of multi-billion-dollar contracts, litigation battles between companies, and government policy directions.
OpenAI's charter defines AGI as "highly autonomous systems that outperform humans at most economically valuable work" — note that this definition revolves entirely around "economic output," with almost no connection to what most people intuitively associate with the term, like "does the machine have consciousness" or "can the machine think for itself." Google DeepMind CEO Demis Hassabis has publicly stated we are "nowhere near" AGI; venture capitalist Marc Andreessen has said AGI "is already here." Both are heavyweight experts in this field, and they obviously can't both be right at the same time — precisely because the two of them very likely have entirely different things in mind when they each say "AGI."
The most dramatic concrete case in this definitional battle played out between OpenAI and Microsoft. According to multiple media reports, the two companies had an "AGI clause" written into their early agreement: once OpenAI's board formally declared its systems had achieved AGI, Microsoft's existing technology license would automatically become void. But for this clause to actually be enforceable, both sides needed a concrete, determinable standard — and according to reporting by The Information, the standard the two companies privately settled on had nothing at all to do with consciousness, reasoning ability, or generality; it was a financial threshold: as long as OpenAI developed an AI system capable of generating at least $100 billion in profit, that would count as achieving AGI. Given OpenAI's financial state at the time — still burning through massive losses, with profitability not expected until 2029 at the earliest — this definition effectively meant AGI's arrival date depended, in a sense, on financial reports rather than on any technical evaluation. This clause went through multiple rounds of renegotiation in 2026, culminating in April when both sides announced Microsoft's license to OpenAI's technology would extend through 2032, with the AGI clause's own effect changing accordingly — this entire business tug-of-war over "what AGI is" is, in itself, the clearest illustration of how unstable this term's definition currently remains.
Setting business contracts aside, the academic and research community is equally divided over AGI's definition. A 2023 paper published by Google DeepMind deliberately shifted the focus away from "economic value," using "versatility" as the core criterion instead — for example, whether a system can rapidly learn new skills with scarce data, rather than merely surpassing humans on a specific task. The most common layperson's definition tends to be "a system that can perform any cognitive task a human can perform, at least at a human level" — but this definition is equally vague: what exactly does "any cognitive task" include? Does "human level" mean an average person's level, or an expert's level in that field? None of these questions currently have a standard answer.
For readers just starting to follow this field, the most practical takeaway from understanding this definitional battle is: whenever you see a headline like "AGI is coming soon" or "we're nowhere near AGI," the first question worth asking isn't "is this claim correct," but "what definition of AGI does the person making this statement have in mind." A lab CEO, an academic researcher, and a lawyer drafting a business contract might each have three entirely different things in mind when they say "AGI." OpenAI's o3 model scored around 87% on a certain public benchmark, seen by some as a major breakthrough toward AGI — but others point out this score required massive Test-Time Compute to achieve, and the method used may not reflect the kind of "efficient generalization" the benchmark's designers originally intended to measure. This is exactly why understanding the definitional disagreement behind the word "AGI" is the most foundational, and most necessary, first step into this entire field — it directly shapes how you should interpret every subsequent article you read on the topic.