09/10 2026
487
On September 3, 2026 (Eastern Time), OpenAI unveiled its new flagship model, GPT-6 Astra. Co-founder and President Greg Brockman immediately declared, “Welcome to the AGI era.” A few days later (September 6), NVIDIA CEO Jensen Huang also posted on social media, stating, “AGI has arrived.”
Yet just two days before the release, OpenAI CEO Sam Altman had publicly stated that AGI was, at best, “a very poorly defined term,” “just a marketing buzzword.” Such public disagreement between the president and CEO on the same issue is rare in Silicon Valley. Even more intriguing than the internal inconsistency is the basis for OpenAI’s claim that AGI has arrived—they scored 99.9% on the ARC-AGI-3 benchmark test.
The problem? When the ARC Prize Foundation re-evaluated the same model using a standardized testing environment, the score plummeted from 99.9% to 62.7%. Same model, same set of questions—a 37-percentage-point difference. The evaluation agency’s stance was clear: even if Astra scored nearly perfect on ARC-AGI-3, they did not believe it proved AGI had been achieved.
This debate reveals a deeper issue than the technology itself: What criteria should we use to determine whether AGI has arrived?
01. A 'Finish Line' Without Universal Standards
There is no universally recognized technical benchmark for AGI. This sounds absurd—the world’s top labs are investing tens of billions of dollars to chase a goal, yet no one can clearly define what should count as reaching it. The Turing Test has long been seen by the public as a benchmark for machine intelligence, but it is essentially a 1950 thought experiment never designed as a practical evaluation standard. While academia has a widely cited definition—a system capable of “understanding, learning, and performing any intellectual task that a human can”—this definition leaves enormous room for interpretation.
OpenAI’s own charter includes a definition: “Highly autonomous systems that outperform humans at most economically valuable work.” This sounds clear but is riddled with ambiguities: What counts as “most”? Which jobs are considered “economically valuable”? What are the criteria for “outperforming”? OpenAI claims to be nearing the most profound milestone in computing, yet the yardstick is self-defined.
The struggle for definition is never just an academic issue. In 2025, Microsoft and OpenAI revised their partnership agreement, stipulating that OpenAI’s declaration of AGI must be verified by an independent expert panel. By April 2026, the two sides amended the agreement again, removing key contractual arrangements tied to AGI. The term AGI serves as both a contractual trigger and a valuation lever—even a marketing pitch. Whoever controls the definition of AGI controls the pricing power of capital.
02. Giving AGI a 'Ruler'
Academia has not stood idly by amid this chaos.
In October 2025, Turing Award winner Yoshua Bengio, along with global research institutions, published the paper A Definition of AGI, proposing a quantifiable definition and testing framework for AGI. Their standard: AGI is “artificial intelligence that matches or surpasses the cognitive versatility and proficiency of a well-educated human adult.”
The research team, drawing on the most empirically validated psychological model—the Cattell-Horn-Carroll (CHC) theory—broke down general intelligence into 10 measurable cognitive domains, including common sense knowledge, reading and writing, mathematics, fluid reasoning, working memory, storage and retrieval of long-term memory, visual and auditory processing, and processing speed. Assessment uses a 100-point scale, with equal weighting across domains (10% each). A total score of 100 corresponds to the AGI level defined in the paper. By this standard, GPT-4 (2023) scored only 27 points, while GPT-5 (2025) improved to 57—rapid progress, but still far from 100.
In March 2026, Google DeepMind proposed its own framework. They published Measuring Progress Toward AGI: A Cognitive Framework, refining AGI evaluation into 10 key cognitive domains and introducing a three-stage assessment protocol. DeepMind, in collaboration with Kaggle, offered a $200,000 prize for researchers worldwide to design truly effective AGI tests.
In January 2026, Andrew Ng proposed a more real-world approach: the Turing-AGI test. This test requires AI or human professionals to complete authentic professional tasks over multiple days using computers and internet tools, with judges designing tasks and evaluating performance throughout. Ng argued that existing AI benchmarks are too narrow and easily gamed, while the Turing-AGI test focuses on AI’s performance in real economic work, aligning better with society’s general understanding of AGI—capable of completing most knowledge-based jobs like humans.
These efforts share a common direction: transforming AGI from a vague philosophical concept into measurable, verifiable technical metrics.
03. The 'Definition War' Has Just Begun
Returning to OpenAI’s AGI declaration. Astra’s 99.9% score on ARC-AGI-3 is not meaningless—the ARC Prize Foundation noted that Astra met or exceeded human benchmarks in 96% of test scenarios. However, ARC Prize cautioned that ARC-AGI-3 puzzles are designed to prevent memorization, testing the model’s exploration and trial-and-error abilities on novel tasks. High scores on puzzles do not prove AI can independently think and create when faced with “real-world problems without standard answers.” In other words, solving puzzles ≠ solving unknown problems in the real world.
A deeper issue is that OpenAI’s Astra introduces a “recurrent depth” architecture, shifting part of its reasoning into hidden states. This means neither the reasoning process nor how answers are derived is transparent or verifiable. How can an intelligent system that cannot be transparently validated be independently judged as having achieved AGI? This is not just a technical issue but a trust issue.
By autumn 2026, the focus of the AGI race had quietly shifted from model capabilities to definition rights. As technical gaps narrow to months, the weight of valuation and narrative rises. Whoever first claims “AGI” as their own asset gains pricing power over capital and markets. Multiple media outlets report that OpenAI’s IPO target valuation nears $1 trillion, though its latest private valuation ($852 billion) still lags behind rival Anthropic’s $965 billion. Announcing “the AGI era is here” at this juncture is transparently motivated by commerce.
Perhaps AGI never had an objective, discoverable “standard answer.” It resembles a moving finish line—just as you think you’re nearing it, the definition recalibrates. As Altman himself said, AGI is largely a “marketing term.” But that does not mean it lacks meaning. On the contrary, the debate over AGI standards itself drives industry progress. Bengio’s quantifiable framework, DeepMind’s cognitive assessment, Ng’s Turing-AGI test—these efforts are transforming AGI from a vague vision into measurable, verifiable technical milestones.
Who defines AGI defines the rules of the next era. And this war over definition has only just begun.
- End -