Earlier this month, speaking at the launch event for OpenAI’s new GPT-6 Astra, Greg Brockman, OpenAI’s president, declared that the “AGI era” had begun. AGI, or artificial general intelligence, generally refers to AI systems that can match or surpass human abilities across virtually every cognitive task.
Long regarded as the holy grail of AI research, AGI has for decades remained a goal somewhere over the horizon. And now, apparently, it’s here—at least, if Brockman is to be believed. Nvidia CEO Jensen Huang, another one of the industry’s most influential voices, seems convinced. He congratulated OpenAI on its achievement, tweeting, “AGI has arrived.”
There is, however, a problem with declaring that AGI has arrived: There is no universally accepted threshold for what counts as AGI, and many researchers would dispute that today’s models have crossed it. That ambiguity also gives companies plenty of room to claim the milestone early, especially when doing so carries obvious marketing value. It feels, in a sense, like the race among cellular carriers to slap the next “G” on their networks before the underlying technology fully satisfies the technical standard.
Now, Astra is an impressive model. It almost aced a benchmark test designed to resist attempts by model makers to train specifically for the test, which has been a major problem with AI benchmarks. The ARC-AGI-3 benchmark, developed by AI researcher François Chollet, tests a model’s ability to encounter a completely unfamiliar situation (in this case a series of video games), figure out how it works, learn to play efficiently, and win.
But Astra’s score depends heavily on the software used to administer the test. Running through OpenAI’s own “harness,” the software layer that lets the model interact with the benchmark, Astra scored 99.9%. Using ARC-AGI-3’s standard testing setup, designed to give models a more uniform interface, Astra scored 62.7%.

In any case, the ARC-AGI-3 benchmark doesn’t represent the finish line in the race for AGI, Chollet says, because it tests a non-exhaustive set of attributes at very small scales.
“The real world features much longer time horizons for continual learning compared to ARC 3 games (decades vs minutes), much larger world modeling complexity, much greater goal ambiguity, more greater exploration spaces, etc.,” Chollet says in an email to Fast Company.
“So solving the benchmark is a strong sign of progress (as prior systems did not exhibit these attributes), but it is not proof of AGI, and that was never the point.”
Continuing a trend
“We are nowhere near AGI,” says New York University professor and noted AI skeptic Gary Marcus in a message to Fast Company. “That’s just marketing by people who either don’t know the original definitions or are deliberately lowering the bar.”
Marcus also points out that Astra doesn’t look like a radical departure from previous models. If Astra truly represented AGI, Marcus explained on his Substack earlier this month, it should be decisively outperforming rival models, not merely matching them in everyday use.

Marcus pointed to Astra’s performance on the Epoch Capabilities Index, a composite measure developed by the independent nonprofit Epoch AI that combines results from numerous AI benchmarks. Astra did set a new record, scoring 169 compared with the previous high of 163. But Epoch’s analysis found that the jump was still consistent with the existing trajectory of AI progress.
In other words, Astra may be better than what came before, but its improvement does not look like the kind of dramatic break from the past that Marcus argues AGI should represent.
Influential AI researcher Andy Konwinski, who confounded Databricks, Perplexity, and Laude, suggests that it’s not necessary to go swimming through benchmark numbers to see the gaps between artificial and human intelligence.
“These systems can’t yet think on their own for long,” he tells Fast Company. “They can build complex software or find a lot of bugs, but that’s still narrow. They’re really good at coding, but most of the world’s value doesn’t come from software engineers. They’re not growing our food, building our solar panels, or running our government.”

Not “general” enough yet
Anthropic researcher Jacob Coxon resigned last week because he believes AI labs currently can’t mitigate the risks of increasingly intelligent and autonomous AI systems. That caused a chorus of voices from inside and outside big AI labs to call for a slowdown in AI capabilities research.
But the most immediate threats from AI systems relate to specific skill sets, such as AI’s ability to identify and exploit software security vulnerabilities, not to a sudden increase in general intelligence.
Ben Goertzel, the data scientist who coined the term “artificial general intelligence,” or AGI, in 2005, says his impression after using Astra is that the model is very strong in some areas and still weak in others.
“It is superhuman at many aspects of math and programming—though not, I think, at radical innovation in math or programming—and it is smarter than me at plenty of other things too,” Goertzel says. “So we have a system that is superhuman at some things and subhuman at others, and comparing it to a human being ends up being complicated rather than a simple yes or no.”
Goertzel says that using a capable harness makes Astra a lot more useful, but, he adds, there remain missing aspects that no amount of software scaffolding supplies. These include “the way Astra models and understands itself . . . the way it remembers its whole life (or doesn’t) and brings those memories to bear on what it says and decides now, and the way it coordinates its different goals and aspirations, to the extent it has any, which isn’t much,” explains Goertzel, who now leads the decentralized AI research organizations SingularityNET and the ASI Alliance.

Brockman isn’t even the first to publicly declare AGI accomplished. Huang made the claim back in March. In July OpenAI CEO Sam Altman said we’re already in the Singularity, a phase when AI models have taken over and accelerated the development of new and better models, leaving humans largely out of the loop.
It’s easier to declare AGI when there’s a lot of debate over what AGI even is. There’s no consensus on a single definition.
The AI labs use different definitions of AGI than independent researchers do—definitions that may make reaching the goal easier. OpenAI has modified its own definition several times. “Highly autonomous systems that outperform humans at most economically valuable work,” the current version reads.
The Center for AI Safety has a more detailed definition that requires a model to hit high thresholds in 10 different cognitive domains spanning reasoning, memory, and perception. And there are many others.
“Nobody has a definition of AGI that’s worth its weight,” Konwinski says, “so who cares whether it’s ‘here’?”