OpenAI says Astra is AGI. Its own launch page never uses the word.

Greg Brockman says GPT-6 Astra has crossed into artificial general intelligence, a term with no agreed definition, no threshold and no independent test to fail.

Share
A man leans back from his desk as a spreadsheet fills itself in on the monitor in front of him
OpenAI says Astra can run a computer the way a person does, filling spreadsheets and forms without step-by-step direction. Photo: The Glass

OpenAI released GPT-6 Astra on September 3 and said the machines had caught up.

"Welcome to the AGI era," president Greg Brockman told reporters at the close of the launch briefing.

Artificial general intelligence is the industry's finish line. It is also the least agreed-on term in it.

Two colleagues look at a wall screen showing two very different scores for the same AI model
Astra scored 99.9 per cent on ARC-AGI-3 through an adapter OpenAI wrote, and 62.7 per cent on ARC Prize's neutral setup. Photo: The Glass

OpenAI's own definition is a system that outperforms humans at most economically valuable work.

There is no test for that, no threshold and no umpire.

The hardest definition OpenAI ever had was in its contract with Microsoft, where AGI meant systems capable of generating about $100 billion in profit. That clause was deleted in April.

Brockman now calls AGI "a mission concept or spiritual concept."

OpenAI's own announcement never uses the word. It calls Astra "the world's most intelligent and aligned model."

What Astra actually does is narrower and far more concrete.

It runs a computer the way a person does, filling spreadsheets, completing forms and moving between web pages, which Brockman says it can do at superhuman speed.

It was trained on more than 100,000 graphics chips at OpenAI's Texas site, the company's largest run.

It is also the first OpenAI model judged able to find and exploit unknown security flaws without a person directing it, which is why the full version is locked to vetted customers.

Then the numbers.

Astra's headline result is 99.9 per cent on ARC-AGI-3, a test built from problems no model has seen before.

That score came through an adapter OpenAI wrote, which lets the model keep its own working between moves.

On ARC Prize's neutral setup, the same model scored 62.7 per cent.

ARC Prize verified both, then said plainly it is not claiming Astra is AGI, because the test has a tightly bounded scope and does not represent the open-endedness of the real world.

It is still a genuine jump. The model Astra replaces scored 7.8 per cent on the same test, and ARC Prize co-founder François Chollet has pulled his own AGI forecast forward.

Everywhere else the picture flattens.

On Artificial Analysis's intelligence index Astra scored 61, level with the model it replaces and behind Anthropic's Claude Fable 5.1 on 66, at 2.5 times the price.

On Humanity's Last Exam, a set of expert-level questions, it scored 57.2 per cent against Fable 5.1's 65.

On a benchmark adapted from OpenAI's own dataset of real work across 44 occupations, Astra went backwards.

OpenAI did not publish its own result on that one, the test written to measure the exact definition of AGI the company uses.

A model that can operate a computer better than you can is real, and it shipped this week.

A word that means whatever the company saying it needs it to mean is not.