
Share
OpenAI's president says the company has crossed into "the AGI era" with a computer-use model that scores 98.6% on ARC-AGI-3. But the omission of its own real-world work benchmark raises questions about what's actually being measured.
OpenAI has released GPT-6 Astra, and the company wants investors and enterprises to understand this launch differently than the eleven that preceded it. Co-founder and president Greg Brockman closed a press briefing Wednesday with a line built for headlines: "Welcome to the AGI era." That is not a phrase OpenAI has used lightly before. It is also, on inspection, a claim resting on a narrower evidentiary base than the framing suggests.
The thesis is straightforward. Astra is designed to operate a computer the way a person does, clicking through browsers, filling spreadsheets, drafting documents, without the API integrations that have defined enterprise AI deployment for the past three years. Brockman calls this the removal of a bottleneck. "We've been bottlenecked over this gigantic era by people writing connectors and very painstakingly building these connections into all these tools that people can already use," he said. With computer use working well enough, an agent can "zip through spreadsheets, fill out forms, navigate across web pages."
That is a real architectural shift, if it holds up in production. It is also a shift enterprises should evaluate on cost and reliability grounds, not rhetoric.
OpenAI reports Astra scoring 97.6% on FrontierMath Tier 4 v2, 74.1% on DeepSWE v1.1, 95.9% on BenchCAD, 96% on GPQA Diamond, a perfect 100% on ExploitBench, and 98.6% on ARC-AGI-3. On an offline OSWorld 2.0 subset, Astra hit 72.6% accuracy in roughly 40 minutes per task, versus GPT-5.6 Sol's 65.7% in about 75 minutes, a 47% reduction in task time. Those are strong numbers on their face.
The ARC-AGI-3 figure deserves scrutiny, and OpenAI's own evaluation notes provide the reason. Astra ran under OpenAI's Responses API harness, while comparison models used different configurations. That is not a trivial footnote. In August, Nvidia reported its AVO architecture hit 100% across all 25 ARC-AGI-3 environments using Claude Opus 5 as the underlying model, a model Nvidia said scored roughly 30% on its own. The gain came entirely from scaffolding: persistent memory, tools, feedback loops, recovery mechanisms. Nvidia's conclusion was direct: long-horizon capability emerged from the complete agent system, not the foundation model.

That distinction matters for anyone trying to price AI capability into a business case. A model that scores well because of an elaborate harness is a different asset than a model that scores well on raw weights. One Reddit commenter reacting to the Nvidia result put it bluntly: "Let's see if the capabilities generalise or if it was just overtrained on this specific benchmark." Others have argued the opposite, that stripping context between actions makes ARC-AGI-3 an unrealistic test of how production agents actually work. Both critiques can be true simultaneously, and neither resolves cleanly in OpenAI's favor or against it.
Aidan Clark, an OpenAI researcher who worked on Astra, described it as the company's largest training run yet, the first pretrained with more than 100,000 DBUs on OpenAI's Stargate infrastructure, and the first where prior models supervised training of the next generation. "Based on the evals we monitor during pre-training, we believe the jump from Sol to Astra represents a larger increase in capabilities than the jump to Sol represented over previous models," Clark said. That is an internal claim, not an independently verified one, and it should be weighted accordingly.
The more conspicuous gap sits elsewhere. OpenAI's own GDPval benchmark, built in 2025 specifically to measure performance on 1,320 real-world tasks across 44 occupations and nine industries, legal briefs, engineering designs, nursing care plans, customer support work, is absent from the Astra launch materials. That is the benchmark OpenAI itself designed to move the AGI conversation away from academic tests and toward observable economic output. If Astra's central claim is that enterprises can now delegate materially more work to AI, GDPval is the company's most directly relevant internal yardstick for proving it. Its absence does not invalidate the other results, but it leaves an analytical hole exactly where the AGI framing needs support most. GDPval's current one-shot design does not capture the long-horizon, multi-application work Astra is pitched to handle, which may explain the omission, but that mismatch cuts both ways: it also means the benchmark most tailored to Astra's use case is the one OpenAI chose not to show.
Pricing reinforces the enterprise-first positioning. API access to gpt-6-astra runs $10 per million input tokens and $50 per million output tokens under Standard pricing, with Fast mode at 2.5x speed for 2x the price. That places Astra well above the budget tier occupied by models like Meta's Muse Spark ($0.30 per million combined) or DeepSeek-V4-Flash off-peak ($0.88 per million), and even above Gemini 3.7/3.8 Flash at $4.50 combined. Astra is not competing on cost. It is competing on the claim that it does more work per dollar by collapsing multiple specialized tools into one general-purpose agent.
Brockman's own framing is worth taking seriously precisely because it is more careful than the "AGI era" soundbite implies. "Everyone has a different definition of AGI," he said. "It's a much more gray, fuzzy thing." Pressed on whether Astra qualifies, he answered: "For me personally, I do think we're there. I think there's a pretty good argument for it." That is a hedge dressed as a declaration, and enterprises evaluating Astra should treat it as such. The computer-use capability is the concrete story here: a model that can navigate software like a human user, reducing integration overhead that has bottlenecked enterprise AI deployment for years. Whether that constitutes AGI is a definitional argument. Whether it reliably completes multistep workflows at a lower cost than existing tool-integrated systems, and whether OpenAI can substantiate that with its own occupational benchmark rather than a curated set of academic scores, is the question that will determine Astra's actual return on investment. Until GDPval numbers appear, the AGI claim rests on a mosaic of specialized tests rather than the broad economic proof OpenAI built its own benchmark to provide.
Tags
Original Sources
'Welcome to the AGI era': OpenAI launches GPT-6 Astra
↗ https://venturebeat.com/technology/welcome-to-the-agi-era-openai-launches-gpt-6-astra
OpenAI launches new Astra model amid growing scrutiny ... - Reuters
↗ https://www.reuters.com/legal/litigation/openai-launches-new-astra-model-amid-growing-scrutiny-over-agents-safety-2026-09-03
AI models are becoming unknowable - Axios
↗ https://www.axios.com/2026/09/04/astra-openai-how-ai-models-think
Bernie Sanders floats ban on superintelligent AI - Axios
↗ https://www.axios.com/2026/09/03/bernie-sanders-superintelligence-ban-ai-pause
"Welcome to the AGI era," OpenAI says as GPT-6 Astra debuts - Axios
↗ https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman
About the author
Marcus began tracking AI's market implications in 2016, noticing AI-related patent filings accelerating ahead of earnings upgrades before most of the sell-side had caught on. A former fixed-income quantitative analyst, he spent two decades building models that priced risk across emerging markets before pivoting to cover the economic impact of AI full-time. His writing translates opaque technical developments into clear risk/reward terms — and he's rarely diplomatic about the gap between AI valuations and underlying fundamentals. He believes most market participants still underestimate AI's long-run deflationary effect on knowledge work.
More from The Analyst →This Week's Edition
6 September 2026
41 articles
Related Articles
Related Articles
More Stories
© 2026 Cedar & Bloom. All rights reserved.