On September 3, 2026, OpenAI dropped a headline most companies would never dare write themselves: GPT-6 Astra might be the first step into the "AGI era." President Greg Brockman said it plainly, calling the model's release a moment when "it's not unreasonable to feel that we are now in the AGI era." Bold claims deserve a close read, so let's separate what the benchmarks actually show from what everyday users and businesses can expect to use this week.
What OpenAI Announced
GPT-6 Astra launched with a rollout plan rather than an instant free-for-all. It reached enterprise customers and OpenAI's Daybreak cybersecurity program first, on September 5, with Plus, Pro, Business, and Enterprise ChatGPT tiers following in the days after, alongside access via the OpenAI API, AWS Bedrock, and Microsoft Azure.
The Benchmark Numbers
The scores OpenAI is leaning on hardest involve reasoning, coding, and offensive security testing. Compared with the prior model, GPT-5.6 Sol, the jump is dramatic on paper:
| Benchmark | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| ARC-AGI-3 (enhanced tools) | 99.9% | 7.8% (standard) |
| ExploitBench | 100% | 78.5% |
| FrontierMath Tier 4 | ~98% | — |
One detail worth flagging: Astra's ARC-AGI-3 score drops to roughly 66% under a standard testing harness, only climbing near 100% with enhanced tooling enabled. That gap matters — it's the difference between a lab-optimized result and what a typical user's setup will produce.
Is This Actually "AGI"?
Brockman himself hedged, admitting that a clean definition of AGI doesn't exist and that progress has arrived "in bits and pieces" rather than a single breakthrough. That's a notably softer claim than the "AGI era" framing suggests. Analysts have also pointed out that OpenAI's launch materials skipped GDPval, the benchmark the company itself designed to measure real-world economic task performance — an odd omission if the goal is proving Astra can replace human labor at scale. Until that number shows up, the AGI claim is more marketing framing than settled fact.
What Computer-Use Actually Means Today
Behind the AGI headline is a more concrete, and arguably more useful, story: Astra can operate a computer the way a person does, rather than just answering questions about one. OpenAI's demos show it:
-
Filling out forms and updating CRM records
-
Working inside spreadsheets and running Python-based data analysis
-
Building and testing websites, and using engineering tools like KiCad and FreeCAD
-
Organizing calendars and conducting multi-step web research
-
Responding to voice instructions to draft legal documents or prototype 3D games
On OSWorld 2.0, a benchmark for exactly this kind of software navigation, Astra scored 72.6% on the offline subset — about 47% faster than its predecessor. That speed claim lines up with reporting that Astra can "zip through spreadsheets" and web pages at what one outlet described as superhuman pace. For businesses, this is the part worth paying attention to now: not whether the model is generally intelligent, but whether it can reliably finish a multistep office task unsupervised.
Pricing and the Shift Away from Tokens
API pricing for Astra sits at $10 per million input tokens and $50 per million output tokens in standard mode, roughly doubling for a faster response mode. But Brockman argued the token-based model is already outdated, saying "pricing tokens doesn't make any sense" when what businesses actually care about is the cost of completing a task. Expect OpenAI to keep pushing toward price-per-outcome billing as computer-use agents take on longer workflows.
The Caveats Getting Less Attention
Astra's ExploitBench score isn't just a math flex — OpenAI says the model meets a "critical cybersecurity capability threshold," meaning it can find and exploit previously unknown security flaws without human oversight. That's why the most capable version is currently restricted to enterprise and Daybreak program customers, with consumer-facing variants built to refuse advanced cybersecurity tasks. OpenAI added extra monitoring following a security incident involving Hugging Face in July, though some outside experts have questioned whether those safeguards go far enough. A government review of the model reportedly took place under a voluntary framework from the Trump administration, but details of that review haven't been made public.
Whether GPT-6 Astra marks the true start of AGI is still an argument, not a fact — but its ability to actually drive software, not just describe it, is the part that will show up in real workflows first. Expect the next few months to be less about benchmark charts and more about how well Astra holds up when it's left alone with a spreadsheet and a deadline.
-EditorZ

Post a Comment