GPT-6 Astra Lands, and OpenAI Says the A in AGI Might Be Due

GPT-6 Astra Lands, and OpenAI Says the A in AGI Might Be Due

Every few months a tech company whispers "this one's different" right before shipping something that is, in fact, mostly the same thing but faster. GPT-6 Astra might actually be the exception — or at least OpenAI is betting its entire PR budget that it is.

Astra Rolls Out, Benchmarks Go Bananas

OpenAI began rolling out GPT-6 Astra on September 3rd, calling it the product of "years of research and big bets" across pretraining, reinforcement learning, and alignment. The headline numbers are genuinely eye-popping: a 98% score on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and a perfect 100% on ExploitBench, OpenAI's own cybersecurity benchmark.

The rollout is deliberately staggered. Companies in OpenAI's application-based cybersecurity program get first access, followed by a phased expansion to ChatGPT Plus, Pro, Business, and Enterprise plans, plus the API and AWS. Enterprise admins have to flip it on manually — it's off by default at launch, which for a model OpenAI itself flagged as hitting the "Critical" threshold for cybersecurity risk, feels less like caution theater and more like an actual seatbelt.

AGI Talk Is Cheap, Root Access Isn't

OpenAI president Greg Brockman called it a possible "generational leap" toward AGI, which is the kind of sentence that ages either brilliantly or hilariously and rarely anything in between. The more interesting tell is buried lower in the announcement: additional safeguards were added to Astra specifically after the Hugging Face breach, meaning this launch was shaped as much by recent security scar tissue as by benchmark chasing.

That's the part worth watching if you run a business rather than write headlines about one. A model this good at "computer use" and "professional work" is also a model that's very good at automating tasks your team currently pays humans to babysit — for better and for worse, often in the same afternoon.

AGI or not, a model that scores 100% on an exploit benchmark is a model everyone's security team should be quietly Googling this week.

If you're weighing where AI actually belongs in your business workflow versus where it's just a shiny distraction, that's a conversation we'd genuinely enjoy having with you.

Source: 9to5Mac