OpenAI says GPT-6 Astra may mark the beginning of the AGI era. The tech press response, according to coverage of the rollout, was “model fatigue.”
Those two things are both true, and the gap between them is the most interesting story in AI right now. Not the model. The gap.
I’ve reviewed enough launches to know the pattern. A company claims state of the art, posts near-perfect benchmark scores, invokes AGI, and waits for the internet to lose its mind. Astra did all of that. Astra got a collective nod and a scroll past. Something has shifted in how we receive these announcements, and it isn’t that the models stopped improving.
The benchmark problem OpenAI actually admitted
Give credit where it’s earned: the most honest thing in Astra’s system card is a confession. OpenAI acknowledged concerns that exposure to historical software vulnerabilities may have affected benchmark results. In plain terms: if your model trained on the internet, and the internet contains write-ups of every famous security bug ever disclosed, then testing it on those bugs doesn’t measure reasoning. It measures recall with extra steps.
So they built something better. Two novel benchmarks, including an internal “ExploitBench” evaluation dataset containing only recent vulnerabilities disclosed after Astra’s training cutoff — the internal port covers June through August 2026. That’s the right instinct. You cannot claim generalization on a test the model may have already read.
This is the part I want more labs to copy. Not the AGI framing. The methodological paranoia. A company that builds a benchmark specifically designed to make its own model look worse is doing science. A company that posts a leaderboard screenshot and calls it a day is doing marketing.
Why near-perfect scores stopped landing
Near-perfect scores on AI benchmarks used to be the headline. Now they read like a warning label. When a model saturates the tests, the tests are done being useful, and everyone in the field knows it. The score tells you the ceiling of the measurement, not the ceiling of the model.
Which puts reviewers like me in an awkward spot. The number I’d normally quote has become the least informative thing in the announcement. What I actually want to know:
- How does Astra perform on problems nobody has written a solution to yet
- Where does its reasoning break down, and does it know when it’s breaking down
- What happens on the security work when the vulnerability is genuinely novel, not a variation on something in the training data
The novel benchmarks are OpenAI’s attempt to answer exactly that. Good. That’s the evaluation that matters, and it’s the one worth pressing them on as independent researchers get their hands on the model.
The cybersecurity capability is the real story
Advanced cybersecurity capability is not a feature. It’s a dual-use problem with a friendly product name. A model that can reason about vulnerabilities well enough to need a purpose-built exploit benchmark is a model that can reason about vulnerabilities. That skill does not check who’s asking.
The reporting around the launch reflects this. Coverage described rising scrutiny and safety concerns circulating on social media, which is roughly what you’d expect when a lab ships stronger security reasoning and simultaneously suggests the AGI era might be starting. Those two claims amplify each other in ways that make people nervous, and the nervousness is proportional.
I’m not going to tell you Astra is dangerous. I haven’t tested it, and I don’t review models I haven’t used. What I’ll say is that the safety conversation around this release deserves more oxygen than the benchmark numbers, and it’s getting less.
On the AGI claim
OpenAI thinks Astra may kick off the AGI era. I think “AGI era” is a phrase that does more work for investor decks than for engineers. There’s no agreed definition, no agreed test, and no way to falsify the claim. That makes it unreviewable, which makes it useless to you.
Here’s a more practical frame. Astra was announced in a week that also put Sam Altman and Dario Amodei on a stage with India’s Prime Minister at the AI Impact Summit in New Delhi. The launch is a policy event as much as a product event. When rival lab CEOs share a photo op with heads of state, the models are no longer competing only on capability. They’re competing for regulatory position.
My verdict, such as it is
Astra looks like a serious model with serious evaluation work behind it, wrapped in a claim I can’t test and a level of hype the audience has grown numb to. The methodology is the highlight. The AGI talk is noise. The security capability is the thing to watch.
Model fatigue isn’t cynicism. It’s the market asking a fair question: what can this do for me that the last one couldn’t? Until independent testing answers that, treat the near-perfect scores as what they are. A measurement of tests that have run out of room.
🕒 Published: