OpenAI is calling GPT-6 Astra its most capable and aligned model so far, with what it describes as major gains in computer use, coding, and scientific reasoning. Bold words. And as Fortune’s Emily Forlini reported when the model launched on September 3, 2026, the company is particularly proud of one specific ability: Astra can use your computer.
Read that again. Not “assist you while you use your computer.” Use it. Click things. Run things. Do the work. That framing tells you exactly where OpenAI thinks the money is, and it’s not in chatbots anymore.
What OpenAI Is Actually Claiming
Strip away the launch-day glow and the pitch comes down to three claims:
- Capability. OpenAI positions Astra as its frontier model, with the strongest gains in computer use, coding, and scientific reasoning.
- Efficiency. The company says Astra was trained to complete tasks in fewer tokens with fewer retries, framing it as more useful work per dollar for customers.
- Alignment. OpenAI calls it the most aligned model it has shipped, which is the kind of claim that’s impossible to verify from the outside and therefore worth exactly what you paid to read it.
The efficiency angle deserves more attention than it’s getting. “Fewer tokens, fewer retries” is not a sexy headline, but it’s the metric that actually decides whether businesses deploy this thing at scale or quietly shelve it after the pilot. An agent that finishes a task on the first attempt is a tool. An agent that needs three retries is a bill.
The Benchmark Asterisk
Here’s a detail I respect, grudgingly. OpenAI acknowledged concerns that exposure to historical software vulnerabilities may have affected benchmark results, and responded by evaluating Astra on novel benchmarks, including an internal one it built called ExploitBench.
Translation: they know the model may have memorized old security flaws from training data, which would inflate its scores on established tests, so they built fresh exams it couldn’t have studied for. That’s the right instinct. It’s also a reminder of how much benchmark theater exists across this industry. When a lab has to construct new tests because the old ones might be contaminated, you should apply that same skepticism to every leaderboard screenshot you see this year.
The catch, of course, is that ExploitBench is internal. An exam written, administered, and graded by the student’s own family is better than no exam, but I’d like to see independent evaluators get their hands on this before anyone updates their org chart.
The Hype Machine Is Already Confused
Want a snapshot of how noisy this launch is? Within days, YouTube channels were pushing videos with titles like “The Biggest AI Revolution | Dangerous | IT Job Ends?” and pulling tens of thousands of views. Meanwhile, some aggregated posts floating around managed to attribute the model to a team of inventors at Amazon, which is flatly wrong — this is OpenAI’s release, per OpenAI itself and every credible outlet covering it.
When the hype cycle spins fast enough that people can’t even agree on who built the thing, that’s your cue to slow down and read primary sources. The “IT jobs are over” genre of content appears within hours of every major model release, and so far the IT jobs have stubbornly continued to exist. That doesn’t mean Astra changes nothing. It means the people shouting loudest have the least information.
What I’d Watch For
The computer-use pitch is where Astra either earns its keep or becomes an expensive demo. If it genuinely completes professional tasks end to end, with fewer retries and lower cost as OpenA
đź•’ Published: