What good is a model that solves Navier-Stokes if you can’t open a terminal and ask it anything?
That’s the question I keep circling back to, and nobody at OpenAI seems interested in answering it. Here’s what we know. On August 1, 2026, OpenAI published a post saying an internal version of its next major model — Astra — produced new results on ten long-standing open problems in mathematics and theoretical computer science. By September 21, the claim had grown to more than 100 resolved problems spanning most areas of mathematics. The Navier-Stokes existence and smoothness problem, one of the Millennium Prize problems, was announced on September 9, reportedly after roughly 88 hours of compute with up to 10,000 coordinating agents.
It’s now October 8. There is no release date. No pricing. No availability details. Nothing.
Reviewing a product that doesn’t exist
This site reviews AI tools. My job is to tell you whether something is worth your money and your workflow. I can’t do that here, because there is no product. There is a series of announcements about a thing that lives inside OpenAI’s own infrastructure, described by OpenAI, verified on OpenAI’s timeline.
Two months passed between “ten problems” and “more than 100 problems.” Twenty-four days of that window produced the bulk of the jump, according to the reporting. That is either the steepest capability curve in the history of computing, or it’s a number that got easier to produce as the definition of “resolved” got looser. I genuinely don’t know which, and that’s the problem. Neither do you.
The coverage was everywhere. CNBC, BBC, ABC News — the ABC segment alone pulled over a million views and 16,000 likes. Scientific American covered it on September 8, a day before the Navier-Stokes announcement, and the framing in their own headline included the phrase “amid swirl of controversy.” When the science press is hedging in the headline, something about the claim isn’t settling cleanly.
The verification question nobody wants to sit with
One detail in the source material is more interesting than the headline numbers. Following an investigation, it was confirmed that mathematician Buckmaster’s Codex prompts over the two months preceding the announcement could not have influenced the system. Read that again. Someone thought it was worth investigating whether the model had been fed the answer through user prompts. That investigation cleared it, which is good. But the fact that it was necessary tells you exactly how much daylight exists between “an AI company says it solved this” and “the mathematical community agrees this is solved.”
Mathematical proof has a verification culture that predates computing by centuries. A proof is not solved when its author says so. It’s solved when other people who understand the domain read it, attack it, and fail to break it. That process takes months or years for a result of this magnitude. OpenAI’s announcement cadence is measured in days.
What I’d actually need to score this
If Astra ships, here’s what I’d be looking for before putting a number on it:
- Compute cost per result. 88 hours and up to 10,000 coordinating agents is not a consumer workload. It’s not an enterprise workload either. If that’s the floor for a hard problem, the pricing conversation gets strange fast.
- Behavior on problems that aren’t famous. Open problems in mathematics are heavily documented, discussed, and partially attacked in public literature. A model that performs on well-known targets may behave very differently on a question nobody has written about.
- Independent verification rate. Of the 100-plus claimed results, how many have been reviewed and accepted by people outside OpenAI? That number, whenever it arrives, is the only score that matters.
- Whether any of this reaches the API. Internal models that stay internal are research, not tools.
My read
I’m not calling this fake. The reporting is real, the outlets are credible, and the investigation into prompt contamination came back clean. If even a fraction of the claimed results hold up under review, it’s the most significant thing to happen in automated reasoning.
But I review things people can use, and the gap between the announcement and the availability is now more than two months wide with no bridge in sight. What OpenAI has shipped is a narrative. The math may well be solid. The product is vapor.
Until a version of this lands somewhere I can actually test it, treat the 100-problem figure as a company’s claim about its own unreleased software. That’s not cynicism. That’s the standard we apply to every other tool on this site, and a Millennium Prize problem doesn’t earn an exemption.
🕒 Published: