Ten thousand concurrent AI agents. That’s the number OpenAI put behind its September 8, 2026 announcement that an internal model had produced a proof resolving the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize Problems. Not one model reasoning carefully. Ten thousand of them, running at once, pointed at a question that has outlasted every mathematician who has ever attacked it.
My first reaction was not awe. It was a very specific kind of suspicion that I’ve developed reviewing AI tools for a living: when a company reaches for a number that large, the number is usually doing rhetorical work. Ten thousand agents is a compute flex. It tells you about the budget, not the proof.
The announcement that came with an asterisk
Here is what I find more interesting than the claim itself. On September 21, less than two weeks later, an independent Advisory Group on Mathematics and AI was formed to advise on the release of these results. Think about the sequence. Announce first, assemble the people qualified to evaluate the announcement second.
That ordering tells you something about how AI labs now operate in scientific domains. The press release is the product. Verification is a follow-up release. In software, we’d call that shipping to production and writing the tests later. In mathematics, where the entire value of a result rests on whether the proof actually holds, it’s a stranger choice.
To be fair to the Advisory Group, its existence is the most encouraging detail in this whole episode. Somebody at the table understood that a Millennium Prize claim cannot be self-certified by the company making it. Whether the group has real authority over what gets released, or whether it functions as reputational cover, is the question I’d want answered before I believed anything.
Twenty-five Fields medalists are not impressed
The mathematical community did not clap politely. Twenty-five Fields Medal winners signed a declaration arguing that the push by AI companies to solve famous problems as a benchmark actively harms the science of mathematics.
That’s worth sitting with, because it is not a Luddite complaint. These are not people who struggle with abstraction or fear new tools. Their objection is about incentives. Famous problems make terrible benchmarks precisely because they are famous. They are legible to journalists, to investors, to anyone who has heard the phrase “Millennium Prize.” They are optimized for headlines, not for the health of a field that advances through thousands of unglamorous results nobody writes about.
When you turn a discipline’s most celebrated open questions into a leaderboard, you don’t accelerate the discipline. You accelerate the part of the discipline that produces announcements. I’ve watched the same distortion play out in coding benchmarks, where models got very good at the specific puzzle format and no better at the actual work.
The unglamorous progress is the real story
Strip away the Navier-Stokes drama and something genuinely useful is happening. Quanta Magazine has documented AI models solving research-level questions across multiple areas of mathematics. In February, a challenge called First Proof gave entrants one week to have their models solve ten research-level problems. That’s not a single trophy result. That’s a format designed to measure capability across a spread of questions, with a clock, which is roughly how you’d evaluate any tool you planned to actually depend on.
Renaissance Philanthropy’s second round of funding points the same direction. Twenty-two grant awards in 2026 for AI for mathematics and theoretical computer science, balanced across moonshots, field-building, benchmarks and datasets to track progress, and infrastructure. Read that list again. Benchmarks and datasets. Infrastructure. These are the boring categories that determine whether a field can tell the difference between real progress and a good demo.
Nobody is going to make a video about grant number seventeen. But that portfolio approach is what a serious effort looks like, and it stands in sharp contrast to pointing ten thousand agents at the most quotable problem on the board.
What I’d actually watch
If you’re trying to track whether AI is becoming useful for mathematics rather than useful for math-adjacent marketing, here’s where I’d look:
- Whether the Advisory Group publishes anything that contradicts the company that convened it
- Whether the Navier-Stokes proof survives independent review, and how long that takes
- Whether benchmark work funded by groups like Renaissance Philanthropy gets adopted by labs, or quietly ignored in favor of trophy problems
- Whether working mathematicians, not CEOs, start describing these tools as part of their daily process
That last one matters most to me. The honest measure of a tool is not what it solved once under maximum compute with a press team standing by. It’s whether people who do the work reach for it on a Tuesday afternoon with no audience watching.
Twenty-five of the most decorated mathematicians alive just told us the current scoreboard is measuring the wrong thing. Before we accept that a famous problem has fallen, it’s reasonable to ask what the scoreboard was built to prove, and who built it.
🕒 Published: