What happens when the most important math release of the decade shows up as a GitHub commit?
That’s not a hypothetical. On October 6, 2026, OpenAI published 722 mathematical manuscripts to GitHub, crediting them to an internal frontier model it has not named and will not release. No press conference. No launch event. Just a repository full of result families, proof artifacts, Lean formalizations, and ten abridged summaries of how the model reasoned its way there.
I review AI tools for a living. My job is to tell you whether something is worth your time, your money, or your trust. This one breaks the format, because there is no tool here. There is only output.
What’s actually in the drop
The release contains 372 result families across the 722 manuscripts, plus supporting proof artifacts and Lean formalizations. That last part matters more than the headline count. Lean is a proof assistant, which means a machine can check the work. You don’t have to take OpenAI’s word that the proofs hold. You can run them.
That is a genuinely unusual move for a company that has spent years asking people to trust benchmark charts. Formal verification is the one form of evidence that doesn’t care who produced it or how good their marketing team is. A Lean proof either compiles or it doesn’t.
OpenAI also made a point about cost, emphasizing that these results came without the multimillion-dollar price tag previously associated with similar breakthroughs. Make of that what you will. Companies rarely volunteer cost numbers unless the comparison flatters them.
The timeline is the real story
Walk it back a few weeks. Around September 9, 2026, OpenAI announced its system had solved the Navier-Stokes existence and smoothness problem after roughly 88 hours of compute, using up to 10,000 coordinating AI agents. Then on September 21, the same model was credited with resolving more than 100 long-standing open problems across most areas of mathematics, in roughly 24 days of training. Then October 6 brought the 722 manuscripts.
Three announcements in under a month, each one larger than the last. Scientific American described a field already in shock. That phrasing is doing a lot of work, and I think it’s accurate. Mathematics operates on a timescale measured in careers. A single open problem can define a researcher’s entire body of work. Compressing that into 88-hour compute runs is not an incremental improvement to a workflow. It’s a different relationship between humans and the subject.
Why the unnamed model bothers me
Here is my actual complaint, and it has nothing to do with whether the math is correct.
You cannot use this. You cannot test it. You cannot probe its limits, check whether it fails gracefully, or find out what it does on problems outside the curated set. OpenAI says it has been consulting with an independent Advisory Group on Mathematics and Artificial Intelligence, which is a reasonable governance step. But an advisory group is not the same as public access.
The Lean formalizations partially solve this. Verified proofs are verified proofs, regardless of their source. Still, there’s a difference between verifying 722 outputs and understanding a system. The first tells you what happened once. The second tells you what to expect next time.
For a site that reviews AI tools without the marketing gloss, this creates an awkward situation. I can evaluate the artifacts. And the agent is the thing that matters for anyone trying to plan around what’s coming.
What I’d watch next
- Independent verification throughput. How fast can mathematicians actually check 372 result families? The Lean proofs help, but someone still has to confirm the formalizations match the stated theorems.
- Which problems got solved. Breadth across most areas of mathematics is a strong claim. The distribution of difficulty inside those 722 manuscripts will tell you whether this is depth or volume.
- Whether access ever opens up. An internal model producing public results is a specific choice. If that stays the pattern, the research community gets outputs it can check but not a system it can question.
- The 10,000-agent architecture. Coordinating that many agents on a single problem is an engineering story as much as a math one. The ten abridged reasoning summaries are the only window into how it works.
My honest read
I’m not going to pretend skepticism I don’t have. Formal proofs are checkable, and OpenAI shipped them. That’s the strongest possible way to make this kind of claim, and the company chose it over a staged demo. Credit where it’s due.
But I’d trade a hundred manuscripts for one afternoon with the model. Results you can verify are valuable. A system you can interrogate is how a field actually moves. Right now mathematicians have the first and not the second, and that asymmetry is the part worth arguing about.
🕒 Published: