722. That’s how many mathematical manuscripts OpenAI dumped on GitHub on October 6, 2026, all credited to an internal frontier model that has no name and no release date. Not a paper. Not a conference. A repository.
I review AI tools for a living. My entire job is to get hands on a thing, push it until it breaks, and tell you whether it earns its price tag. So let me be upfront about the problem with this story: there is no tool here to review. There’s an output pile and a promise.
What actually happened
The October drop included results touching Hilbert’s 10th problem and the Kakeya conjecture. It followed a September 8, 2026 announcement that an internal model had produced a proof resolving the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize Problems. OpenAI said it is committed to responsibly releasing the model behind the work, while confirming in the same breath that the model stays unnamed and unreleased.
Read that pairing twice. The commitment and the withholding arrived together, in the same communication. That’s not a roadmap. That’s a posture.
Scientific American’s coverage described a field already in shock before this latest batch landed. Reaction inside the mathematical community has split, with some people crediting real advances and others worried about what this does to how research gets done. Both camps are reasonable, and I’d argue both are responding to the same underlying gap.
The review problem nobody solved
Mathematics has a verification culture that predates computers by centuries. A proof is not a result until other humans have chewed on it. That process is slow on purpose. It’s slow because being wrong in public is expensive and because subtle errors hide in long arguments.
Now consider the arithmetic of 722 manuscripts arriving at once. Even if every single one is correct, the checking labor involved is enormous and falls on people who did not ask for it and are not paid for it. The cost of the claim got pushed onto the field. That asymmetry is the actual story, more than any individual theorem.
And here’s where my professional irritation kicks in. When a company ships a coding assistant, I can benchmark it. I can run the same prompts across competitors, count failures, check whether the marketing matches reality. With this, the evaluation surface is:
- Static text files, with no ability to probe the system that generated them
- No way to test whether the model produces garbage on adjacent problems
- No visibility into how many attempts, hints, or human interventions went into each result
- No independent party able to reproduce the pipeline end to end
Those last two matter more than people realize. The difference between a model that proves things and a model that occasionally proves things under heavy human steering is the entire difference. From outside, those two scenarios look identical.
Why GitHub instead of a stage
I keep coming back to the delivery channel. A repository drop is a legitimate choice for math, since that’s where collaborative work and code already live. It’s also the lowest-ceremony way to put 722 claims into the world without standing behind any one of them in front of a microphone. I’m not accusing anyone of cowardice. I’m pointing out that the format shifts narrative control to whoever reads first and shouts loudest, which in practice means Reddit threads and press coverage set the framing before experts finish page one.
What would change my mind
Access. Not a demo, not a curated showcase. Give working mathematicians the ability to pose their own problems to the system and publish whatever comes back, including the failures. Document the human involvement per result. Let someone outside the building measure the error rate.
Until then, the honest review is a non-review. I can tell you 722 manuscripts exist. I can tell you they cover serious problems that serious people have worked on for decades.
My actual take
If even a fraction of these results hold, it’s a genuine milestone and the people who built it deserve credit. That’s not in dispute. What I’m pushing back on is the structure of the announcement, where the most important artifact, the model, stays locked up while its claimed output gets distributed for free.
Science runs on the ability to check things. A release that gives you conclusions but withholds the instrument is an interesting publication and a weak contribution to the record. Mathematicians are going to spend months sorting the signal from whatever else is in that repository, and the company that made the work is not on the hook for any of that labor.
Impressive? Probably. Reviewable? Not yet. I’ll update when someone hands me a keyboard.
🕒 Published: