\n\n\n\n 722 Proofs and Not One Model You Can Actually Use - AgntHQ \n

722 Proofs and Not One Model You Can Actually Use

📖 4 min read•788 words•Updated Oct 7, 2026

Ten thousand concurrent AI agents. That’s the number OpenAI put on the compute behind its September 8, 2026 announcement that an internal model had produced a proof resolving the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize Problems. Ten thousand agents, one of the hardest open questions in mathematics, and a model the public cannot name, test, or query.

I review AI tools for a living. My job is to put a thing in front of readers and tell them whether it works. So let me be blunt about where this leaves me: there is nothing here to review. There is an announcement, a claim, and a repository.

What actually shipped

In October 2026, OpenAI published 722 mathematical manuscripts on GitHub, crediting the work to an unnamed internal frontier model it has not released. The company also claimed the same model resolved hundreds of additional problems, including a solution to the four-dimensional Kakeya conjecture and improvements on some of the world’s most important computer algorithms.

Read that list again and notice what’s missing. Not the results — the results are there, in volume. What’s missing is the artifact. No model card. No weights. No API endpoint. No name. The manuscripts went onto GitHub rather than through a press conference, which some people read as humility. I read it as a company that wants the credit for the output without the scrutiny that comes with shipping the thing that produced it.

This is the pattern I keep running into across the AI tools space, and math is just where it gets the most absurd. A vendor demonstrates a capability, declines to release the system, and the demonstration itself becomes the product. You’re not being sold software. You’re being sold a reputation.

The verification problem nobody wants to own

Mathematics has an unusually honest quality control system. A proof is correct or it isn’t, and the people qualified to check are the people who spent decades on the problem. That process is slow by design. It is also the only thing that converts a claim into knowledge.

Dropping 722 manuscripts at once does something strange to that system. It transfers an enormous verification burden onto a small community of specialists who did not ask for it and are not paid to absorb it. If even a fraction of those documents are substantive, checking them is years of unfunded labor. If a fraction are subtly wrong, finding out which ones is also years of unfunded labor. Volume isn’t a gift when the recipient has to do the work.

And the pushback was immediate. Twenty-five Fields Medal winners argued that the push by AI companies to solve famous problems as a benchmark actively harms the science of mathematics and the mathematical community. That is not a fringe reaction. That is roughly the top of the discipline saying the incentive structure is broken.

Why famous problems are a terrible benchmark

Here is what I think the Fields medalists are getting at, and it maps onto something I see constantly in AI evaluation.

Famous problems make great headlines precisely because they are legible to outsiders. Navier-Stokes is a Millennium Prize Problem; everyone understands that a prize is a big deal. So when a lab wants a demonstration that lands, it reaches for the trophy, not the useful thing.

  • Trophy problems measure one narrow capability and get mistaken for general ability
  • Optimizing for them pulls research attention toward whatever is marketable rather than whatever is unresolved and important
  • They reward announcement speed over verification, because the news cycle closes before the referees do
  • They let a vendor claim a win without ever handing you a tool you can apply to your own problems

The parallel in my world is benchmark scores. A model tops a leaderboard, the press release writes itself, and then you use the thing for a week and discover it cannot hold a thought across three files. Trophy math is the same trick with higher stakes and a more defenseless audience.

What I’d want before calling this real

I’m not arguing AI has nothing to contribute to mathematics. The opposite, probably. A system that can grind through proof search at scale is genuinely useful, and some of those 722 manuscripts may hold up beautifully.

But “useful” and “credible” both require access. Release the model, or at minimum name it and let independent researchers probe it. Submit the significant results through normal channels and accept the timeline. Fund the verification work instead of outsourcing it to people whose field you just disrupted for a news cycle.

Until then, my review stands at the only honest rating available. Unverifiable. Ten thousand agents is an impressive number. It is not a result, and a repository is not a product.

đź•’ Published:

📊
Written by Jake Chen

AI technology analyst covering agent platforms since 2021. Tested 40+ agent frameworks. Regular contributor to AI industry publications.

Learn more →
Browse Topics: Advanced AI Agents | Advanced Techniques | AI Agent Basics | AI Agent Tools | AI Agent Tutorials
Scroll to Top