OpenAI's 722 Math Proofs: What Its Secret Model Actually Found
OpenAI released 722 math manuscripts from an unreleased model touching the Riemann hypothesis and Navier-Stokes. Here's what it actually shows.

What did OpenAI actually release?
OpenAI published 722 mathematical manuscripts produced by an unreleased internal model, posting proofs, conjectures, reasoning traces, and Lean verification files to a public GitHub repository (github.com/openai/math). The drop spans 17 subject areas, from prime number theory to plasma physics to quantum circuits. None of it solves a Millennium Prize problem outright, but several results chip away at the edges of some of the hardest open questions in math, including work related to the Riemann hypothesis, the Navier-Stokes equations, the Birch and Swinnerton-Dyer conjecture, and the Goldfeld conjecture.
TL;DR
- OpenAI’s release follows a sharp acceleration curve: roughly 10 results in August, 100-plus in September, and 722 by early October, a pace that looks close to 10x growth per month.
- The model produced a partial Riemann hypothesis result, nicknamed the quasi Riemann hypothesis, pushing a key boundary to 0.875 compared to Riemann’s target of 0.5 and the prior human-verified position at 1.0.
- On the Birch and Swinnerton-Dyer conjecture, the model reportedly established a complete formula for a broad class of cases, an area where the last major breakthroughs date back decades.
- Each result was generated using roughly three hours of ChatGPT Pro-level thinking compute, a tiny fraction of the time human mathematicians typically spend on comparable problems.
- None of the 722 manuscripts has been independently verified by the mathematics community yet, and that verification bottleneck is now the central open question.
- The releases include machine-readable Lean proofs, which makes some of the verification process faster than traditional peer review, though not instant.
- The bigger story isn’t any single proof, it’s that math is becoming a testing ground for what happens when AI models hit the frontier of a checkable discipline, with other fields like biology and physics likely to follow.
How fast is this actually moving?
The pace is the part worth paying attention to. In August, OpenAI put out around 10 results. In September, that jumped to over 100, including a Navier-Stokes result. By October 6, the count hit 722. That’s roughly a 10x increase month over month, compressed into about two months of visible output.
Context matters here: Navier-Stokes describes fluid motion and is one of the seven original Millennium Prize problems, each carrying a $1 million prize for a full solution. OpenAI has not claimed the prize, and a full solution wasn’t reached. But the pattern across the 722 manuscripts is consistent: AI-generated work narrowing in on problems that have resisted full human solutions for decades, sometimes over a century.
What is the “quasi Riemann hypothesis” result?
The Riemann hypothesis concerns the distribution of prime numbers, specifically a conjecture about where the zeros of the Riemann zeta function lie. Mathematicians believe that a hidden pattern governs prime number behavior, and Riemann’s conjecture targets a precise line, the 0.5 line, where all meaningful zeros are predicted to sit.
A useful way to picture progress on problems like this: imagine a mountain whose peak is hidden in cloud. Nobody knows how tall it really is. Climbing partway up still tells you the mountain is “at least this tall,” even if you can’t see the summit. On this framing, human mathematicians had climbed to a boundary around 1.0. The internal OpenAI model reportedly pushed that boundary down to 0.875, closer to Riemann’s 0.5 target than anyone had gotten before. That’s not a solved Riemann hypothesis. It’s a partial result, sometimes called the quasi Riemann hypothesis, that narrows the space where a full proof would need to live.
What about Birch and Swinnerton-Dyer and Goldfeld?
The Birch and Swinnerton-Dyer conjecture is another Millennium Prize problem, also carrying a $1 million reward for a complete proof. The model reportedly established a complete formula for a broad class of cases within the conjecture, which, if it holds up, would be a meaningful advance in a field where the last major progress happened in the 1980s.
On the Goldfeld conjecture, which predicts that a certain family of equations should split roughly 50/50 between having finite and infinite solutions, the model reportedly supplied additional cases that move the evidence closer to that predicted split.
In each case, the pattern is the same: not a full solution, but measurable, specific progress on problems that have stumped specialists for a long time, generated by a general-purpose language model rather than a system built specifically for mathematics.
Why is verification the real bottleneck?
Dropping 722 manuscripts in a single release creates an immediate and obvious problem: who checks it? The pool of mathematicians qualified to review frontier work in any one of these 17 subject areas is small, often numbering in the hundreds worldwide for a given subfield, and reviewing a single dense paper can take a week or more of a specialist’s time.
Some of the work includes Lean-verified proofs, meaning the logical steps can be checked by a proof assistant rather than relying purely on human review. That speeds things up, but it doesn’t eliminate the need for mathematicians to confirm that a formalized proof actually captures a meaningful, correctly stated result. The scale of 722 manuscripts arriving at once makes this a genuine bottleneck, not a formality.
The stakes of getting verification right are high. If the overwhelming majority of the results hold up, this becomes one of the largest single contributions to mathematical knowledge in a short window, the kind of output that eventually reshapes how topics are taught. If a meaningful share turns out to be flawed, it still has value: it shows researchers where current AI reasoning breaks down, which feeds directly back into model training.
Did AI steal credit from mathematicians?
This question came up after OpenAI’s September Navier-Stokes related release, when some coverage suggested the AI had appropriated a human mathematician’s existing work rather than generating it independently. That framing did not hold up. With 722 manuscripts now in public view, the same scrutiny is likely to repeat, and it’s a reasonable instinct: extraordinary claims about AI mathematical output deserve scrutiny over provenance, not just correctness.
Why is OpenAI doing this?
Part of the explanation is pragmatic. Frontier AI labs are running short of benchmarks hard enough to meaningfully differentiate model capability. Unsolved math problems, particularly ones tied to Millennium Prizes, are naturally difficult, well-defined, and checkable, which makes them attractive as a proving ground. Math also has a property few other fields share: a proof is either valid or it isn’t, especially once formalized in a language like Lean. That makes it a cleaner test of whether a model’s reasoning is actually sound rather than merely fluent.
That checkability is also why some observers expect math to be the first domain where this kind of AI-driven flood of results shows up, with biology, physics, and economics likely to follow once equivalent verification tools mature in those fields.
Will any of this matter outside of math departments?
The honest answer is: not clearly yet, and not immediately. None of the Millennium Prize problems were fully solved in this release. The 722 manuscripts represent narrowing progress around hard problems rather than resolution of them, though the pattern of incremental results closing in from multiple directions suggests some of these problems may fall sooner than expected.
Where the impact is most plausible in the near term is inside AI research itself. Mathematical breakthroughs tend to feed back into the tools used to build and train models, including the underlying linear algebra and optimization techniques that make large models work. Projects like Google’s AlphaEvolve and research from Sakana AI have already shown that AI-discovered mathematical and algorithmic improvements can be fed back into making better models. It would be more surprising if these math results had no effect on model training than if they did.
Frequently Asked Questions
Did OpenAI’s model solve the Riemann hypothesis?
No. It produced a partial result, sometimes called the quasi Riemann hypothesis, that narrows the boundary closer to Riemann’s original target without proving the full conjecture.
Did OpenAI solve Navier-Stokes and claim the Millennium Prize?
No. OpenAI published a Navier-Stokes related result in September but has not claimed the Millennium Prize, and the full problem remains unsolved.
Remy is new. The platform isn't.
Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.
How much compute did each result take?
Each of the 722 manuscripts used roughly the equivalent of three hours of ChatGPT Pro-level thinking compute, according to OpenAI’s release.
Has the math community verified these 722 manuscripts?
Not yet. Verification is ongoing and is expected to take significant time given the small number of specialists qualified to review work in each of the 17 subject areas covered.
Why does this matter beyond mathematics?
Math is one of the few fields where results are fully checkable, which makes it a likely first domain to show what happens when AI models reach the frontier of a field. Similar effects are expected eventually in biology, physics, and other research-heavy disciplines.

