Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Scott AaronsonOpenAI mathAnthropic math breakthroughs

Are AI Labs Hiding Solved Math Problems? Scott Aaronson's Claims Explained

Scott Aaronson says OpenAI and Anthropic may be sitting on unpublished math breakthroughs after backlash over a Navier-Stokes proof claim.

Edited by Luis Chavez-Mattos, Director of Product RSS
Are AI Labs Hiding Solved Math Problems? Scott Aaronson's Claims Explained

What is the claim about AI labs hiding math breakthroughs?

Theoretical computer scientist Scott Aaronson wrote a blog post arguing that OpenAI and Anthropic are likely withholding solutions to significant open problems in theoretical computer science, based on rumors he says he has heard. His reasoning connects to a real, documented controversy: an OpenAI model was reported to have produced a solution related to the Navier-Stokes equations, one of the Clay Mathematics Institute’s Millennium Prize Problems. That claim triggered immediate pushback, with critics arguing the model had repackaged existing human work rather than generated anything new. Aaronson’s point is that labs, having been burned once, may now be sitting on further results rather than risk another round of credit disputes and public skepticism.

TL;DR

  • Scott Aaronson, a theoretical computer scientist who has worked with OpenAI and was involved in verifying Google’s quantum supremacy claims, is the source of the claim that frontier labs may be holding back solved math problems.
  • The rumor centers on OpenAI and Anthropic specifically, described as the two labs most likely running dedicated, well-funded efforts aimed at cracking longstanding open problems.
  • The trigger for this caution was the backlash following a reported Navier-Stokes Millennium Prize Problem result, where outlets like Scientific American and the Guardian framed the achievement as controversial, credit-stealing, or overstated rather than as a genuine breakthrough.
  • Aaronson argues the objection “it just stole human work” doesn’t hold up logically, because if that were the whole story, the same models shouldn’t be able to generate novel solutions to other unsolved problems, yet the same pattern reportedly keeps happening.
  • He frames the real issue as backlash management, not safety: nobody is claiming these results are dangerous, the problem is reputational and institutional, involving how mathematicians get credit and how the field processes AI-generated proofs.
  • One publicly visible example he cites is an Anthropic employee’s disproof of a mathematical conjecture shared informally on social media, which he uses as evidence that AI-assisted math results are already circulating outside formal publication channels.
  • Aaronson’s broader argument is that the AI research community keeps setting goalposts (“wake me up when AI solves a Millennium Problem”) and then moving them the moment those goalposts are reached.
Cursor
ChatGPT
Figma
Linear
GitHub
Vercel
Supabase
goremy.ai

Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

Why would a lab withhold a solved math problem?

The logic, as Aaronson lays it out, isn’t about secrecy for competitive advantage in the traditional sense. It’s about avoiding another credibility fight. When the Navier-Stokes-related claim surfaced, the reaction wasn’t celebration. It was accusations that the model had lifted a human mathematician’s unpublished or overlooked work and presented it as a novel discovery. Major outlets ran headlines questioning the legitimacy of the result and emphasizing human contribution over the AI’s role.

That kind of reception creates a strong incentive for labs to slow down before publishing the next one. If a result invites accusations of plagiarism, or triggers disputes among mathematicians about who deserves authorship credit, a lab has to decide how to release it responsibly, meaning with enough context, verification, and attribution that it doesn’t blow up into another controversy. Aaronson’s framing is that this isn’t a safety problem at all. Nobody is arguing these proofs are dangerous. It’s a social and institutional problem: how do you credit a discovery when the entity that found it isn’t a person?

What problems are these labs supposedly working on?

According to Aaronson, the rumors he has heard aren’t about the biggest, most famous open questions like P vs NP or other complexity class separations. Those remain unsolved. Instead, the rumored results involve other longstanding open problems in theoretical computer science and mathematics that have resisted human effort for decades. He suggests both OpenAI and Anthropic have dedicated internal efforts, with real compute budgets, aimed specifically at these kinds of problems, and that the publicized Navier-Stokes-adjacent result was likely not an isolated event but one piece of a larger pattern.

He also points to a concrete, publicly visible data point: an Anthropic employee reportedly shared a disproof of the Jacobian conjecture on social media, generated with the help of an AI model referred to informally as “Claude Fable.” That single tweet, in Aaronson’s telling, is a signal that AI-assisted mathematical results are already leaking into public view faster than the formal academic system can process them.

Why is the media response to this so skeptical?

Aaronson highlights a pattern across major publications. Scientific American reportedly framed a math result as technically eligible for a $1 million prize but only through what it called a “controversial loophole.” The Guardian ran a piece arguing that humans remain essential to mathematics and that tech companies were reluctant to acknowledge that. In both cases, the coverage leans toward minimizing or discrediting the achievement rather than treating it as evidence of a real capability jump.

Aaronson’s broader argument is that this reflects a recurring pattern in how AI progress gets received. Every time a system clears a bar that skeptics previously set as the line for “real” intelligence or capability, from passing the Turing test to writing working code to now producing novel mathematical proofs, the response isn’t to update the underlying belief. It’s to move the bar. He compares this to describing today’s AI capabilities to a computer scientist from twenty years ago: even the most conservative skeptics of that era would have said flatly that a system solving Millennium Prize-level problems would be undeniable proof of a major breakthrough. Now that something resembling that has arguably happened, the reaction has been to argue about credit and legitimacy instead.

Is there a legitimate concern buried in the backlash?

Yes, and Aaronson doesn’t dismiss all of it. Some of the concerns raised by mathematicians are substantive: how do you verify a machine-generated proof rigorously? How do you assign authorship when a model surfaces a solution based on a human’s suggested framing or unpublished partial work? Should a paper list a model as a co-author, and if so, how does peer review even function in that case? These are real, unresolved questions about process, credit, and verification standards in a field built entirely around human-to-human attribution.

Where Aaronson pushes back is on the instinct to use those procedural concerns as a reason to deny the underlying capability. A messy authorship question is not evidence that a proof is fake. Conflating the two, he argues, is how the field ends up in a “shell game” where every new capability gets waved away with a technicality, only for the goalposts to quietly move again once the next result lands.

Frequently Asked Questions

Who is Scott Aaronson and why does his claim carry weight?

Aaronson is a theoretical computer scientist who has worked on AI alignment research at OpenAI and played a significant role in verifying Google’s quantum computational supremacy claims. His background in both AI and rigorous mathematical verification gives him credibility to speak on both what these models are technically doing and how the math community is likely to react.

Did an AI actually solve a Millennium Prize Problem?

A reported result connected to the Navier-Stokes equations, one of the seven Millennium Prize Problems, drew major controversy. Critics argued the achievement relied on prior human work rather than being fully original, and outlets covered it skeptically rather than as a confirmed solved problem. It remains disputed rather than formally settled.

Why would solving a math problem cause backlash instead of celebration?

Because credit and originality are central to how mathematics as a field operates. If a result appears to lean heavily on unpublished or overlooked human work, mathematicians raise concerns about proper attribution. Add an AI system into the mix and questions about authorship, verification, and who “discovered” the result become even more contested.

Are OpenAI and Anthropic confirmed to be withholding results?

No. This is based on rumors Aaronson says he has heard, not on any official statement from either company. The claim should be treated as reported speculation from a credible but not neutral source, not as confirmed fact.

No. Aaronson is explicit that nobody involved is arguing these math results are dangerous. The hesitation to publish, if real, is about avoiding reputational blowback and navigating credit disputes within the mathematics community, not about withholding something hazardous.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.