Who Gets Credit When AI Solves a Math Problem Nobody Could?
AI models are solving open math problems, sparking fights over attribution, authorship, and whether humans still need to understand the proofs.

What’s actually being disputed
When an AI model produces a solution to a longstanding open problem, nobody agrees on what happened next. Did the model solve it, or did it recombine a human’s unpublished insight? Does the person who prompted the model deserve authorship? Should the model itself be listed as a co-author? These aren’t hypothetical questions anymore. They came to a head after a widely discussed case involving an OpenAI model and a proof connected to the Navier-Stokes problem, one of the Clay Millennium Prize problems, which set off a public fight over whether the result counted as a genuine breakthrough or a repackaged version of existing human work. Theoretical computer scientist Scott Aaronson wrote about this dynamic in a blog post, and his account (along with the broader online reaction) shows the fight has moved past one paper and into a full argument about how mathematics as a field handles machine-generated results.
TL;DR
- Attribution is contested because it’s often unclear whether an AI model generated a genuinely new proof or restated a solution that existed somewhere in unpublished human work.
- AI labs are reportedly holding back results, according to Aaronson, sitting on solutions to significant open problems in theoretical computer science because the backlash to earlier claims made them wary of how to release new math responsibly.
- Media coverage has skewed toward skepticism, with outlets like Scientific American and the Guardian framing AI-assisted proofs as loopholes or overhyped rather than treating them as real progress.
- The math community has no shared norm yet for whether a paper should list a model like “GPT-6” or “Claude” as an author, how citations should work, or whether results should be shared informally on social media before formal publication.
- Human understanding is the deeper worry, separate from credit: even if a proof is correct, the question of whether any human actually understands why it’s true is becoming its own open problem.
- The incentive structure is shifting, since researchers now have to decide whether announcing an AI-generated result invites scrutiny that a human-only result wouldn’t face.
Built like a system. Not vibe-coded.
Remy manages the project — every layer architected, not stitched together at the last second.
How did this attribution fight start?
The immediate spark was a claimed AI-assisted result tied to the Navier-Stokes equations, one of the seven Clay Millennium Prize problems. When the claim surfaced, critics argued the model had effectively “stolen” a human researcher’s unpublished reasoning rather than independently deriving the result. Scientific American covered it with a headline suggesting the achievement was “eligible” for the million-dollar prize only through a “controversial loophole,” while the Guardian ran a piece arguing humans remain essential to the field and that tech companies were unwilling to admit it.
Aaronson’s post pushes back on that framing. His argument is straightforward: if a model’s output were purely derivative of a specific human’s hidden work, it wouldn’t generalize to other unsolved problems. But according to the rumors he references, AI systems have been credited with progress on other longstanding problems in theoretical computer science too, not just one isolated case. If that pattern holds, it undercuts the “it just copied someone” explanation, because a copy of one paper doesn’t explain solutions to unrelated problems.
Why would AI labs sit on solutions instead of publishing them?
Based on Aaronson’s account, the reasoning is reputational rather than technical. After the hostile response to the Navier-Stokes claim, he says AI companies (the ones best positioned to be doing this kind of work at scale, with dedicated research efforts and heavy compute spend, are widely assumed to be OpenAI and Anthropic) became more cautious about how they announce new mathematical results. The concern isn’t that the results are unsafe or dangerous. Nobody in this discussion is arguing the math itself poses a risk. The concern is about how to present machine-generated proofs in a way that doesn’t immediately trigger another round of “this isn’t real, it’s stolen” backlash, and how to do it before a rival lab publishes a similar result first.
That last point matters for timing. Part of what pushed the original claim out into public view, in this account, is competitive pressure between labs racing to be first, not necessarily a fully worked-out plan for how to communicate AI-generated math to the field.
Does the math community have norms for crediting AI-generated proofs?
Not yet, and that’s a big part of the tension. Aaronson lays out the live questions researchers are actually asking each other: Do you get credit for a paper you publish if an AI model produced the core argument? How would anyone verify that the human author, rather than an AI agent, actually made the discovery? Should a paper list a model like “GPT-6” or “Claude” as a co-author and let it thank the human in the acknowledgments for posing an interesting question? Is it acceptable to just tweet a result informally, the way one case involving a disproof of the Jacobian conjecture reportedly circulated, rather than going through peer review first?
One coffee. One working app.
You bring the idea. Remy manages the project.
None of these questions have settled answers. Traditional math publishing assumes a human worked through a proof, understood every step, and stands behind it professionally. AI-generated proofs complicate all three of those assumptions at once. A human can prompt a model, receive a valid proof, and not fully understand why it works, and existing norms have no clean way to handle that combination of correct-but-not-fully-understood results with a credit system that assumes human authorship implies human comprehension.
Is human understanding still necessary?
This is the part of the debate that goes beyond bureaucracy. A correct proof doesn’t automatically come with a human who understands why it’s correct. Historically, that gap has been unusual, most human-derived proofs get understood by the people close to them simply because deriving something forces you to reason through it. AI-generated proofs can break that link. A model can produce a valid argument that’s long, dense, or structured in a way no human easily follows, and now the field has to decide whether that’s a problem.
For some areas of math, an automated verification (checking the logic step by step) might be enough. For others, especially results people hope to build on, mathematicians want the conceptual insight, not just confirmation that the steps hold. If AI systems keep producing more of these results, entire subfields could accumulate correct but poorly understood proofs, and it’s unclear who has the incentive to go back and build human-level understanding of something a machine already “solved.”
Why does the media coverage lean skeptical?
A recurring pattern in coverage of these results is to look for a reason to discount them: a loophole, a stolen idea, an overstated claim. That skepticism isn’t baseless (extraordinary claims deserve scrutiny), but it also tends to move the goalposts. Aaronson’s framing is that each time an AI system does something that previously would have been treated as an unambiguous sign of major progress, critics respond by saying it doesn’t really count, and that the real threshold is whatever the AI hasn’t done yet. He argues that if you’d shown the same set of recent results to computer scientists twenty years ago, even skeptical ones would have called it a clear sign of a major shift. Whether or not that specific comparison holds, it points to a real dynamic: the goalposts for what counts as “real” progress keep moving in step with the progress itself.
Frequently Asked Questions
Did an AI model actually solve the Navier-Stokes Millennium Prize problem?
The claim involved an AI-assisted result connected to Navier-Stokes that generated significant controversy. Critics argued it relied on unpublished human work rather than being an independent AI derivation, and outlets like Scientific American framed it as only technically eligible for prize consideration through a contested interpretation. It remains a disputed case rather than a settled one.
Are AI labs really withholding solved math problems?
According to Scott Aaronson’s blog post, rumors suggest OpenAI and Anthropic have solutions to several major open problems in theoretical computer science that haven’t been publicly released, reportedly because of caution after the backlash to earlier claims. This is based on rumor and reporting from someone close to the field, not confirmed publication.
Should AI models be listed as authors on math papers?
There’s no agreed standard yet. Some informal cases have had models effectively “thank” the human collaborator in acknowledgments while producing the core argument, which sidesteps rather than resolves the authorship question. Formal journals and institutions haven’t settled on rules for this.
Why does it matter if humans understand an AI-generated proof?
A proof can be logically valid while still being something no human fully grasps conceptually. That matters because mathematics traditionally advances by humans building on understood ideas, not just verified conclusions. If proofs accumulate without human comprehension, it could slow the field’s ability to extend or apply those results.
Is this attribution problem unique to mathematics?
No. The transcript frames it as a preview of a broader issue: any field that relies on credited, understood, peer-reviewed work (science, law, engineering) will eventually face the same questions about crediting AI contributions and verifying human understanding of AI-produced output.