Nine of the world's top mathematicians have agreed to help OpenAI stop fumbling its math breakthroughs. The OpenAI math advisory group, formally called the Advisory Group on Mathematics and Artificial Intelligence, was announced on September 21 in a guest post on Fields Medalist Terence Tao's blog. Hosted at the Institute for Advanced Study in Princeton, the nine-member panel includes Timothy Gowers, Martin Hairer, Edward Witten, and Melanie Matchett Wood, drawn from institutions including Stanford, Harvard, Oxford, and Cambridge. Its mandate is to advise AI labs on reviewing and communicating math results. The reason for the rush is clear: OpenAI says its unreleased model has resolved more than 100 long-standing open problems across most areas of mathematics, and it wants to start publishing them.
Nine elite mathematicians, 100 unsolved problems, and a company that controls the pace.
Mathematicians who spoke to The Verge described a messy rollout that confused the very community it was meant to reassure. The announcement's timing, the unclear provenance of the group, and the looming wave of results have created a tense atmosphere. Researchers fear their life's work could be dispatched overnight by a company they suspect is treating mathematics as a publicity stunt.
What went wrong with OpenAI's earlier math announcements?
OpenAI's relationship with the mathematics community has been deteriorating for months. The company's announcement of a solution to the Navier-Stokes Millennium Prize problem descended into fights over credit, scooping, and poor treatment of mathematicians whose prior work was relevant. The Clay Mathematics Institute established seven Millennium Prize Problems in 2000, each carrying a $1 million award. Only one had been solved before OpenAI entered the picture: the Poincare conjecture, resolved by Grigori Perelman in 2003.
Beyond the Navier-Stokes controversy, mathematicians cited specific scholarly failures. Researchers described poorly written manuscripts with scant engagement with surrounding literature, making it difficult to assess a result's significance or credit prior work. The company quietly altered documents after release without announcing changes or leaving a clear record. Hairer separately described instances where OpenAI modified manuscripts "sneakily" in response to criticism, calling the practice "shoddy" and "really bad and sloppy scholarship."
This matters beyond academic etiquette. If a model produces a proof but the accompanying manuscript is incoherent, the mathematical community cannot verify, build on, or properly cite the result. The solution exists in a liminal state: claimed but not integrated into the body of mathematical knowledge. University of Connecticut professor Alvaro Lozano-Robledo told The Verge that AI companies fundamentally misunderstand the burden they are creating. "The burden is that they are producing a solution," he said, but solutions are often less important than the understanding that comes with them. "They need our expertise," he added. "They need us to celebrate that solution." AI can generate results, but determining their significance still requires human mathematicians.
The pattern echoes broader problems with AI scientific discovery claims that we have tracked before, where biologists pushed back on similar overreach from AI labs promising breakthroughs without the scholarly infrastructure to support them.
Who sits on AGMAI and what power do they actually have?
The nine initial members are Francois Charles (ENS-PSL), Camillo De Lellis (IAS, GSSI), Timothy Gowers (College de France, Cambridge), Martin Hairer (EPFL, Imperial College London), Nikhil Srivastava (Berkeley, Simons Institute), Ulrike Tillmann (Oxford, INI), Ravi Vakil (Stanford), Edward Witten (IAS), and Melanie Matchett Wood (Harvard). The Next Web reported the full roster, which includes multiple Fields Medalists and MacArthur fellows.
The group's independence is the contested part. AGMAI's own website says it formed after OpenAI approached some members about establishing an external advisory board. Those mathematicians then decided to create an independent group instead and invited others to join. It remains unclear which of the nine were initially contacted by OpenAI. Hairer insisted to The Verge that the group receives no financial, technical, or other support from the company, and that members have signed no restrictive agreements beyond standard confidentiality for early access to research. "We don't work for OpenAI, are not paid by them, and it's totally independent," he said. He also said the group is open to working with other frontier AI labs and had already begun conversations with some, though he declined to name them.
OpenAI stressed in its own announcement that members can challenge the company publicly, publish their advice, and offer guidance it did not request. But the company also drew a hard line: the group "will not be responsible for advising us on how to pace our internal progress on mathematics." AGMAI gets a say in how results are communicated. It has no authority over whether or how fast they are produced. The Decoder noted this limitation explicitly, reporting that the model produced the 100-plus results after roughly a month of training.
That boundary is the crux. A group that can advise on presentation but cannot slow production functions as a communications consultancy with star power but no governance authority. For builders, this is a familiar pattern: an advisory board with no veto power, assembled after a crisis, with enough prestige to lend credibility but insufficient authority to change the underlying behavior. Hairer acknowledged the risk that OpenAI may spin his involvement to its advantage. "They're not going to sort of damage me within the math community," he said. "In some sense, I don't really care about what they say. But what I do worry about is the math community as a whole."
The chart below shows the scale gap: OpenAI claims over 100 open problems solved by its unreleased model, compared to 7 Millennium Prize problems in total and 1 previously announced. The advisory group tasked with managing this deluge has 9 members and no authority over production pace.

What does the 100 results claim signal about model capability?
The 100-plus figure is the number that should get your attention. OpenAI says these results came after roughly a month of training, covering "most areas of mathematics." Even allowing for OpenAI's tendency to overstate, the scale suggests a model that can reliably produce novel mathematical results across domains, not just in one narrow specialty.
For anyone building with AI, that capability ceiling matters. If a model can generate publishable math results across fields, it has strong reasoning, long-chain verification, and the ability to produce outputs that experts consider non-trivial. The same capabilities that solve open math problems also power code generation, formal verification, scientific reasoning, and complex multi-step planning. The math results are a proxy for general reasoning depth, not just a niche academic curiosity.
But the claim comes with heavy caveats. OpenAI has not released the results, named the specific problems, or provided manuscripts for independent review. The 100-plus number is a self-reported metric from a company with a track record of botched math announcements and quietly edited papers. Lozano-Robledo told The Verge that the company is "feeding the frenzy" by hinting at another Millennium Prize problem without revealing what it is. "Like, what other Millennium problem are they talking about when they say 'We have almost solved another Millennium problem'?" he said. "That sentence makes no sense in mathematics."
Should builders care about AI doing mathematics?
If you build AI products, the OpenAI math saga is a case study in what happens when capability outpaces accountability. Three things are worth your attention:
-
The governance gap is a product risk. OpenAI assembled elite advisors but gave them no power over the core decision: when and how fast to release results. If you are building AI tools that produce consequential outputs, whether in research, legal, medical, or financial domains, an advisory board with no veto is a liability. Design your review process with actual stopping power, or do not pretend you have one.
-
Verification infrastructure is the bottleneck. The math community's complaint is that results arrive without the surrounding scholarship needed to verify them: proper literature engagement, stable manuscripts, clear provenance. If your product generates novel claims, the verification pipeline matters more than the generation pipeline. This connects to the prompt data privacy questions we raised when OpenAI's last math scoop raised concerns about how inputs and prompts were handled.
-
The solution without understanding problem is coming to your domain. Lozano-Robledo drew a sharp distinction: in mathematics, solutions are often less important than the understanding that accompanies them. AI can produce an answer without producing insight. If you are building tools for knowledge workers, your users will face the same disconnect. A model that answers correctly but cannot explain why is a tool that breaks trust the first time it is wrong.
What happens when the first batch of results lands?
The real test for AGMAI is the first release cycle. If OpenAI dumps 100 results with poorly written manuscripts and no literature engagement, the advisory group will have failed on its first day, regardless of how independent it claims to be. If the group manages to slow the release, demand proper scholarship, and coordinate with affected researchers, it will have justified its existence.
Colva Roney-Dougal, a professor at the University of St Andrews, told The Verge she is "unclear whether I should be trying to rush out as many papers as possible, or passively waiting to see what these results are, or carrying on as normal." That is the human cost of a company that treats mathematics as a capability demo. Hairer said the group started two days before being asked about its plans and laughed: "So it's not like we have a big master plan." Nine people, no plan, 100 results to review, and a company that controls the pace. The math community is left to hope that star power and good faith are enough. They rarely are.
Sources
- The Verge: OpenAI keeps bulldozing mathematicians
- The Verge: OpenAI wants to consult elite mathematicians about how to not fumble again
- The Decoder: OpenAI says its internal model solved over 100 long-standing math problems
- The Next Web: Top mathematicians will advise OpenAI on releasing its AI maths results
- Clay Mathematics Institute: Millennium Problems
