
OpenAI published AI-generated full or partial solutions Tuesday to more than 370 outstanding mathematical problems, including some that have long been considered grand challenges in the field.
The volume of results stunned many mathematicians, while the way OpenAI has gone about tackling the problems and publishing the solutions divided the field. Some said they were enthusiastic about the results, seeing huge new areas for mathematicians to explore. Others said the approach OpenAI and other AI companies have taken to solving mathematical problems constitutes an assault on mathematics as a human academic discipline.
OpenAI said it achieved the results using an unreleased internal AI model. It said that on average the model took about three hours of computing time to arrive at each solution.
The massive cache of new solutions includes full or partial results for many of the problems mathematicians have considered the most important to the field. The results come weeks after OpenAI said it had used an unreleased internal model to solve the Navier-Stokes equations, one of the seven Millennium Prize problems for which the Clay Mathematics Institute offers a $1 million award. In the most recent batch of results, OpenAI said it had made progress on three other Millennium Prize problems but had not fully solved them.
AI companies have been targeting mathematical problems as a way of showcasing the capabilities of their models. AI researchers have also said that training their AI models on difficult math problems may help them learn many skills that generalize to other domains in the real world. For instance, it may help teach the models logical reasoning skills as well as how to be persistent in the face of difficult problems. It may also teach the models to do well in domains such as physics or economics that involve a lot of mathematics—although so far, it is unclear exactly how a model’s mathematical capabilities may generalize to domains, such as law or business strategy, which involve logical reasoning, but do not have objectively verifiable correct solutions.
Meanwhile, some of the traits learned in tackling very difficult mathematical problems—such as persistence—may increase safety risks. In recent “rogue AI” incidents, AI agents went to extreme lengths to achieve results in an evaluation, including taking unauthorized and illegal actions. Faced with a seemingly impossible challenge, a human might simply give up rather than resort to these kinds of unauthorized steps.
Dan Litt, a professor of mathematician at the University of Toronto, told Fortune he was excited about OpenAI’s results. “My view is that this is great for mathematics,” he said, adding that there were several solutions that OpenAI published that impacted problems he was interested in and was eager to understand the solutions OpenAI’s AI model found. “I think that it’s great to have new solutions to questions that I and others are interested in.”
Litt cautioned, however, that he is worried about the effect the solutions may have on the field of mathematics, especially if a perception that AI has “solved math” leads funding organizations to withdraw support for mathematical research or discouraged promising young mathematicians from entering the profession. “It’s important that society reaffirms support for human mathematical expertise if we want to get anything out of the progress on these problems that AI has made.”
Showing the work
When OpenAI published its Navier-Stokes solution, two mathematicians, who had also been working on a solution to the problem using AI tools, including OpenAI’s, accused the company of either intentionally or inadvertently feeding their work in progress to its AI model, helping point it in the direction of the solution. OpenAI denied this was the case, saying it did not feed its model the two mathematicians’s work and that the model could not have picked up any clues about their research from its training data because the cutoff for that data preceded the date on which the two mathematicians had begun using OpenAI’s Codex AI product to work on Navier-Stokes.
In response to the latest results, Tristan Buckmaster at New York University, one of the mathematicians involved in the earlier controversy, told the New York Times that it remained unclear whether mathematicians using OpenAI’s models had inadvertently helped point the company’s internal AI system toward the solutions it found. “There’s likely to be a bunch of results where they take someone’s work and then take it to completion,” he told the Times. Given the number of results being released simultaneously, he said “I don’t think they’ve done their sort of due diligence at all” to ensure the AI model had not plagiarized anyone’s work.
Last month, following criticism from mathematicians in the wake of its Navier-Stokes solution, OpenAI said it was forming an independent advisory group on mathematics and artificial intelligence hosted at the Institute for Advanced Study in Princeton, N.J.
Late last month, the group released a set of recommendations for the publication of AI-generated mathematical proofs. The recommendations included that AI-generated proofs should be published following the conventions of a traditional mathematical research paper, so that human mathematicians could more easily scrutinize and learn from the results. It also recommended that for each solution, an AI company should make public the name of the model used, the prompts used, the model’s “chain of thought” (or an output of its reasoning steps), the time it took the model to arrive at the solution, and an approximation of how much that computing time cost. It said that the company should also disclose how it decided to have the AI try to solve that particular problem and, if many results were published at once, that the company should publish a report detailing why those problems were targeted and how many other problems of comparable difficulty the model tried and failed to solve.
OpenAI published the latest mathematical solutions to GitHub, the code repository site. It followed some, but not all, of the steps the advisory group had recommended. The group published a statement on Tuesday saying “we reaffirm our published recommendations on responsible release.” It said its discussions with OpenAI had been “constructive” but that “ultimately it is up to the mathematical community to assess the extent to which our recommendations were followed successfully, and whether there are others we should suggest.”
The company released a blog post on Tuesday in which it said it had “drawn on” the advisory group’s advice about how to publish the solutions. “For future releases, we are committed to further improving the quality of the papers via the citations, mathematical exposition, and presentation of the results for better understanding,” OpenAI said. It said it was sharing formalizations of the proofs for many of the problems—these are versions of the proof that can be verified by specialized computer software—and would share more of these as it obtained them. It also said that for 10 problems it was publishing summaries of its model’s reasoning, estimates of the compute spent, and statistics about the number of attempted problems.
Litt, who was not a member of the advisory group, told Fortune he approved of most aspects of how OpenAI published the solutions. Having them on GitHub made them easily accessible for other mathematicians to study, he said, and he credited the company for not making too much of any particular advance in a blog post or marketing material intended for a non-technical audience. He also said he thought OpenAI lacked the capability to publish all the results in research papers that would meet rigorous academic standards, both because the AI models don’t write mathematical exposition well enough and struggle to cite prior mathematical work, and because OpenAI doesn’t employ enough mathematicians with expertise in enough areas to understand all the proofs the AI models can generate.
While some mathematicians have complained that AI-generated proofs, such as OpenAI’s Navier-Stokes solutions, are difficult to follow, making it hard for mathematicians to build on the results, Litt said he thought such concerns were “overstated.” He said mathematical writing was often difficult to follow any way. “I think to extract understanding from [the OpenAI results], there will be a huge amount of human labor involved, but it’s not so different from the labor that mathematicians have been doing forever,” he said.
OpenAI said it wanted its solutions “to push the frontier of human knowledge and enable further progress in mathematics.” It said it would be funding a series of workshops, conferences, and programs around helping mathematicians understand the results its AI system had generated.
The end of ‘Math 1.0’
The independent math advisory group said in its statement on OpenAI’s release that “the future of mathematical research cannot consist only of understanding results produced by AI labs. Mathematicians must be able to formulate their own questions, develop their own approaches, and explore directions that have not been selected as examples of an AI system’s capabilities. Equitable access to powerful research tools and adequate computational resources are essential to that freedom.”
Terence Tao, a UCLA mathematics professor considered among the world’s greatest living mathematicians, has been increasingly critical of the way AI companies have gone after mathematical problems, arguing that it is the process of arriving at solutions—not so much the solutions themselves—that advances mathematical understanding, and that by solving so many interesting problems so quickly, AI companies are discouraging students from becoming mathematicians, robbing the field of its future.
In a social media post on Mastodon Tuesday, Tao reiterated these criticisms. “Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is ‘solved,’ and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field,” he wrote. “Many fewer seminars, workshops, collaborations, or other activities are being generated from these results compared to traditional breakthroughs; few people are joining the community around the field as a consequence; and promising open directions are now being withheld from the public in fear that this will cause their own research to be ‘scooped.’”
Tao said that OpenAI’s mass publication of math solutions marked the end of “Math 1.0,” in which finding solutions to unsolved conjectures and problems, even if those solutions could not easily be understood at first, served as the field’s engine. He said there would now need to be a “Math 2.0” era that “will need to decenter the role of raw problem solving and value mathematical progress more holistically—for instance by elevating the role of exposition, but also that of community building and opening up new directions of study.”
Litt said he agreed with Tao that the field must change. He said OpenAI’s publication of such a massive set of solutions would help get the entire field “on the same page and understanding that we need to be a little bit radical about rethinking” things such as what kinds of contributions it rewards and how it trains PhD. students. And while Tao has generally sounded wistful about this transition, Litt said he was “optimistic” about it.
“One of my collaborators told me, I feel like I’ve been crawling my entire life, and now I can fly,” Litt said of the advent of AI as a tool for solving mathematical problems. “It’s like incredible what what we can do now.” He said he thought AI would enable human mathematicians to engage in much more “open-ended exploration” than was possible before. “We should expect mathematicians to be like way more productive in the future,” he said.










