OpenAI Publishes 722 Papers Claiming Progress on 372 Major Math Problems

openai publishes 722 papers claiming progress on 372 major math problems OpenAI has uploaded 722 manuscripts to GitHub that, taken together, claim either solutions to or progress on 372 major math problems. None of the papers has yet passed review by mathematicians.

OpenAI has uploaded 722 manuscripts to GitHub that, taken together, claim either solutions to or progress on 372 major math problems. None of the papers has yet passed review by mathematicians.

The company announced the release itself. According to OpenAI, the work was produced by an unreleased ChatGPT pioneer model, and its new advisory group confirmed the breakthroughs. The release is likely to add fuel to the Navier-Stokes AI controversy, which is already causing turmoil in the mathematical community.

One prompt, one agent

OpenAI is making bold claims. It says it has solved the four-dimensional Kakeya conjecture, improved key computer algorithms and made progress toward the extremely difficult Riemann hypothesis.

The way the results were produced is just as notable. OpenAI told Scientific American that almost every paper came from one prompt given to one AI agent. The company said the “average result” needed about three hours of ChatGPT Pro use.

Skeptics are pushing back hardest on that one-prompt, one-agent claim.

Mathematicians want more than OpenAI’s word

The value of these results won’t be clear until mathematicians have assessed them. Many scientists are still wary after the Navier-Stokes imbroglio.

MIT mathematician Andrew Sutherland put the doubt plainly to Scientific American: “Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified.”

That standard is reasonable. The model that produced the papers isn’t public, so for now no one outside OpenAI can repeat the work.

OpenAI followed only some of its advisers’ recommendations

OpenAI said the papers were released under guidelines recommended by its independent Advisory Group on Mathematics and Artificial Intelligence (AGMAI). Those guidelines ask for results to be released promptly through traditional academic channels. They also ask for details such as the name of the model used, the prompts and the compute costs.

The release does include the model’s reasoning, compute estimates and figures on how many problems were attempted. OpenAI ignored some of the board’s suggestions, though. It didn’t give specific compute times for individual problems, and it didn’t say which prompts it used.

Those omissions are important. Outside researchers would need the prompts and the per-problem compute figures to test the single-prompt claim.

“For this release, we’re publishing the results in a GitHub repository, with protocols for paper revisions and citations,” the company said. “We’re continuing to explore other community-hosted alternatives for this release which meet the committee’s guidelines. For future releases, we are committed to further improving the quality of the papers via the citations, mathematical exposition, and presentation of the results for better understanding.”

The release was expected

The release didn’t come out of nowhere. Last month, OpenAI said its models had resolved “more than 100 long-standing open problems across most areas of mathematics.” The new release provides the paperwork behind that claim and goes beyond it.

If you want to judge the work yourself, skip the press release and go straight to the GitHub repository. Start with the Kakeya paper, then follow how working mathematicians respond to it over the next few weeks. Until someone outside OpenAI reproduces even one of these results, Sutherland’s word for them still applies: unverified.