Skip to main content

OpenAI Math Results: Hundreds of New Claims Released, and Mathematicians Say Not Like This

OpenAI's math results include claimed proofs of famous open problems. Mathematicians are split on how they were released, and most have not been checked by humans yet.

By Mohammad Muneer Ahmed
Published: Oct 08, 2026
6 mins read
👁️ 14 Unique Views
OpenAI Math Results: Hundreds of New Claims Released, and Mathematicians Say Not Like This
The scale of inference: Optimized for multimodal workloads.
Premium Insight

Why It Matters

India has a large community of mathematicians and computer scientists, and the questions in this release, such as the Unique Games Conjecture, are central to computer science research worldwide. The claims are OpenAI's and have not been broadly checked, so researchers and students should treat them as claims until independent experts confirm them. The dispute over how to release AI-written proofs, including prompts, costs, and readable write-ups, will affect any lab or university that publishes AI-assisted results.

On October 6, 2026, OpenAI posted a huge collection of math papers on GitHub. It says an unreleased internal model wrote them. The papers claim answers to hundreds of open problems, including the Unique Games Conjecture, a famous question in computer science. A group of mathematicians says the way OpenAI released them was wrong.

Almost none of it has been independently checked yet. OpenAI itself says some results may have errors.

How Many Results Are There?

The numbers look confusing, but they count different things. OpenAI's catalogue lists 719 manuscripts, which are separate papers. Early coverage said 722. OpenAI groups related papers into 372 families, and 372 is the number it uses for new results. The Association for Human Mathematics (AHM) says "over 700 files," which is the paper count. Scott Aaronson used both 372 and 376 in his posts, so even experts are not sure of the exact number.

OpenAI says nearly all results came from the same process on one internal model. Scientific American reported that a company spokesperson said nearly all came from a single prompt given to a single AI agent, and some took several tries. OpenAI says the average result used about three hours of ChatGPT Pro thinking time, and that the model was given about 4,000 problems.

Aaronson wrote that the model tried about 8,000 problems and solved about 5%. OpenAI's own number would put the rate closer to 9%, by my arithmetic, though the repository counts families, not problems. I could not find which figure is right.

What Is Claimed and What Is Checked

The headline claims are big. The Unique Games Conjecture is about how hard it is to get good approximate answers to many optimization problems. Aaronson also lists L=BPL, which asks whether randomness helps computers that use very little memory, and a way to multiply integers faster than the long-standing barrier. These are OpenAI's claims.

Verification is mixed. OpenAI says about 42% of its top results come with Lean proofs. Lean is a program that checks each logic step of a proof, a bit like spell-check for logic. It does not tell you whether a paper is clear, or whether the formal statement says what the paper claims. Aaronson says the Unique Games proof has a Lean certificate, but not every result does. OpenAI's README warns that some results without Lean could have issues.

Reading them is hard. Aaronson's wife, the complexity theorist Dana Moshkovitz, texted him that the Unique Games paper felt "written by someone on psychedelics" and was nearly impossible to read without AI help. A day later, with an AI tool's help, she said she mostly understood the proof and was impressed by its new ideas. That is one expert's view, not a verdict. Aaronson says almost no human had understood these proofs yet.

Why Mathematicians Pushed Back

The AHM statement, reposted as a guest post on Terence Tao's blog, calls the release "a demonstration of power." It says mathematicians did not ask for this work, rejects OpenAI's claim that it advances the field, and urges mathematicians to stop working with OpenAI. It is the AHM's statement, not Tao's own.

The fight centers on the Advisory Group on Mathematics and AI (AGMAI), which is linked to the Institute for Advanced Study. In September, AGMAI said it does not endorse labs testing advanced math problems on private models, and it asked them to stop. It also set release rules. For each result, labs should publish the prompts, a summary of the model's reasoning, the time taken, and the cost. Papers should be written in a readable style, and labs should say how many other problems the model failed on.

OpenAI says it drew on that advice. It published 10 reasoning summaries, average compute, and a rough problem count. In the repository's README and OpenAI's post, I found no prompts, no per-result times, and no model name. AGMAI itself did not condemn the release. It called its talks with OpenAI constructive and said the community must judge whether its advice was followed.

What Could Come Next

In the near term, OpenAI says it will keep adding Lean proofs, is exploring community-hosted repositories, and will fund workshops and conferences. Moshkovitz wants to give talks on the Unique Games proof soon. OpenAI also says it is working to release the model responsibly.

Aaronson points to a different route. Anthropic helped researchers Virginia Williams and Josh Alman write up and announce a separate result, in exchange for payment, rather than posting raw proofs. He says both approaches have trade-offs.

The long-term picture is speculation. If these proofs hold, math may shift toward humans explaining and guiding machine-found results. If many fail, this release may be remembered as a lesson in how not to share AI claims. Either way, the open question is who decides what counts as understood.

Found this analysis insightful?

Share with colleagues, engineers, and your network.

Link copied to clipboard!