OpenAI Drops 372 Math Results on GitHub. Mathematicians: "Show Your Work"

OpenAI Drops 372 Math Results on GitHub. Mathematicians: "Show Your Work"

Every math teacher's favorite line is "show your work." This week OpenAI answered with a GitHub repo holding hundreds of manuscripts, and the math world collectively replied: "Great. Now show us the model."

A Semester's Worth of Proofs in One Push

OpenAI published 722 math manuscripts grouped into 372 result families in a public GitHub repository, all produced by an unreleased internal frontier model. According to a company spokesperson quoted by Scientific American, the model produced almost every result from a single prompt given to a single AI agent, though some took multiple attempts. Each result used about three hours of ChatGPT Pro-level thinking compute on average. That's a dramatic drop from the 10,000-agent swarm OpenAI reportedly needed for an earlier Navier-Stokes result.

The claimed results include improvements to major computer algorithms and progress related to the Riemann hypothesis. Many proofs have been formalized in Lean, a language whose proofs can be checked by machine, and more formalization is planned. OpenAI also says it consulted the Institute for Advanced Study's Advisory Group on Mathematics and AI and plans to fund workshops on making sense of AI-produced results.

Correct Isn't the Same as Important

The reaction has been somewhere between "wow" and "excuse me?" Lean can confirm that a proof is logically valid, but it can't tell you whether the result is original or even worth knowing. Twenty-five Fields Medal winners warned in an open letter that problem-solving is "merely a tool" and that the real goal is conceptual understanding. In other words, a firehose of proofs that nobody has time to read might do more harm than good.

The sharper critique is about reproducibility. The model itself, the exact prompt, and per-problem compute weren't released, only averages. MIT's Andrew Sutherland put it plainly: until people can replicate the results, claims about one-shotting problems with a single agent should be treated as unverified. That's not hostility. That's just how science works.

The bottleneck in AI-assisted work is no longer producing output. It's checking it, and that's true whether you're proving theorems or writing product descriptions.

If your team is turning out AI-generated content faster than anyone can review it, we can help you build review and verification steps into the workflow so speed doesn't cost you accuracy.

Source: The Decoder