Hundreds of mathematical solutions published by OpenAI in early October fall short of the field's standards — that assessment was voiced on October 8 by prominent members of the mathematics community and independent analysts. At the center of the criticism are three issues: the model's chain of thought is almost entirely undisclosed, much of the proof corpus has not been formalized, and inconsistencies were found in the Navier–Stokes solution.

The Core of the Criticism

During the week of October 5, 2026, OpenAI published hundreds of solutions to the world's hardest mathematical problems. The company said it had collaborated with an advisory group of elite mathematicians to avoid repeating a previous controversy. But analysis of the published materials tells a different story: according to TechCrunch, only 10 of 719 manuscripts openly disclose the model's chain of thought, and 42% of the proofs have not been formalized at all.

Formalization is the process of converting a proof into a form a computer can verify automatically — for example, the Lean programming language. This is precisely the step that exposes hidden errors and gaps inside a proof, which is why the community treats it as standard practice.

The analysis was published on October 8 on the TechCrunch website. It notes that OpenAI consulted an advisory group of elite mathematicians before publication, aiming to avoid a repeat of the earlier controversy — yet the published materials do not fully align with that group's recommendations.

Discrepancies in the Navier–Stokes Proof

On October 6, King's College London researcher Alexandros Bastounis and Cambridge University scientists Fabian Sirselli and Anders Hansen examined OpenAI's Navier–Stokes solution in detail in a paper published on arXiv. The authors documented at least two discrepancies between the natural-language proof and its corresponding Lean code, concluding that the formalized Lean proof does not match the natural-language proof.

If the formalized version does not match the original text, it remains unclear exactly which part of the proof was verified — and that directly affects the scientific credibility of the published result.

How the Recommendations Were Followed

The AGMAI group (Advisory Group on Mathematics and Artificial Intelligence), based at the Institute for Advanced Study at Princeton University and comprising nine prominent researchers, published recommendations for leading labs in late September. The first recommendation was to stop benchmarking advanced mathematical problems on closed models; another key requirement was to make formalization standard practice.

The published materials fall short of these requirements: 42% of the proofs were not formalized, and the model's chain of thought was disclosed in only 10 of 719 manuscripts. Whether AGMAI's recommendations were followed is ultimately for the mathematics community itself to assess, the group stated:

"Ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully." — AGMAI statement.

The group also suggested that OpenAI help fund the work of human mathematicians needed to make its solutions meaningful.

Tao's Critique and the Wider Debate

Fields Medalist Terence Tao criticized OpenAI's approach. He wrote that problems are being solved autonomously by "AI prompters" who do not understand the result deeply enough, cannot answer questions about the obtained result, or present it before field specialists.

OpenAI emphasized that this time it worked with an advisory group of elite mathematicians to avoid repeating the earlier controversy. Yet community criticism began within days of publication — evidence that the debate over the criteria for publishing AI-derived mathematical results continues.

AGMAI's guidelines, published in late September, also called for making formalization the standard. Yet 42% of the published proofs were not formalized — noted as the largest gap between recommendation and practice.

Key Numbers

Of the 719 published manuscripts, the model's chain of thought is disclosed in only 10 — not even 1.5% of the total. 42% of the proofs were not formalized, meaning nearly every second proof remained in a form that has not passed automated computer verification.

Given AGMAI's recommendation to make formalization standard practice, these figures starkly illustrate the gap between the group's recommendations and the published practice. Analysts note that the undisclosed chain of thought sharply limits outside experts' ability to independently verify the model's path to its conclusions.