OpenAI Criticized for the Reliability of 719 Mathematical Proofs

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Researchers from Cambridge and King's College report discrepancies between natural language demonstrations and their Lean versions in solutions published by OpenAI. The AGMAI advisory group highlights unmet requirements, including the lack of metadata and the use of proprietary models. Only 10 out of 719 publications reveal the chain of thought, and 42% of the proofs have not been formalized.
Discrepancies Identified Between Text and Lean Code
Mathematicians from the University of Cambridge and King's College London have identified at least two discrepancies between a proof written in natural language and its encoding in Lean for an OpenAI solution to a problem derived from the Navier-Stokes equations. The process followed by the models involves first producing a natural language explanation, then attempting to formalize it in Lean, a language that theoretically allows for the certification of a proof's validity through compilation. These discrepancies do not necessarily undermine the solutions, but they raise doubts about the reliability of self-formalization without human intervention. They believe that natural language proofs produced by OpenAI and other self-formalized proofs in Lean should not be considered reliable without peer review and a rigor equivalent to that applied to traditional proofs.
Human Understanding and Responsibility Still Uncertain
Melanie Wood, a professor at Harvard, emphasizes that a solution generated by a model lacks any human understanding at the time of publication, and that the work of interpretation remains to be done afterward. Terence Tao criticizes an approach where AI users solve problems without commitment to the field, nor the ability to explain, teach, or discuss the results. Mathematicians remind us that in traditional research, authors assume responsibility for their discoveries and disseminate them through articles, conferences, and seminars, which fosters understanding and application of the solutions. It is not established that OpenAI takes responsibility for ensuring this human understanding at the time of publishing its proofs, as outlined by AGMAI principles.
Several Gaps with AGMAI Recommendations
AGMAI has called for the inclusion of machine-readable metadata linking natural language and formal artifacts, which OpenAI has not done in its recent publications. AGMAI's first recommendation was also not to test advanced mathematical problems on proprietary models, while OpenAI explicitly states it uses its proprietary models on open research problems. In terms of transparency, only 10 out of 719 publications contain a model's chain of thought, and 42% of the published proofs have not been formalized. AGMAI specifies that it is up to the mathematical community to assess the implementation of its recommendations and suggests that OpenAI contribute to funding the work of mathematicians necessary to interpret these solutions.
What OpenAI Delivered and What Remains in Debate
OpenAI has published 719 manuscripts claiming solutions to difficult mathematical problems, rapidly disseminating its results and providing information on the methods employed by its models. The lab claims to have consulted an advisory group of mathematicians, AGMAI, hosted by the Institute for Advanced Study in Princeton and composed of nine researchers. This group published guidelines at the end of September but did not provide a thorough evaluation of this wave of proofs. Discrepancies exist between the natural language formulation and the formalization for a problem presented as worth a million dollars. Mathematicians believe that OpenAI has not met the necessary standards, particularly regarding human understanding of the results.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.