OpenAI Astra: Ten Advances and Debate on Attribution

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI initially claimed that no progress had been made on ten problems before quietly correcting the attribution of at least one of them, which has frustrated researchers. Nevertheless, the company is claiming ten breakthroughs all at once on topics that mathematicians consider central. These announcements come at a time when models remain weak in arithmetic and trivial tasks, while appearing to produce professional-level results in abstract mathematics. The contrast fuels a debate about the future of work and knowledge in mathematics in the age of models.
Quietly Corrected Attribution and Researchers' Reservations
OpenAI initially asserted that no progress had been made on the ten problems presented over the last decade, before one of the associated texts clarified that it relied on advancements made by two researchers. This presentation was then discreetly modified, which several researchers found unconvincing. Questions of attribution have thus arisen, especially since the new evidence and methods are based on prior work. According to mathematicians, the problems in question are indeed at the heart of their interests, unlike previous breakthroughs criticized as peripheral, and specialists had already encountered these issues without success.
Ten Claimed "Breakthroughs" and Extensive Documentation
A few weeks ago, OpenAI published a blog post accompanied by several hundred pages titled "10 Breakthroughs in Mathematics and Theoretical Computer Science." The company claims that its model Astra has somehow solved ten problems spanning multiple disciplines, including quantum game theory and sphere packing in dimensions higher than three. The announcement has caused quite a stir, partly because it contrasts with the pace of previous breakthroughs, which were often isolated, by presenting ten results at once. Researchers explain that solving even one of these problems by a human would already be impressive, and that an individual achieving all ten alone would seem almost unbelievable.
A Precedent in May: The Unitary Distance Conjecture
In May, an internal model from OpenAI refuted the unitary distance conjecture, an 80-year-old problem. Since then, OpenAI has introduced Astra, perceived by some as the trigger for an existential crisis within the field. Robert Hart believes that Astra is likely also behind the May refutation, although the company had only referred to an unnamed internal model at that time and did not respond to questions about it.
Contrasting Capabilities: Weakness in Calculation, Strength in Abstraction
Models continue to struggle with arithmetic and even trivial benchmarks like the days of the week, with a close associate of Robert Hart noting that they often respond that it is Wednesday. Elissa Welle has written that ChatGPT does not know how to tell time, and according to Robert Hart, this has not changed. Meanwhile, these systems are making strides in high-level abstract mathematics, to the point that their performance can appear professional-level in certain areas. The result is a heterogeneous landscape where basic weaknesses coexist with marked strengths in abstraction.
Why It Works Sometimes: Connections and Capacity Thresholds
Observers describe models capable of linking different domains and applying old methods in new ways, suggesting that a form of capacity threshold has been crossed for certain types of mathematics. The discipline itself is not a homogeneous block; it encompasses very distinct subdomains. AI remains particularly weak in counting, and according to testimonies from mathematicians, it is still quite poor in topology, a point that Robert Hart indicates he cannot verify. Overall, performance varies significantly across subdomains.
Anticipated Impact on Research, Jobs, and Verifiability
The discussion extends beyond mere results to touch on an existential crisis: what are mathematics, what do mathematicians do, and what will their role be in the future? This concern manifests both in areas where models seem to produce work equivalent to that of good practitioners and in others where they fail. Questions arise regarding employment and funding, while the observed progress mainly concerns problems that lend themselves to more computation and a certain verifiability. Laboratories highlight self-contained theoretical problems for which models generate proofs or solutions previously unattainable, then attempt to verify them and repeat the cycle. This acceleration fits within a rapid transition over the last six to twelve months, while other fields have evolved over several years, as until 2024, the prevailing idea remained that models failed at tasks as simple as counting letters in a word. Robert Hart further asserts that very specific cases, like counting strawberries, could be hard-coded, reminding that despite the importance of reasoning in mathematics, academic papers themselves exhibit few numbers.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.