Brief IA

OpenAI and GPT-5.6 Sol Ultra: A Mathematical Breakthrough After 50 Years

🤖 Models & LLM·Tom Levy·

OpenAI and GPT-5.6 Sol Ultra: A Mathematical Breakthrough After 50 Years

OpenAI and GPT-5.6 Sol Ultra: A Mathematical Breakthrough After 50 Years
Key Takeaways
1OpenAI's GPT-5.6 Sol Ultra model has solved the Double Cover Cycle Conjecture, a mathematical problem that has remained unsolved for 50 years.
2In less than an hour, 64 sub-agents collaborated to produce a proof described as "surprising" and "elementary" by mathematician Thomas Bloom.
3However, Bloom criticized the lack of references to prior work in the proof generated by the AI.
💡Why it mattersThis breakthrough raises questions about AI's ability to innovate beyond mere recombination of existing knowledge.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI and GPT-5.6 Sol Ultra: Mathematical Breakthrough After 50 Years

OpenAI's AI model, GPT-5.6 Sol Ultra, has produced a proof of the Double Cover Cycle Conjecture using 64 sub-agents working in parallel. Mathematician Thomas Bloom praises the proof but criticizes the lack of citations.

OpenAI announced that GPT-5.6 Sol Ultra generated a complete proof of the Double Cover Cycle Conjecture, which had remained unproven for about 50 years. The AI model took just under an hour to accomplish this task.

In simple terms, the conjecture addresses a fundamental question in graph theory: is it possible to find a set of cycles in any network of vertices and edges that traverses each individual edge exactly twice? The problem was independently formulated by several mathematicians in the 1970s. Since then, many partial solutions have been proposed for specific cases, but no generally accepted proof has been found.

Machine Persistence

According to OpenAI, the proof comes entirely from GPT-5.6 Sol Ultra. The article was written by GPT-5.6 Sol. Thomas Bloom from the University of Manchester describes it as a "very beautiful proof," noting that the solution is "short, elementary, and could have been discovered in the 1980s." It does not require any new mathematical theory but skillfully combines known tools.

Why haven’t humans found it? Bloom suspects that a small counterintuitive twist in the reasoning played a key role. A human mathematician would likely have tried the obvious approach, seen that it failed, and moved on. The AI, on the other hand, does not get discouraged; it keeps trying small variations until one works.

"You can imagine first trying the natural labeling, checking the linear algebra, and when that fails, shrugging and thinking 'well, I expected to fail, I guess this can't be done so easily' - while the AI does not get discouraged and continues to try small variations," Bloom writes.

Bloom's initial assessment is the most detailed to date; a complete mathematical verification by the scientific community is still pending.

AI Still Doesn't Cite Its Sources

Bloom points out that the fundamental mathematical ideas behind the proof date back at least to a 1983 paper by Bermond, Jackson, and Jaeger. He criticizes the fact that OpenAI's article does not mention this prior work at all, so anyone reading just the article might think that the AI invented the underlying strategy itself.

"I assume that this previous work had a significant influence on OpenAI's proof, and it is unfortunate that they are not mentioned at all […]", Bloom writes. "[…] This is a common problem with proofs and articles generated by AI: they use ideas and proof strategies drawn from the literature without proper citation." The mathematician doubts that the AI found the solution by itself, "given that its first instinct for problem-solving is usually to search for all related articles on a problem and read them."

This is an ongoing debate surrounding reasoning models. Do they "simply" find existing knowledge and recombine it? Or do they actually produce something new through creative work? For this proof, Bloom seems to lean towards the first option.

AI Shows What Humans Could Have Solved With More Patience

Bloom compares the result to the unit distance conjecture, which OpenAI also recently solved. Both were major open problems "that turned out to be much easier than expected - no major new theory was needed, and one can imagine many alternative histories where these proofs would have been found decades earlier," he writes.

He expects AI systems to solve more conjectures like this one, "those whose solutions require only existing, well-developed theory, plus a lot of patience and belief." But according to Bloom, "this is probably only a small proportion of open problems, and we do not know in advance which ones they are."

"But in this strange world where large AI companies spend a lot of time and money tackling many open problems at once (and of course only report successes), we will soon discover more of what was within our reach all along," he writes.

How to Formulate a Complex Mathematical Proof?

Part of the solution lies in the prompt crafted by humans. It essentially engineers the type of persistence that Bloom describes as key to finding the proof. First, the prompt asks the model to assume that a complete proof exists, thus cutting off its most honest response: that the conjecture is open. Then, it forbids the model from searching the internet to check if the conjecture has already been solved and to respond that the conjecture is unresolved. Thus, the model has virtually nowhere to go except to solve the problem.

The verification is equally strict. Partial results, reductions to other unproven conjectures, summaries of the current state of research, and explanations of the problem's difficulty have all been rejected as insufficient. The model cannot respond until a complete proof is ready and has passed an adversarial test.

The rest of the prompt resembles more the guidelines of a research lab than a typical AI prompt. Most of the 64 agents are deliberately kept in the dark about the currently most promising approach to encourage "independent thinking." Adversarial agents then check each candidate proof against a detailed list of typical errors, looking for elements such as incorrectly identified closed paths as cycles or reductions that accidentally create new bridges in the graph.

The model was instructed to compute for at least eight hours before it could even consider giving up. It finished in one hour.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.