Brief IA

OpenAI: GPT-5.6 Sun Surpasses Opus 5 with an Optimized API

💻 Code & Dev·Tom Levy·

OpenAI: GPT-5.6 Sun Surpasses Opus 5 with an Optimized API

OpenAI: GPT-5.6 Sun Surpasses Opus 5 with an Optimized API
Key Takeaways
1OpenAI claims that GPT-5.6 Sol surpasses Opus 5 on ARC-AGI-3 with a score of 38.3%.
2This high score was achieved through the use of specific features of OpenAI's API, rather than the standard test configuration.
3The ARC prize may have used an outdated API, calling into question the neutrality of its testing environment.
💡Why it mattersThis situation raises questions about the reliability of performance comparisons between AI models, influenced by the tools used.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI: GPT-5.6 Sol Surpasses Opus 5 with an Optimized API

OpenAI claims that GPT-5.6 Sol outperforms Opus 5 on the ARC-AGI-3 benchmark thanks to its latest API and two additional parameters.

The co-founder of the ARC Prize, François Chollet, responded to OpenAI's results by distinguishing between two types of test configurations. He stated that configurations "customized to solve the benchmark or containing knowledge about the benchmark format" are not acceptable. In contrast, general-purpose API parameters "that were not developed for ARC-AGI-3 and are available to all API users" are considered valid. Chollet thus concedes that GPT-5.6 Sol's score in the context of the ARC Prize disadvantaged OpenAI.

He noted that the ARC Prize has had "many exchanges with OpenAI about the best way to test their models, particularly regarding compaction," and welcomed the fact that the company "is starting to find answers." Chollet also emphasized that the use of different parameters by different providers creates "a potential parity issue," but he considers this acceptable "as long as the parameters and costs are clearly reported."

OpenAI asserts that it can compete on ARC-AGI-3. After Claude Opus 5 from Anthropic quadrupled the record score on the logical benchmark, OpenAI now shows that GPT-5.6 Sol achieves 38.3% with two API parameters, surpassing 30.2% of Opus 5.

The scores of ARC-AGI-3 for GPT-5.6 Sol increase significantly when using OpenAI's Responses API with "Retained Reasoning," which maintains the model's thought chain between steps, and "Compaction," which summarizes old context instead of truncating it. In the official testing environment, GPT-5.6 Sol only scored 7.8% as the model's reasoning is discarded after each action.

OpenAI argues that benchmarks never measure just the model, but also the technical configuration surrounding it. This is true, and ARC-AGI-3 is designed to test the pure performance of the model. The official ARC scores use a standardized approach without provider-specific parameters to ensure fair comparisons, the ARC Prize stated in response to OpenAI's results. The point of contention is whether the ARC Prize used an older "OpenAI-style completion API" that lacked features already offered by the Claude API, which would make the comparison unfair for OpenAI.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.