OpenAI: GPT-5.6 Sun Surpasses Opus 5 with an Optimized API

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI: GPT-5.6 Sol Surpasses Opus 5 with an Optimized API
OpenAI claims that GPT-5.6 Sol outperforms Opus 5 on the ARC-AGI-3 benchmark thanks to its latest API and two additional parameters.
The co-founder of the ARC Prize, François Chollet, responded to OpenAI's results by distinguishing between two types of test configurations. He stated that configurations "customized to solve the benchmark or containing knowledge about the benchmark format" are not acceptable. In contrast, general-purpose API parameters "that were not developed for ARC-AGI-3 and are available to all API users" are considered valid. Chollet thus concedes that GPT-5.6 Sol's score in the context of the ARC Prize disadvantaged OpenAI.
He noted that the ARC Prize has had "many exchanges with OpenAI about the best way to test their models, particularly regarding compaction," and welcomed the fact that the company "is starting to find answers." Chollet also emphasized that the use of different parameters by different providers creates "a potential parity issue," but he considers this acceptable "as long as the parameters and costs are clearly reported."
OpenAI asserts that it can compete on ARC-AGI-3. After Claude Opus 5 from Anthropic quadrupled the record score on the logical benchmark, OpenAI now shows that GPT-5.6 Sol achieves 38.3% with two API parameters, surpassing 30.2% of Opus 5.
The scores of ARC-AGI-3 for GPT-5.6 Sol increase significantly when using OpenAI's Responses API with "Retained Reasoning," which maintains the model's thought chain between steps, and "Compaction," which summarizes old context instead of truncating it. In the official testing environment, GPT-5.6 Sol only scored 7.8% as the model's reasoning is discarded after each action.
OpenAI argues that benchmarks never measure just the model, but also the technical configuration surrounding it. This is true, and ARC-AGI-3 is designed to test the pure performance of the model. The official ARC scores use a standardized approach without provider-specific parameters to ensure fair comparisons, the ARC Prize stated in response to OpenAI's results. The point of contention is whether the ARC Prize used an older "OpenAI-style completion API" that lacked features already offered by the Claude API, which would make the comparison unfair for OpenAI.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.