OpenAI's GPT-5.5: A Costly Yet Promising Agentic Advancement
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
An Ambitious Launch for OpenAI
On April 23, OpenAI unveiled GPT-5.5, an artificial intelligence (AI) model distinguished by its ability to perform tasks autonomously. This model is presented as a new generation of agentic AI, capable of planning, using tools, verifying its results, and completing tasks without human intervention. This advancement marks a turning point in how OpenAI envisions the use of AI for practical applications.
GPT-5.5 is the first model to be retrained since the GPT-4.5 version, in collaboration with NVIDIA's rack-scale systems GB200 and GB300 NVL72. This partnership has enabled the creation of a model that reduces the need for human interventions to correct the course of tasks. Available to Plus, Pro, Business, and Enterprise users via ChatGPT and Codex, GPT-5.5 was also integrated into the API starting April 24.
Impressive Performance on Benchmarks
OpenAI highlights the performance of GPT-5.5 through Terminal-Bench 2.0, a benchmark that evaluates the models' ability to handle command-line workflows. GPT-5.5 achieved a score of 82.7%, surpassing GPT-5.4, which reached 75.1%, and Claude Opus 4.7 with 69.4%.
Regarding SWE-Bench Pro, which assesses problem-solving on GitHub, GPT-5.5 scored 58.6%, demonstrating its ability to solve more problems in a single attempt compared to previous versions. OpenAI also introduced Expert-SWE, an internal benchmark where tasks are estimated to have a median human completion time of 20 hours. GPT-5.5 scored 73.1% there, up from 68.5% for GPT-5.4.
In the realm of long-context reasoning, the MRCR v2 benchmark, which tests a model's ability to retrieve specific information from a large document, saw GPT-5.5 achieve 74.0%, compared to 36.6% for GPT-5.4.
However, on the MCP Atlas benchmark from Scale AI, which evaluates the use of context protocols, Claude Opus 4.7 remains in the lead with 79.1%. GPT-5.5 does not yet have a recorded score on this test, but OpenAI remains confident in the overall performance of its model.
Cost and Token Efficiency
Access to the GPT-5.5 API is priced at $5 per million input tokens and $30 per million output tokens, which is double the rates of GPT-5.4. OpenAI justifies this cost by the increased efficiency of GPT-5.5, which accomplishes the same tasks with fewer tokens, making the effective costs about 20% higher, a claim supported by the independent lab Artificial Analysis.
The Pro version of GPT-5.5, aimed at Pro, Business, and Enterprise users, is priced at $30 per million input tokens and $180 per million output tokens. This model stands out for its ability to perform additional calculations in parallel on complex problems, ranking first on BrowseComp, OpenAI's agentic web browsing benchmark, with a score of 90.1%.
Before committing to a model change, it is advisable to test token efficiency against real workloads. At a volume of 10 million output tokens per month, GPT-5.5 costs $300 compared to $250 for Claude Opus 4.7, a 20% difference that is only justified if the superior agentic performance of the model reduces iterations and retries.
Adoption and Future Prospects
OpenAI reports that over 85% of its employees now use Codex each week, particularly in engineering and marketing departments. For instance, the communications team utilized GPT-5.5 to analyze six months of speaking request data, allowing the model to create an evaluation and risk framework to automate low-risk approvals.
Greg Brockman, co-founder of OpenAI, described this advancement as a "real step forward" towards the computing of the future, while Jakub Pachocki, chief scientist, noted that the progress of models over the past two years has seemed "surprisingly slow."
OpenAI asserts that GPT-5.5 maintains the same latency per token as GPT-5.4 in production, while offering a higher level of intelligence. Larger and more powerful models are often slower, but this trade-off has been avoided here.
The question of whether benchmark scores translate into production gains for teams using real agentic pipelines remains open and will require a few weeks of observation. The high score on Terminal-Bench is promising for DevOps automation and unattended terminal agents, while the gap on MCP Atlas should be monitored by those heavily relying on tool usage orchestration.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.