OpenAI and the Era of Autonomous AI: GPT-5.4 Redefines Standards
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Launch of GPT-5.4 by OpenAI
OpenAI recently unveiled its latest artificial intelligence model, GPT-5.4, on March 5. This model stands out for its work-oriented design, incorporating advanced features that make it particularly suitable for professional applications. GPT-5.4 is available in several forms, including GPT-5.4 Thinking in ChatGPT, gpt-5.4, and gpt-5.4-pro in the API, as well as in Codex. This variety of formats allows for flexible use tailored to the specific needs of users.
The model integrates the coding capabilities of GPT-5.3-Codex and adds features such as native computer usage and tool searching. It also offers an impressive context window of 1 million tokens, with a default parameter of 272K tokens. This extended capacity enables users to handle complex tasks with increased efficiency.
The pricing for using GPT-5.4 has been set at $2.50 / $15 per million tokens for the base model, and $30 / $180 for the Pro version. While these costs may seem high, the increased efficiency of tokens more than offsets the expenses, as demonstrated by tests conducted by OpenAI. Queries exceeding 272K input tokens result in doubled costs, highlighting the importance of effective resource management.
Accelerated Release Pace
The pace of model releases by OpenAI is also noteworthy. Since December, several versions have been successively launched: GPT-5.2, GPT-5.3-Codex on February 5, followed by Codex-Spark on February 12, then GPT-5.3 Instant on March 3, before the arrival of GPT-5.4. A staff member from OpenAI confirmed on the developer forum that monthly releases are now the norm. The progress made is attributed to post-training evaluation loops, reasoning time checks, tool selection, memory compaction, and product integration. While the race for the base model remains crucial, the fastest gains are being made in the surrounding engineering.
Performance of GPT-5.4
GPT-5.4 represents a significant leap in many dimensions, although it is not a clear knockout. On the Intelligence Index from Artificial Analysis, it ties with Gemini 3.1 Pro Preview with a score of 57. On LiveBench, GPT-5.4 Thinking xHigh slightly outperforms Gemini 3.1 Pro Preview, scoring 80.28 against 79.93. On the Vals benchmark grid, the picture is mixed: GPT-5.4 dominates ProofBench, IOI, and Vibe Code Bench, while Gemini 3.1 Pro leads in other areas such as LegalBench, GPQA, MMLU Pro, LiveCodeBench, and Terminal-Bench 2.0. Claude Opus 4.6 and Claude Sonnet 4.6 dominate SWE-bench and the broad composite of Vals and Finance Agent, respectively. There is no longer a single leading model, as each AI has its own areas of strength.
The story of OpenAI's benchmarks this time is particularly focused on the workplace. On GDPval, which tests real knowledge work across 44 professions, GPT-5.4 achieves 83.0% compared to 70.9% for GPT-5.2. In internal spreadsheet modeling tasks, it scores 87.3% against 68.4%. On OSWorld-Verified for desktop navigation, it scores 75.0%, surpassing the human benchmark of 72.4% and nearly doubling GPT-5.2's 47.3%. On BrowseComp, it scores 82.7%, with Pro reaching 89.3%. OpenAI claims a 33% reduction in false statements and 18% fewer responses containing errors compared to GPT-5.2. Mainstay reported that across approximately 30,000 HOA and property tax portals, GPT-5.4 achieved 95% success on the first try and 100% in three tries, being about 3 times faster while using 70% fewer tokens. Harvey’s BigLaw Bench: 91%.
Interface Gap and Microsoft's Reaction
Despite ongoing progress on GDPval, OpenAI faces an interface gap for office work. The preamble and ongoing response direction of GPT-5.4 are genuinely useful. ChatGPT for Excel and the new financial data integrations are a smart approach for high-value workflows. However, OpenAI still lacks a non-developer surface as user-friendly as Claude Cowork for delegating real and complex office tasks. Codex and the API now have serious computing usability, but the overall experience remains more technical than it probably should be if OpenAI aims to dominate daily office work.
Microsoft reacted quickly this week with Copilot Cowork. The company announced that it is integrating the technology behind Claude Cowork directly into Microsoft 365 Copilot, with enterprise controls, security positioning, and pricing under the existing umbrella of Microsoft 365 Copilot. This gives Microsoft a clear distribution advantage, as Word, Excel, PowerPoint, Outlook, and Teams are already key tools where much office work takes place. However, Microsoft's execution so far has often given the impression of a company with perfect distribution but intermittent product urgency. In contrast, OpenAI and Anthropic have generally been more effective at encouraging users to actually use their products. Microsoft still has an installed base. The question is whether it can convert that into genuine product demand before model labs sell their own work agents directly to businesses.
Andrej Karpathy's Autoresearch
The other significant story of the week, although it seems smaller on paper, is Andrej Karpathy's autoresearch experiment. Karpathy publicly reported that after about two days of self-tuning on a small nanochat training loop, his LLM agent found about 20 transferable additive changes from a depth 12 proxy model to a depth 24 model, reducing the "Time for GPT-2" from 2.02 hours to 1.80 hours, an improvement of about 11%. The autoresearch repository describes the setup: giving an AI agent a small but real LLM training environment, allowing it to modify code, run short experiments, check if validation improves, and repeat the process overnight.
Prospects for Autonomous Improvement
Many people immediately raised the idea that "this is just hyperparameter tuning." I think that misses the economic point. If a swarm of agents can reliably explore optimization parameters, attention adjustments, regularization choices, data mixing recipes, initialization patterns, and architectural details on inexpensive proxy runs, then promote promising changes to larger scales, it is already an extremely valuable research process, even if it doesn't resemble a synthetic scientist inventing an entirely new paradigm from scratch. Cutting-edge research is full of limited research problems with delayed but measurable feedback. This is exactly the terrain where agents can start to accumulate.
Here is the trajectory I expect from now on. Labs will give swarms of agents significant GPU budgets to conduct thousands of small and medium experiments on proxy models. They will seek better attention mechanisms, better optimization schedules, better training curricula, better post-training recipes, and better evaluation methods. Promising ideas will then be promoted through increasingly larger training runs. Human experts will remain involved at obvious bottlenecks: deciding which metrics are important, spotting false positives, designing new research spaces, choosing which ideas deserve costly scaling, and co-designing higher-stakes modifications once dealing with real parameter counts and serious training budgets. But the internal loop of "propose, implement, test, compare, iterate" seems increasingly automatable.
We already have hints that labs are on the first rung of this ladder. OpenAI stated that GPT-5.3-Codex was the first model "instrumental in its own creation," with initial versions used to debug its own training, manage deployment, and diagnose evaluations. To be precise, OpenAI has been much more explicit publicly about autonomous development in GPT-5.3-Codex than in GPT-5.4 itself. But the direction of the journey is hard to miss.
There is also an important nuance in OpenAI's system map for GPT-5.4. The company claims that GPT-5.4 Thinking does not meet its threshold for high-capacity autonomous AI improvement, which it defines as being roughly at the level of a mid-career performing research engineer. I think this distinction is important, but probably in the opposite sense to what some skeptics assume. The threshold for economically useful autonomous improvement is much lower than the threshold for cutting-edge autonomous research. A model does not need to be a synthetic principal scientist to improve prompts, evaluations, tools, scaffolding, training recipes, and smaller model experiments around it. This lower threshold is the one that accelerates everything else.
Why This Matters
The center of gravity in AI has shifted from "smart chatbot" to "reliable operator." The winning system is no longer the one that writes the most beautiful single response. It is the one that can stay focused for an hour, use the right tools without drowning in token fees, make unattractive software work that no one has exposed through clean APIs, compress its own history, and allow a human to steer without restarting all the work. GPT-5.4, Codex, the agent teams of Opus 4.6, Gemini CLI, Microsoft's Copilot Cowork, and Karpathy's autoresearch all point in the same direction.
That is why GDPval is more important than GPQA or MMLU. The trajectory from 12.4% with GPT-4o to 83.0% with GPT-5.4 in about 18 months does not measure chatbot intelligence. It measures how close AI is to replacing the actual output of knowledge workers on well-defined tasks. We have crossed the halfway mark, and the curve is steepening. That said, GDPval still has obvious limitations, and we hope the project will receive more funding from OpenAI to expand the benchmark and test more multi-stage and long-term agentic tasks.
Karpathy's autoresearch extends the same logic inward. If agents can reliably improve the training stack itself, the rate of improvement accumulates. I expect Frontier Labs to give swarms of agents significant GPU budgets this year to explore attention mechanisms, optimization variants, and data set recipes on small proxies before scaling the winners. Human researchers will co-design at scale. My hypothesis is that by the end of the year, we could well see a leading model whose development has been materially shaped by this type of autonomous AI research loop. I do not mean fully autonomous in the science fiction sense. I mean that a significant fraction of attention adjustments, optimization choices, data recipe changes, post-training methods, and evaluation corrections will have been discovered, filtered, and...
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.