GPT-5.6: More Efficient Agents, Lower Costs

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI highlights quantified gains for its new GPT-5.6 models and the associated API, featuring extended caching, preserved reasoning mechanisms, and native multi-agent orchestration. Customer feedback cites increases in scores, reductions in latency and tokens, while Luna and Terra are presented as cheaper alternatives to previous flagship models.
ARC-AGI-3 and 30-Minute Cache Show Gains
OpenAI reports that with the standard architecture, GPT-5.6 Sol achieved 13.3% on ARC-AGI-3, and by activating preserved reasoning and compaction, the score rose to 38.3%, with approximately six times fewer output tokens, without modifying the model. Meanwhile, the prompt cache lifespan has been extended to at least 30 minutes, and deterministic breakpoints can be placed within the context window, which, according to the publisher, has improved cache success rates at startups.
At Ploy, Lorenzo Gentile reports adding breakpoints and workspace-specific keys to a shared prompt of 29,000 tokens, reducing uncached input by 28%. The 30-minute window allowed for the same context to be reused across executions. OpenAI specifies that an appropriate cache key increases the likelihood of reusing the same inference engine for a given prefix, which reduces latency.
GPT-5.6 Integrates Preserved Reasoning, Orchestration, and Tools
OpenAI claims to have trained GPT-5.6 end-to-end with three complementary components: preserving reasoning from turn to turn and natively compacting long conversations to maintain coherence; natively orchestrating multiple agents in parallel; and moving deterministic work into code via programmatic tool calls. The latter is used to filter, aggregate, and orchestrate outputs outside the context window, reserving tokens for judgment and lowering cost, latency, and context degradation.
The publisher distinguishes judgment tasks from data movement, filtering, and combination operations; for example, when retrieving 100 records, the model does not need to reason about each intermediary if tools can be orchestrated in JavaScript, including with parallel calls and out-of-context processing. At Rogo, Alex Wang reports that GPT-5.6 with this programmatic call matched the quality of their grid while consuming 21% fewer input tokens.
Luna and Terra Approach GPT-5.5 for Much Less
OpenAI explains that the historical strategy of prioritizing the flagship model for maximum reasoning was due to its mastery of long contexts and tool calling; with the 5.6 family, Luna and Terra, boosted by computation at test time, are presented as often close to GPT-5.4 and 5.5 while remaining significantly cheaper. At Hypha, Serhii Shchoholiev mentions a retention of 98% of GPT-5.5's extraction accuracy for a cost of one eighteenth, making high-quality document analysis feasible in many cases.
Browser Use claims to have submitted Luna to 106 difficult navigation tasks: 78% success for about $14, while a SOTA model achieved 80% for about $235. PlayerZero has made Luna its default model for code retrieval and decision modeling tasks; on a key multi-agent system code exploration task, the team reports a 64% reduction in inference costs, a 90% reduced response time, and a five-point increase in F1 score.
The API Activates Multi-Agent; ChatGPT Leverages "Ultra"
OpenAI describes a schema in which a main agent orchestrates sub-agents working in parallel, then aggregates their results for a final synthesis. This multi-agent orchestration can be directly activated in the response API and also feeds into ChatGPT's "ultra" capacity parameter.
E Chi, founder of Quadrillion, states that Qualia manages teams of agents on open research topics and notes a marked improvement of GPT-5.6 Sol over GPT-5.5, with faster completions than almost all other tested models; GPT-5.6 Sol has become their reference OpenAI model. At Obvious, Jon Bell reports launching six simultaneous specifications and describes GPT-5.6 as the best orchestrator seen at OpenAI, with no observed degradation in quality.
Startups Report Lower Costs with Less Effort
OpenAI positions GPT-5.6 as a continuation initiated with GPT-5: handling long tasks with fewer tokens, for more efficient agents at lower costs without disrupting the architecture. The publisher asserts that accuracy improves by reducing reasoning effort, citing a final review of agents where GPT-5.6 Sol, with low effort, outperformed GPT-5.5 at high effort with constant architecture. Production teams report lowering their costs by adjusting this parameter; at Hex, Izzy Miller explains that this configuration avoids false leads, detects data absence, and converges with fewer tokens.
OpenAI emphasizes that startups are now combining a finer selection of models with new controls—continuity of reasoning, native multi-agents, and programmatic tool calls—to build faster and more capable agents at a fraction of the cost. The publisher concludes that the economics of agent building have changed: cases that once required a flagship model at every step could achieve comparable, if not better, results by relying on smaller models, adjusting reasoning effort, and making effective architectural choices.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.