Brief IA

GPT-5.6: More Efficient Agents, Lower Costs

🤖 Models & LLM·Tom Levy·

GPT-5.6: More Efficient Agents, Lower Costs

GPT-5.6: More Efficient Agents, Lower Costs
Key Takeaways
1On ARC‑AGI‑3, GPT‑5.6 Sol increases from 13.3% to 38.3% with retained reasoning and compaction, using ~6× fewer output tokens
2Luna claims 98% extraction accuracy of GPT‑5.5 at one eighteenth the cost; 78% success for ~14$ on 106 tasks
3The prompt cache increases to a minimum of 30 minutes; Ploy reduces uncached input by 28% on a 29,000 token prompt
💡Why it mattersPreviously exclusive use cases for flagship models are becoming feasible at lower costs through new API controls and smaller models, according to OpenAI.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI highlights quantified gains for its new GPT-5.6 models and the associated API, featuring extended caching, preserved reasoning mechanisms, and native multi-agent orchestration. Customer feedback cites increases in scores, reductions in latency and tokens, while Luna and Terra are presented as cheaper alternatives to previous flagship models.

ARC-AGI-3 and 30-Minute Cache Show Gains

OpenAI reports that with the standard architecture, GPT-5.6 Sol achieved 13.3% on ARC-AGI-3, and by activating preserved reasoning and compaction, the score rose to 38.3%, with approximately six times fewer output tokens, without modifying the model. Meanwhile, the prompt cache lifespan has been extended to at least 30 minutes, and deterministic breakpoints can be placed within the context window, which, according to the publisher, has improved cache success rates at startups.

At Ploy, Lorenzo Gentile reports adding breakpoints and workspace-specific keys to a shared prompt of 29,000 tokens, reducing uncached input by 28%. The 30-minute window allowed for the same context to be reused across executions. OpenAI specifies that an appropriate cache key increases the likelihood of reusing the same inference engine for a given prefix, which reduces latency.

GPT-5.6 Integrates Preserved Reasoning, Orchestration, and Tools

OpenAI claims to have trained GPT-5.6 end-to-end with three complementary components: preserving reasoning from turn to turn and natively compacting long conversations to maintain coherence; natively orchestrating multiple agents in parallel; and moving deterministic work into code via programmatic tool calls. The latter is used to filter, aggregate, and orchestrate outputs outside the context window, reserving tokens for judgment and lowering cost, latency, and context degradation.

The publisher distinguishes judgment tasks from data movement, filtering, and combination operations; for example, when retrieving 100 records, the model does not need to reason about each intermediary if tools can be orchestrated in JavaScript, including with parallel calls and out-of-context processing. At Rogo, Alex Wang reports that GPT-5.6 with this programmatic call matched the quality of their grid while consuming 21% fewer input tokens.

Luna and Terra Approach GPT-5.5 for Much Less

OpenAI explains that the historical strategy of prioritizing the flagship model for maximum reasoning was due to its mastery of long contexts and tool calling; with the 5.6 family, Luna and Terra, boosted by computation at test time, are presented as often close to GPT-5.4 and 5.5 while remaining significantly cheaper. At Hypha, Serhii Shchoholiev mentions a retention of 98% of GPT-5.5's extraction accuracy for a cost of one eighteenth, making high-quality document analysis feasible in many cases.

Browser Use claims to have submitted Luna to 106 difficult navigation tasks: 78% success for about $14, while a SOTA model achieved 80% for about $235. PlayerZero has made Luna its default model for code retrieval and decision modeling tasks; on a key multi-agent system code exploration task, the team reports a 64% reduction in inference costs, a 90% reduced response time, and a five-point increase in F1 score.

The API Activates Multi-Agent; ChatGPT Leverages "Ultra"

OpenAI describes a schema in which a main agent orchestrates sub-agents working in parallel, then aggregates their results for a final synthesis. This multi-agent orchestration can be directly activated in the response API and also feeds into ChatGPT's "ultra" capacity parameter.

E Chi, founder of Quadrillion, states that Qualia manages teams of agents on open research topics and notes a marked improvement of GPT-5.6 Sol over GPT-5.5, with faster completions than almost all other tested models; GPT-5.6 Sol has become their reference OpenAI model. At Obvious, Jon Bell reports launching six simultaneous specifications and describes GPT-5.6 as the best orchestrator seen at OpenAI, with no observed degradation in quality.

Startups Report Lower Costs with Less Effort

OpenAI positions GPT-5.6 as a continuation initiated with GPT-5: handling long tasks with fewer tokens, for more efficient agents at lower costs without disrupting the architecture. The publisher asserts that accuracy improves by reducing reasoning effort, citing a final review of agents where GPT-5.6 Sol, with low effort, outperformed GPT-5.5 at high effort with constant architecture. Production teams report lowering their costs by adjusting this parameter; at Hex, Izzy Miller explains that this configuration avoids false leads, detects data absence, and converges with fewer tokens.

OpenAI emphasizes that startups are now combining a finer selection of models with new controls—continuity of reasoning, native multi-agents, and programmatic tool calls—to build faster and more capable agents at a fraction of the cost. The publisher concludes that the economics of agent building have changed: cases that once required a flagship model at every step could achieve comparable, if not better, results by relying on smaller models, adjusting reasoning effort, and making effective architectural choices.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.