OpenAI: Ultrafast in Beta, GPT-5.6 14× Faster

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI offers limited access to Ultrafast for GPT-5.6 Sol and plans to expand the offering as capabilities increase. Meanwhile, the company is already using it for incident response and internal research, while pilot customers are testing the promised speed, powered by Cerebras, of up to 750 tokens per second.
Access is limited and will expand with capacity
GPT-5.6 Sol in Ultrafast mode is offered in a limited preview for a selected group of customers. OpenAI plans to broaden access as capacity increases. Companies can sign up to be notified when this access expands. The registration aims to inform when access will be extended.
OpenAI uses it for incident management and research
A group of developers at OpenAI has tested Ultrafast to identify workflows that benefit from real-time responses. The team is already using it for incident response. During alerts, engineers need to establish an accurate picture while the system and evidence are still evolving. Teams rely on Ultrafast to sift through logs, examine traces, summarize discussions, identify the next checks to perform, and contribute to the preparation or validation of a fix, all in significantly reduced time. According to OpenAI, this approach shortens the period between detecting a signal, verifying a hypothesis, and making decisions on the next steps, while keeping engineers in charge of judgment and deployment. In terms of research, the team leverages Ultrafast to identify sources, query datasets, and collect, structure, and synthesize information using connected tools. Where a series of experiments used to be launched overnight for review in the morning, OpenAI finds that the loop is tightening to allow for multiple iterations within the same day.
Pilot customers are experiencing Ultrafast in production
OpenAI has deployed GPT-5.6 Sol in Ultrafast mode to an initial group of companies, covering areas such as software development, commerce, financial analysis, support, and other interactive uses. According to the company, starting with business processes allows for observing these situations in real-world production contexts. OpenAI states that these experiments help identify cases where significant acceleration provides the most benefits and observe product evolution as the model aligns with user pace. During the preview phase, OpenAI collaborates with a small group of clients and seeks to determine in which cases speed has the most pronounced impact, so that this feedback gradually informs its products.
Four testimonials: Jane Street, Podium, Basis, Rogo
At Jane Street, John Crepezzi finds the speed increase brought by Cerebras impressive and believes it opens new ways to use models, making developers' work more focused and productive. Courtland Lykins at Podium explains that Ultrafast has proven invaluable in the voice stack and that the speed completely changes the calling experience for complex tasks. For Mitch Troyanovsky, co-founder of Basis, the solution enables synchronous experiments that were previously limited by intelligence and combines throughput in tokens per second with model intelligence. Alex Wang at Rogo states that speed changes what users can reasonably leverage, making complex financial research comparable to real-time interaction.
Announced performance: 14× and 750 tokens/s, via Cerebras
OpenAI announces that Ultrafast runs GPT-5.6 Sol up to 14 times faster than standard processing, with an initial launch in the OpenAI API. Powered by Cerebras, the offering could generate up to 750 tokens of output per second. The company claims that this speed makes its most intelligent model accessible to time-sensitive products and workflows, and that GPT-5.6 improves efficiency at all levels of infrastructure, making advanced intelligence more affordable and useful for more people. OpenAI argues that the pursuit of real-time speed often involved resorting to smaller or specialized models and presents Ultrafast as a step towards more useful work per second. According to the company, no longer sacrificing intelligence for speed allows for integrating AI into the most time-constrained segments. OpenAI says it has already observed encouraging use cases: incident response and reliability, financial research and security, customer support and voice, commerce, as well as live research and experimentation.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.