⚡
Brief IA
›

OpenAI: Ultrafast in Beta, GPT-5.6 14× Faster

🤖 Models & LLM·Tom Levy·

OpenAI: Ultrafast in Beta, GPT-5.6 14× Faster

OpenAI: Ultrafast in Beta, GPT-5.6 14× Faster
⚡
Key Takeaways
1Limited preview: Ultrafast for GPT-5.6 Sol is accessible to a select group, with expansion planned as capacity allows.
2Announced performance: up to 14× faster and 750 tokens/s, powered by Cerebras, initially in the OpenAI API.
3Drivers and use cases: testing with clients and internally for incident response and research.
💡Why it matters — OpenAI claims that combining speed and intelligence allows for the integration of AI into real-time tasks (incidents, finance, support) and enables more iterations throughout the workday.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI offers limited access to Ultrafast for GPT-5.6 Sol and plans to expand the offering as capabilities increase. Meanwhile, the company is already using it for incident response and internal research, while pilot customers are testing the promised speed, powered by Cerebras, of up to 750 tokens per second.

Access is limited and will expand with capacity

GPT-5.6 Sol in Ultrafast mode is offered in a limited preview for a selected group of customers. OpenAI plans to broaden access as capacity increases. Companies can sign up to be notified when this access expands. The registration aims to inform when access will be extended.

OpenAI uses it for incident management and research

A group of developers at OpenAI has tested Ultrafast to identify workflows that benefit from real-time responses. The team is already using it for incident response. During alerts, engineers need to establish an accurate picture while the system and evidence are still evolving. Teams rely on Ultrafast to sift through logs, examine traces, summarize discussions, identify the next checks to perform, and contribute to the preparation or validation of a fix, all in significantly reduced time. According to OpenAI, this approach shortens the period between detecting a signal, verifying a hypothesis, and making decisions on the next steps, while keeping engineers in charge of judgment and deployment. In terms of research, the team leverages Ultrafast to identify sources, query datasets, and collect, structure, and synthesize information using connected tools. Where a series of experiments used to be launched overnight for review in the morning, OpenAI finds that the loop is tightening to allow for multiple iterations within the same day.

Pilot customers are experiencing Ultrafast in production

OpenAI has deployed GPT-5.6 Sol in Ultrafast mode to an initial group of companies, covering areas such as software development, commerce, financial analysis, support, and other interactive uses. According to the company, starting with business processes allows for observing these situations in real-world production contexts. OpenAI states that these experiments help identify cases where significant acceleration provides the most benefits and observe product evolution as the model aligns with user pace. During the preview phase, OpenAI collaborates with a small group of clients and seeks to determine in which cases speed has the most pronounced impact, so that this feedback gradually informs its products.

Four testimonials: Jane Street, Podium, Basis, Rogo

At Jane Street, John Crepezzi finds the speed increase brought by Cerebras impressive and believes it opens new ways to use models, making developers' work more focused and productive. Courtland Lykins at Podium explains that Ultrafast has proven invaluable in the voice stack and that the speed completely changes the calling experience for complex tasks. For Mitch Troyanovsky, co-founder of Basis, the solution enables synchronous experiments that were previously limited by intelligence and combines throughput in tokens per second with model intelligence. Alex Wang at Rogo states that speed changes what users can reasonably leverage, making complex financial research comparable to real-time interaction.

Announced performance: 14× and 750 tokens/s, via Cerebras

OpenAI announces that Ultrafast runs GPT-5.6 Sol up to 14 times faster than standard processing, with an initial launch in the OpenAI API. Powered by Cerebras, the offering could generate up to 750 tokens of output per second. The company claims that this speed makes its most intelligent model accessible to time-sensitive products and workflows, and that GPT-5.6 improves efficiency at all levels of infrastructure, making advanced intelligence more affordable and useful for more people. OpenAI argues that the pursuit of real-time speed often involved resorting to smaller or specialized models and presents Ultrafast as a step towards more useful work per second. According to the company, no longer sacrificing intelligence for speed allows for integrating AI into the most time-constrained segments. OpenAI says it has already observed encouraging use cases: incident response and reliability, financial research and security, customer support and voice, commerce, as well as live research and experimentation.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.