Brief IA

OpenAI: Jalapeño Promises Reduced Latency and Gradual Rollout

🤖 Models & LLM·Tom Levy·

OpenAI: Jalapeño Promises Reduced Latency and Gradual Rollout

OpenAI: Jalapeño Promises Reduced Latency and Gradual Rollout
Key Takeaways
1OpenAI plans to release Jalapeño in small quantities by the end of the year, with a ramp-up in 2027
2On InferenceX, OpenAI claims 1.5 to 1.9× per watt and 1.7 to 3.6× in latency
3Jalapeño is an inference ASIC developed with Broadcom, with no total replacement of the existing fleet planned
💡Why it mattersOpenAI announces efficiency and latency gains on benchmarks while preparing for a limited rollout and maintaining its computing partnerships.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI claims that Jalapeño, its inference chip designed with Broadcom, outperforms Nvidia systems in energy efficiency and latency based on InferenceX tests. The company is preparing a small-scale launch by the end of the year, with increased volume expected in 2027, and states that it does not intend to replace its entire fleet or do without its partners.

OpenAI plans a small deployment by the end of the year, with volume increase in 2027

OpenAI plans to roll out Jalapeño in small quantities by the end of this year, followed by an increase in volume in 2027, according to Richard Ho. The company does not specify how many chips it plans to deploy next year. Despite the announced gains, OpenAI does not expect to replace its entire lineup with Jalapeño and maintains a computing strategy that involves partners, including Nvidia. Richard Ho adds that OpenAI will continue to develop the second and third generations of this chip.

InferenceX Tests: 1.5 to 1.9× per watt, 1.7 to 3.6× in latency

To evaluate the chip, OpenAI used the InferenceX platform, comparing Jalapeño to the best results recorded at the time by Nvidia superchips GB200 or GB300. According to OpenAI, Jalapeño delivered 1.5 to 1.9 times more AI work per watt than these systems on GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, and an end-to-end latency that is 1.7 to 3.6 times lower. Richard Ho states that these measurements translate into faster responses, more responsive agents, and more reliable access. He argues that Jalapeño combines low latency and high throughput, whereas typically, one has to trade off between these two criteria.

Jalapeño is an inference ASIC designed with Broadcom

Introduced in June, Jalapeño is an application-specific integrated circuit designed for inference. The chip was developed in partnership with Broadcom. OpenAI states that it performs tasks more efficiently and responds faster than other AI systems. The company published a blog post about this on Tuesday, and Richard Ho elaborated on these points during a briefing with journalists.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.