AMD and Cerebras Challenge Nvidia with AI Inference

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
AMD and Cerebras: A Strategic Alliance for the Future of AI
AMD recently announced a strategic partnership with chip startup Cerebras, marking a significant turning point in the field of AI inference. The company is betting on a new method called "disaggregated inference," which involves distributing workloads across multiple types of hardware rather than relying on a single chip. This innovative approach is embodied by AMD's Helios server system, which promises to compete with Nvidia's solutions in terms of performance and cost.
A New Approach to AI Inference
During an announcement made by AMD CEO Lisa Su, it was revealed that the company is partnering with Cerebras to adopt this new approach to AI inference. AI inference is the process by which artificial intelligence models generate responses from data. Traditionally, the same type of hardware was used to process queries and generate responses. However, AMD argues that these two tasks require distinct types of hardware. The Helios system, designed to handle a high volume of queries, works in tandem with Cerebras' giant chip, which excels at rapidly generating responses.
Helios in Cerebras' Data Centers
The partnership between AMD and Cerebras plans to integrate the Helios system into Cerebras' data centers by the end of the year. This collaboration comes at a time when the demand for AI chips is skyrocketing, with companies like Nvidia and Broadcom leading the way in chip design for AI training. Competition is intensifying as the industry increasingly turns towards the practical application of AI models.
A Paradigm Shift in the Industry
The agreement between AMD and Cerebras is part of a broader movement towards disaggregated inference, a shift that analysts are already observing. In June, UBS highlighted that the limitations of current architectures are driving this transition. Nvidia, for example, has integrated the startup Groq to explore similar configurations, and Amazon Web Services is also following this trend to improve efficiency and reduce costs. However, this approach presents challenges, particularly in terms of orchestration, which refers to the effective coordination between different chips to ensure smooth operation.
AMD Unveils Helios at the Advancing AI Event
At the Advancing AI event, AMD showcased Helios, its latest server system, which combines multiple types of AI chips. This announcement follows the unveiling of Nvidia's Vera Rubin NVL72 rack. AMD also announced multi-billion dollar infrastructure partnerships with cloud giants and AI labs such as OpenAI, Meta, Microsoft, Oracle, and Anthropic.
AMD Challenges Nvidia on Performance Grounds
AMD took the opportunity at the event to launch a direct attack on Nvidia, claiming that Helios offers up to 30% more inference tokens per dollar compared to Nvidia's Vera Rubin NVL72 rack. According to Lisa Su, each Helios system can deliver enhanced performance for larger models, extended capacity for longer contexts, and bandwidth capable of spanning thousands of racks.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.