DiLoCo Decoupled: Google Revolutionizes Global AI Training
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
DiLoCo Decoupled: A Breakthrough in AI Training
Training state-of-the-art artificial intelligence models has traditionally relied on systems where chips must operate in perfect synchronization. This method, while effective for current models, is becoming increasingly complex to maintain as model scales grow. Google offers an innovative solution with DiLoCo Decoupled (Distributed Low-Communication), which divides training sessions into decoupled "islands" of computation, allowing for asynchronous data flow between them.
This new architecture provides increased resilience by isolating local disruptions, enabling other parts of the system to continue learning effectively. DiLoCo Decoupled thus overcomes the limitations of previous distributed methods, such as Data-Parallel, which were hampered by global communication delays.
Towards Asynchronous Fault-Tolerant Training
DiLoCo Decoupled builds on two prior advancements: Pathways, which introduced a distributed AI system based on asynchronous data flow, and DiLoCo, which significantly reduced the bandwidth required between distributed data centers. By combining these innovations, DiLoCo Decoupled allows for asynchronous training across separate learning units, ensuring that a chip failure in one area does not interrupt the progress of others.
During testing, a method called "chaos engineering" was employed to introduce artificial hardware failures during training sessions. DiLoCo Decoupled continued the training process after the loss of entire learning units and seamlessly reintegrated them when they came back online. This self-repairing capability ensures continuity of learning, even in the face of hardware failures.
Performance and Testing Results
Trials conducted with the Gemma 4 model confirmed the effectiveness of DiLoCo Decoupled. The system maintains greater availability of learning clusters than more traditional training methods while ultimately delivering the same level of machine learning performance. DiLoCo Decoupled requires orders of magnitude less bandwidth than conventional methods, making it extremely efficient.
With increasing levels of hardware failures, DiLoCo Decoupled continues to provide high "goodput," or useful training, while that of other approaches declines. Notably, the system successfully trained a 12 billion parameter model across four distinct regions of the United States using 2-5 Gbps of broadband network. This result was achieved over 20 times faster than conventional synchronization methods.
Evolution of AI Training Infrastructure
At Google, the full-stack approach to AI training, integrating hardware, software infrastructure, and research, is exemplified by DiLoCo Decoupled. This system allows for the utilization of unused computing resources worldwide, transforming idle capacity into useful computational power.
DiLoCo Decoupled also enables mixing different generations of hardware, such as TPU v6e and TPU v5p, within the same training session. This extends the lifespan of existing hardware and increases the total computational capacity available for model training. In our experiments, chips from different generations operating at different speeds consistently matched the performance of single-chip training sessions, ensuring that even older hardware can significantly accelerate AI training.
Moreover, as new generations of hardware do not arrive everywhere simultaneously, the ability to train across generations can alleviate recurring logistical and capacity bottlenecks. As we push the boundaries of AI infrastructure today, we continue to explore approaches for resilient systems necessary to unlock the next generation of AI.
Contribution and Support
This project was led by a team from Google DeepMind and Google Research, with key contributions from Arthur Douillard, Keith Rush, Yani Donchev, Zachary Charles, Ayush Dubey, Blake Woodworth, Ionel Gog, Josef Dean, Nova Fallen, and Zachary Garrett. Operational support was provided by Nate Keating and Jenny Bishop, with additional guidance from Jeff Dean, Marc’Aurelio Ranzato, Raia Hadsell, Arthur Szlam, Edouard Yvinec, Henry Prior, Paul Barham, Michael Isard, Daniel Ramage, Brendan McMahan, Chase Hensel, and Zoltan Egyed.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.