Brief IA

Google DeepMind: AI Security and Alignment, a Strategic Review

🔬 Research·Tom Levy·

Google DeepMind: AI Security and Alignment, a Strategic Review

Google DeepMind: AI Security and Alignment, a Strategic Review
Key Takeaways
1Google DeepMind has reinforced the importance of the chain of thought, influencing the industry consensus on its value.
2Google's Frontier Safety Framework has evolved to include measures against misalignment, involving multiple teams.
3DeepMind has published key papers on AGI safety, directing its internal efforts towards specific goals.
💡Why it mattersThese initiatives from Google DeepMind aim to anticipate and mitigate potential risks from advanced AIs, influencing the entire industry.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Introduction to the AGI Safety and Alignment Team

Within Google DeepMind, the AGI Safety and Alignment Team (ASAT) is dedicated to developing technical strategies to minimize existential risks associated with advanced artificial intelligence systems. Last year, the team published a foundational paper titled "An Approach to Technical AGI Safety and Security," which remains an essential reference for understanding their comprehensive approach to AI safety.

Evolution of Standards around Chain of Thought

One of ASAT's major contributions has been transforming the perception of the chain of thought within the industry. Initially viewed as unreliable and therefore not useful, this approach is now recognized as a valuable tool that deserves preservation. This paradigm shift has led to a preliminary consensus within the industry on the importance of the chain of thought. In-depth technical research has been published to help companies maintain the transparency of this chain longer than previously possible.

Strengthening the Frontier Safety Framework

Google DeepMind has also strengthened its Frontier Safety Framework (FSF), becoming the first company to include a dedicated section on misalignment within such a framework. Over the past two years, the scope of this framework has significantly expanded, involving numerous teams within Google. The FSF has played a crucial role in maintaining Google's vigilance against serious risks by developing mitigation measures well before they are needed.

Technical Planning and Strategic Documents

To structure its efforts, the team has published key documents such as "An Approach to Technical AGI Safety and Security" and the "GDM AI Control Roadmap." These publications detail specific areas where it is best to focus on AGI safety and control. They have proven very useful for organizing internal work and ensuring that efforts progress toward the appropriate goals.

Control and Monitorability of Agents

Monitoring agents has become a priority, as it may soon be necessary in practice. Although the chain of thought has been deemed unreliable for simple tasks, it is crucial for complex tasks requiring deep reasoning. The team co-authored "Chain-of-Thought Monitorability: A New and Fragile Opportunity for AI Safety," highlighting the importance of this approach. Capability assessments for stealth and situational awareness have also been developed to evaluate models' ability to evade monitors.

Deep Alignment: A New Approach

Deep Alignment is a relatively new area for the team, which is adapting to this intermediate phase. Historically, the team has focused on conceptual approaches, but it has recently adopted a new strategy. Current models are considered sufficiently similar to future AI systems for today's advancements to be transferable. Post-training analysis work, such as "SFT Drives Gemini’s Safety Properties" and "Why Do Naive SFT Filters For Safety Properties Fail?" has been conducted to explore these dynamics.

Improving the Interpretability of Language Models

Research on the interpretability of language models aims to deepen the scientific understanding of these models and use that understanding to enhance their safety. A pragmatic approach has been adopted to directly address significant issues related to interpretability. Tools and infrastructures have been developed to support internal research and the external ecosystem, with publications on building production-ready probes.

Advances in Amplified Supervision

Work on Amplified Supervision aims to enable weak supervisors to provide good incentives to powerful AI systems during training. This could potentially allow for the alignment of superhuman AI systems. The majority of efforts are focused on empirical work on various forms of debate, seeking to align AI systems with human objectives.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.