Brief IA

AI Models: Struggling to Conceal Their Reasoning

🤖 Models & LLM·Tom Levy·

AI Models: Struggling to Conceal Their Reasoning

AI Models: Struggling to Conceal Their Reasoning
Key Takeaways
1AI agents must be closely monitored to ensure their safety and alignment with user objectives.
2The ability of AI models to conceal their reasoning could complicate monitoring, but it remains limited at present.
3CoT-Control, an evaluation tool, shows that models struggle to follow complex reasoning instructions.
💡Why it mattersThe difficulty of AI models in masking their reasoning ensures better transparency and safety for users.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

The Crucial Importance of Monitoring AI Agents

With the rapid evolution of artificial intelligence agents, now capable of performing increasingly complex tasks autonomously, the need to monitor their behavior has become paramount. In the context of an iterative deployment approach, it is essential to understand how these systems behave in real-world environments. This allows for the continuous adjustment of security measures based on technological advancements. To support this effort, a defense-in-depth strategy is implemented, integrating several layers of complementary security. Among these are security training, behavioral testing, agent code review, and chain of thought (CoT) monitoring. The latter involves analyzing the reasoning steps generated by agents while executing tasks. These reasoning traces provide valuable signals, both during training and deployment, helping to identify potentially dangerous behaviors or those inconsistent with the goals set by the user.

The Challenges of Reasoning Model Monitorability

Currently, the reasoning of models is generally interpretable and easy to monitor. However, this monitoring capability could deteriorate in the future for various reasons. A concerning scenario would be if agents become capable of deliberately modifying or obscuring their reasoning when they know they are being monitored. This could lead to an overestimation of a system's alignment or safety, making monitoring systems less reliable. In this study, we examine whether current reasoning models can control their chain of thought in a way that reduces their monitorability. Understanding this capability is crucial to ensure that CoT monitoring remains a robust safeguard as AI systems improve in performance. The results show that current models struggle to control their CoTs, even when informed of the monitoring. While controllability is higher for larger models, it decreases when models are asked to reason for longer periods or undergo adjustments after training. These observations suggest that CoT controllability does not currently pose a major risk to monitorability, although ongoing assessment is necessary as models evolve.

Defining "CoT Controllability"

The chain of thought refers to the intermediate reasoning steps that an AI agent generates while solving a task. These reasoning traces provide useful safety signals, helping monitoring systems evaluate whether an agent's behavior is safe or aligned. CoT controllability is the ability of a reasoning model to manipulate the properties of its CoT in a way that reduces its monitorability. Measuring this capability within the framework of realistic incentives to avoid monitors is complex. Therefore, we use a practical proxy: the model's ability to follow user instructions regarding its CoT. For example, we test whether a model can adhere to instructions affecting its internal reasoning, such as using only lowercase letters, avoiding certain keywords, or adding specific phrases, rather than simply modifying its visible output.

Launching CoT-Control

To deepen this study, we introduced CoT-Control, an open-source evaluation suite designed to measure the controllability of the CoT of reasoning models. CoT-Control includes over 13,000 tasks based on established benchmarks such as GPQA, MMLU-Pro, HLE, BFCL, and SWE-Bench Verified. Each task is created by pairing a benchmark problem with a CoT instruction, such as avoiding certain relevant keywords in the CoT, using only lowercase letters, or adding a given word to each sentence.

Results: Low CoT Controllability

Evaluations conducted on 13 reasoning models reveal that they all struggle to reliably control their CoTs. We measure controllability as the fraction of reasoning traces where the model successfully complies with CoT instructions. Across the leading models, CoT controllability scores are low, ranging from 0.1% and not exceeding...

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.