Brief IA

Researchers Identify Internal Patterns in AI Reasoning

⚖️ Regulation & Ethics·Tom Levy·

Researchers Identify Internal Patterns in AI Reasoning

Researchers Identify Internal Patterns in AI Reasoning
Key Takeaways
1Reasoning operations leave distinct signatures in the internal states of several language models
2The separation is maximal in the intermediate layers and cannot be explained by word choice
3The effect has been validated on other models and datasets, but remains limited to mathematical tasks
💡Why it mattersThis mapping of internal steps could open new avenues for AI supervision and safety, linking the produced text to internal computation.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Research conducted by KAIST and Naver AI Lab indicates that operations such as deduction or calculation leave a distinct signature in the internal representations of language models. The effect peaks in the intermediate layers, has been replicated across other models and datasets, but remains confined to mathematical tasks. This mapping is of direct interest to supervision and security approaches that link produced text and internal computation.

Out-of-Sample Validated Scope, Assumed Limits, and Security Stakes

Additional tests have verified the separability of reasoning operations. This phenomenon has also been observed with Llama-3-8B, and it has been possible to successfully transfer classifiers trained on Qwen3-8B to GPQA-Diamond and MATH-500. However, it is noted that the trials focus on mathematical tasks and a limited number of models, and the use of these signals to detect errors or guide generation remains to be explored further. This field is directly related to security, as the relationship between textual output and internal computation is deemed crucial. OpenAI considers the reading of the chain of thought as one of the few available supervision tools, while Anthropic has shown that models reveal only 25 to 39% of the clues actually used. Other work indicates that Claude Opus 4.6 processes more information than what appears in its displayed reasoning, and that the Recurrent Depth in Astra shifts part of the reasoning to internal representations, precisely the space explored by this study.

Eight Operations, Three Models, and Automated Labeling

The researchers defined eight recurring operations, such as extraction, decomposition, formula recall, deduction, and calculation. They subjected three models (Qwen2.5-7B, Qwen3-8B, and Gemma4-31B) to mathematical problems, then segmented the solution paths and labeled each segment using GPT-5. The goal was to determine whether the model's internal states retraced these steps. They found that the same type of response generates a distinct activation pattern depending on the operation engaged, with consistent groupings in the representation space.

Clear Signatures at the Core of the Network, Beyond the Words Used

Reasoning operations are reliably distinguishable in the internal representations of the three models, and the separation peaks in the intermediate layers. The researchers ruled out an explanation based solely on word choice: a classifier considering only tokens performs worse than one that leverages internal representations. Similarly, the position of a segment in the solution path does not explain the separation. Therefore, internal states carry information about the nature of the reasoning step that goes beyond the textual surface, an effect particularly pronounced in the intermediate layers of the network.

Same Word, Different Role Depending on the Step and Context Dependence

Common terms such as "a," "is," or "the" appear at very different moments in the reasoning process. In the initial layers, their internal representations remain intertwined, but they become individualized in the intermediate and final layers based on the operation performed. Thus, the same term receives a distinct internal representation depending on the step. Through targeted intervention, the researchers blocked attention on the previous 30 tokens: the signal of the operation weakens, indicating a dependence on context. Finally, even when the problem is poorly solved, the executed step remains identifiable, including for an incorrect calculation that retains its internal calculation signature.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.