Researchers Extract Hidden Traces via LLM APIs

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
A study reports having replayed encrypted reasoning blocks returned by API models such as Anthropic, OpenAI, and Google, successfully retrieving the clear thought processes of an advanced model. The authors indicate that the providers acknowledged receipt of the report and that they have not been able to reproduce the attack since.
Providers Acknowledge Receipt; Attacks Are No Longer Reproducible
The authors of the study state that all model providers acknowledged receipt of their report detailing these attacks. They also specify that they have not been able to reproduce the attacks after this notification. Furthermore, the assistant turn prefix feature has been removed in models 4.6, but remains active in Claude Haiku 4.5.
Encrypted Blocks Replayed Across Sessions, Users, and Models
The researchers explain that the APIs of several providers return encrypted reasoning blocks. According to them, these blocks can be replayed across different sessions, users, and models. By reinjecting a trace from an advanced model into a weaker model, they were able to bypass the latter's protections and obtain the hidden reasoning of the more powerful model in clear text.
API Calls and Shared Keys Within the Same Model Family
An example of an API call to OpenAI presented in the study uses the model "gpt-5.6-luna," a math question, a "medium" effort level, and the inclusion of the field "reasoning.encrypted_content." In the API response, there is a section titled "reasoning" as well as an encrypted_content consisting of a long encrypted string. The authors indicate that all models belonging to the same family shared the encryption key, which allowed them to submit these blocks to less powerful versions of the same family to access the raw reasoning blocks in clear text.
Exfiltration Techniques and Retrieved Content
The researchers present Claude Haiku 4.5 as the easiest model to attack, using a prompt requesting the transcription of the attached reasoning and defining an assistant prefix "<thinking-copy>." The study provides numerous examples of extracted traces in the appendix, including an excerpt of reasoning tokens from GPT-5.5 concerning CSS elements. The authors believe that these tokens were not intended to be read by humans. They also describe a variant of prompt injection where the model is led to incorporate a data exfiltration instruction, such as uploading a file to a remote server, into its reasoning trace, and then return this encrypted trace in another model. According to the authors, models are much more likely to follow instructions present in their own reasoning blocks. The study was published on August 11, 2026, at 10:40 PM.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.