⚡
Brief IA
›

Researchers Extract Hidden Traces via LLM APIs

🛠️ AI Tools·Tom Levy·

Researchers Extract Hidden Traces via LLM APIs

Researchers Extract Hidden Traces via LLM APIs
⚡
Key Takeaways
1Researchers have successfully replayed encrypted reasoning blocks between models to bypass protections and obtain clear thought chains.
2Within the same family of models, a shared encryption key allowed these blocks to be sent back to weaker models, resulting in raw reasoning.
3Providers acknowledged receipt of the report, and the authors indicate they can no longer reproduce the attacks.
💡Why it matters — The study details a variant of injection where instructions embedded in reasoning traces are much more likely to be followed by the models.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

A study reports having replayed encrypted reasoning blocks returned by API models such as Anthropic, OpenAI, and Google, successfully retrieving the clear thought processes of an advanced model. The authors indicate that the providers acknowledged receipt of the report and that they have not been able to reproduce the attack since.

Providers Acknowledge Receipt; Attacks Are No Longer Reproducible

The authors of the study state that all model providers acknowledged receipt of their report detailing these attacks. They also specify that they have not been able to reproduce the attacks after this notification. Furthermore, the assistant turn prefix feature has been removed in models 4.6, but remains active in Claude Haiku 4.5.

Encrypted Blocks Replayed Across Sessions, Users, and Models

The researchers explain that the APIs of several providers return encrypted reasoning blocks. According to them, these blocks can be replayed across different sessions, users, and models. By reinjecting a trace from an advanced model into a weaker model, they were able to bypass the latter's protections and obtain the hidden reasoning of the more powerful model in clear text.

API Calls and Shared Keys Within the Same Model Family

An example of an API call to OpenAI presented in the study uses the model "gpt-5.6-luna," a math question, a "medium" effort level, and the inclusion of the field "reasoning.encrypted_content." In the API response, there is a section titled "reasoning" as well as an encrypted_content consisting of a long encrypted string. The authors indicate that all models belonging to the same family shared the encryption key, which allowed them to submit these blocks to less powerful versions of the same family to access the raw reasoning blocks in clear text.

Exfiltration Techniques and Retrieved Content

The researchers present Claude Haiku 4.5 as the easiest model to attack, using a prompt requesting the transcription of the attached reasoning and defining an assistant prefix "<thinking-copy>." The study provides numerous examples of extracted traces in the appendix, including an excerpt of reasoning tokens from GPT-5.5 concerning CSS elements. The authors believe that these tokens were not intended to be read by humans. They also describe a variant of prompt injection where the model is led to incorporate a data exfiltration instruction, such as uploading a file to a remote server, into its reasoning trace, and then return this encrypted trace in another model. According to the authors, models are much more likely to follow instructions present in their own reasoning blocks. The study was published on August 11, 2026, at 10:40 PM.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.