Vulnerabilities on Azure Following OpenAI's Ban

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI claims to have halted a large-scale distillation campaign targeting the thought processes of its models at the end of July. However, researchers note that protections are inconsistent across platforms: by mid-September, extraction was still possible on Microsoft Azure against several models, including GPT-6 Astra.
On Azure, extraction still worked in mid-September
On September 13, tests conducted by researchers showed that the attack was blocked on the APIs of OpenAI and Anthropic, but remained possible on Microsoft Azure. This allowed for the extraction of the reasoning from each tested OpenAI model, including GPT-6 Astra, as well as from Anthropic models up to Sonnet 5. A single attempt was sufficient to obtain the reasoning in text form. Researchers emphasize that the same models benefit from different protections depending on the platform hosting them. According to their timeline, OpenAI only added protective measures to the Azure endpoint starting September 27, and extraction was no longer reproducible on Azure for Anthropic models from September 28. They also report that GPT-6 Astra had been launched on third-party platforms without protections in place.
Two vectors: cross-decryption and virtual notebook
Attackers copied the encrypted reasoning from a conversation and submitted it in a separate session to another model for decryption and restitution. Joachim Schaeffer and his team explain that providers send these steps in the form of encrypted packets with shared keys, reusable across sessions, users, and models of the same family. A weaker and less expensive model can then serve as a "decryption oracle" and print out the hidden steps verbatim. A second method, publicly presented by developer Can Bölük, involves providing the model with a virtual notebook and asking it to write its reasoning there, readable by the user. According to researchers, this method worked on all tested OpenAI models, as well as on Opus 4.8 and Sonnet 5. Only Opus 5, Fable 5, and Fable 5.1 did not reveal their reasoning. The output obtained through this method resembled that of the decryption attack and would likely be just as useful for distillation.
OpenAI strengthens its defenses and shares its lessons
OpenAI credits the researchers and claims to have confirmed the reality of the attack paths described, which allowed it to deploy countermeasures more quickly. The company states that it has banned fraudulent accounts, strengthened registrations, and closed the possibility of reusing and reading encrypted reasoning not originating from the user. It now continuously filters outputs and retains them if they risk revealing reasoning steps. OpenAI specifies that it has shared its lessons through the Frontier Model Forum and government channels, while acknowledging that models hosted by partners must have equivalent protections and that the work is not finished.
A massive campaign dismantled at the end of July
OpenAI states that it has ended an adversarial distillation campaign aimed at extracting the internal steps of its models. In this type of operation, a model learns from the complete output of another, including its thought processes, which the company considers sensitive as they may contain information excluded from responses and facilitate the reproduction of capabilities. The activity, which appeared at low volume on July 1, peaked on July 24 and 25 with 16,000 requests from over 4,000 users, relying on a typical extraction model. OpenAI reports having identified a network of over 15,000 related accounts and shut it down on July 28, clarifying that these were extraction attempts, without guaranteeing their success.
Partial attribution and calls for rules for clouds
OpenAI links a core activity to individuals associated with Moonshot AI, the developer of the Kimi model, while clarifying that it is not established that all actors come from a single source. Anthropic recently reported similar attempts made by Chinese AI companies. Researchers have published an update indicating that they have successfully extracted reasoning again, with Joachim Schaeffer emphasizing that securing one's own API is not enough if cloud providers distribute the same models with lesser protections. They describe still-fragmentary fixes, often based on a fragile matching of request patterns, and sometimes deployed late on cloud platforms. Schaeffer calls for protections covering every type of attack and every hosting cloud. The team goes further by arguing that clouds that do not impose equivalent protections should not be allowed to serve reasoning models and warns that open backdoors would allow circumventing export controls at the API level. Finally, researchers report that the trick remained effective for weeks on Azure.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.