Grok: Data Leak via Encrypted Injection, xAI Alerted

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
A separate team managed to extract conversations and other personal information from Grok by hiding malicious commands. The exposure was ongoing at the time of publication, even though xAI had been notified in June. This scenario follows a similar case involving Microsoft 365 Copilot earlier in the week.
In the face of injections, LLMs execute hidden commands and rely on safeguards
Prompt injections exploit the tendency of models to respond to user queries whenever possible. Attackers sneak malicious instructions into emails or web pages that the assistant is supposed to summarize. The models do not reliably differentiate this content from instructions typed directly by the user and comply with these hidden commands. To date, the only defense described for Grok and other assistants relies on safeguards designed to detect these suspicious commands and block their execution.
On Grok, exfiltration persists despite a report to xAI in June
At the time of publication, Grok continued to disclose data even though xAI had been informed in June. A separate team devised an attack that forces the model, owned by Elon Musk and operated by xAI, to extract user discussions and other personal information. The exfiltration occurs when malicious instructions are encrypted, using a trick presented as very simple.
Microsoft 365 Copilot experienced a precedent earlier this week
Earlier in the week, researchers detailed an attack against Microsoft 365 Copilot for businesses. It relied on a secret input provided by the assistant and allowed for the extraction of a password found in the user's inbox.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.