Mistral AI Launches Shieldstral, Open Source Moderation Model

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Mistral AI Introduces Shieldstral, an AI Model for Content Moderation
The French company Mistral AI has recently launched Shieldstral, an artificial intelligence model designed for the moderation of potentially harmful content, whether textual or visual. This model is available to the public for free and immediately.
An Adaptable and Personalized Moderation System
Shieldstral is described by Mistral AI as a multimodal safety classifier with three billion parameters. It is designed to moderate various types of content using a binary question-answering method that results in simple "yes" or "no" responses.
In practice, Shieldstral analyzes a given piece of content and determines whether it poses a risk, for example, by promoting violence or being inappropriate for a young audience. Functioning like a filter, similar to anti-spam systems, it assigns a risk score to each piece of content, allowing the company to decide on the action to take: allow, block, or submit to a moderator.
What sets Shieldstral apart from other moderation tools is its flexibility. Unlike traditional systems that rely on a predefined rule list requiring complex retraining to be modified, Shieldstral allows users to define monitoring criteria using natural language questions at the time of use.
Shieldstral's approach enables questions such as "Does this content incite violence?" or "Is this image appropriate for a minor?". The model then evaluates the content and generates a safety score based on the probabilities of "yes" and "no" responses.
With this method, the same model can be applied to different contexts without the need for retraining. Thus, content acceptable in a cybersecurity research context might be deemed inappropriate on a platform intended for the general public, depending on the defined policies. "Every product integrating a model must answer this type of question; the appropriate response depends on the product, target audience, and context," Mistral AI emphasizes.
The Structure of a Shieldstral Query
According to Mistral AI, each query addressed to the Shieldstral model consists of three elements:
- The evaluation context, which determines the severity level and what should be considered dangerous.
- The closed question, to which the model responds with "yes" or "no," such as "Does this content incite violence?".
- The content to be examined, whether it is text, an AI-generated response, or an image accompanied by text.
After analysis, the model provides a risk score. It is then up to the company to define the threshold for action and the steps to take.
Mistral AI claims that Shieldstral outperforms moderation tools up to seven times larger, according to the benchmarks they have published.
An Open Source Model Under Apache 2.0 License
Shieldstral is distributed in open weights under the Apache 2.0 license, making it compatible for commercial use.
The model is downloadable from the Hugging Face platform and is part of Mistral AI's open source approach, already applied to its other models like Devstral and the latest generation Mistral 3, all also available under the same license.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.