Brief IA

Mistral AI Launches Shieldstral, Open Source Moderation Model

🛠️ AI Tools·Tom Levy·

Mistral AI Launches Shieldstral, Open Source Moderation Model

Mistral AI Launches Shieldstral, Open Source Moderation Model
Key Takeaways
1Mistral AI has launched Shieldstral, an open-source AI model for moderating risky texts and images.
2Shieldstral uses a binary question-answering system to evaluate content and assign a risk score.
3The model is available under the Apache 2.0 license, allowing commercial use via Hugging Face.
💡Why it mattersShieldstral provides a flexible and accessible solution for content moderation, suitable for various contexts and industrial needs.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Mistral AI Introduces Shieldstral, an AI Model for Content Moderation

The French company Mistral AI has recently launched Shieldstral, an artificial intelligence model designed for the moderation of potentially harmful content, whether textual or visual. This model is available to the public for free and immediately.

An Adaptable and Personalized Moderation System

Shieldstral is described by Mistral AI as a multimodal safety classifier with three billion parameters. It is designed to moderate various types of content using a binary question-answering method that results in simple "yes" or "no" responses.

In practice, Shieldstral analyzes a given piece of content and determines whether it poses a risk, for example, by promoting violence or being inappropriate for a young audience. Functioning like a filter, similar to anti-spam systems, it assigns a risk score to each piece of content, allowing the company to decide on the action to take: allow, block, or submit to a moderator.

What sets Shieldstral apart from other moderation tools is its flexibility. Unlike traditional systems that rely on a predefined rule list requiring complex retraining to be modified, Shieldstral allows users to define monitoring criteria using natural language questions at the time of use.

Shieldstral's approach enables questions such as "Does this content incite violence?" or "Is this image appropriate for a minor?". The model then evaluates the content and generates a safety score based on the probabilities of "yes" and "no" responses.

With this method, the same model can be applied to different contexts without the need for retraining. Thus, content acceptable in a cybersecurity research context might be deemed inappropriate on a platform intended for the general public, depending on the defined policies. "Every product integrating a model must answer this type of question; the appropriate response depends on the product, target audience, and context," Mistral AI emphasizes.

The Structure of a Shieldstral Query

According to Mistral AI, each query addressed to the Shieldstral model consists of three elements:

  • The evaluation context, which determines the severity level and what should be considered dangerous.
  • The closed question, to which the model responds with "yes" or "no," such as "Does this content incite violence?".
  • The content to be examined, whether it is text, an AI-generated response, or an image accompanied by text.

After analysis, the model provides a risk score. It is then up to the company to define the threshold for action and the steps to take.

Mistral AI claims that Shieldstral outperforms moderation tools up to seven times larger, according to the benchmarks they have published.

An Open Source Model Under Apache 2.0 License

Shieldstral is distributed in open weights under the Apache 2.0 license, making it compatible for commercial use.

The model is downloadable from the Hugging Face platform and is part of Mistral AI's open source approach, already applied to its other models like Devstral and the latest generation Mistral 3, all also available under the same license.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.