Gemini 3.1 Flash-Lite: Google's Fast and Affordable AI

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Introduction of Gemini 3.1 Flash-Lite
Google has announced the launch of Gemini 3.1 Flash-Lite, the fastest and most cost-effective model in the Gemini 3 series. This model is specifically designed for developers handling large workloads. It is now available in a preliminary version for developers via the Gemini API in Google AI Studio and for businesses through Vertex AI.
A Cost-Effective Model Without Compromising Performance
The cost of Gemini 3.1 Flash-Lite is particularly attractive, with a rate of $0.25 per million input tokens and $1.50 per million output tokens. This model offers improved performance compared to the 2.5 Flash, with an initial response time 2.5 times faster and a 45% increase in output speed, according to the benchmark Artificial Analysis. Despite its lower cost, it maintains similar or superior quality, making it ideal for developers looking to create responsive, real-time experiences.
Impressive Performance
Gemini 3.1 Flash-Lite boasts an Elo score of 1432 on the Arena.ai ranking, surpassing other models of similar caliber in reasoning and multimodal understanding benchmarks. It achieves 86.9% on GPQA Diamond and 76.8% on MMMU Pro, even outpacing some larger Gemini models from previous generations like the 2.5 Flash.
Adaptive Intelligence at Scale
Beyond its raw performance, the 3.1 Flash-Lite comes with levels of reasoning in AI Studio and Vertex AI, providing developers with the control and flexibility needed to choose the model's reasoning level for a task. This is essential for managing high-frequency workloads. The 3.1 Flash-Lite can tackle large-scale tasks, such as high-volume translation and content moderation, where cost is a priority. It can also handle more complex workloads requiring in-depth reasoning, such as generating user interfaces and dashboards, creating simulations, or following instructions.
Real-World Applications and User Feedback
The 3.1 Flash-Lite is capable of instantly populating an e-commerce framework with hundreds of products across various categories. It can generate dynamic weather dashboards in real-time, using live forecasts and historical data. Additionally, it can create a SaaS agent capable of executing versatile, multi-step tasks for a business, and quickly analyze and sort large amounts of content, such as images.
Early access developers on AI Studio and Vertex AI, as well as companies like Latitude, Cartwheel, and Whering, are already using the 3.1 Flash-Lite to solve complex large-scale problems. Early testers have highlighted the efficiency and reasoning capabilities of the 3.1 Flash-Lite, stating that it can handle complex inputs with the precision of a higher-tier model while following instructions and maintaining compliance.
Google is eager to see the innovations that developers will create with the 3.1 Flash-Lite and other models in the Gemini 3 series.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.