Brief IA

Google Revolutionizes Enterprise AI with Gemini 3.6 Flash

🛠️ AI Tools·Tom Levy·

Google Revolutionizes Enterprise AI with Gemini 3.6 Flash

Google Revolutionizes Enterprise AI with Gemini 3.6 Flash
Key Takeaways
1Google has launched Gemini 3.6 Flash and 3.5 Flash-Lite to optimize costs and latency for AI agents in enterprises.
2Gemini 3.6 Flash reduces output tokens by 17% compared to the previous version, enhancing efficiency.
3Figma, Hebbia, and Harvey are integrating Gemini 3.6 Flash to accelerate prototyping and document analysis.
💡Why it mattersThese innovations enable companies to reduce operational costs while increasing the performance of AI agents.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Google Introduces Gemini 3.6 Flash for Enterprises

Google has recently unveiled two new artificial intelligence models, Gemini 3.6 Flash and 3.5 Flash-Lite, specifically designed to meet the needs of businesses in terms of reducing latency and the costs associated with tokens. These models aim to optimize the efficiency of AI agents used in professional environments.

In the context of autonomous software agents, operational economics relies on a complex equation. Few providers directly address this issue, but it is essential for a model to handle complex multi-step tasks while minimizing costs and delays. Each additional token generated by the model increases expenses and slows down the process, which can be problematic when these operations are repeated thousands of times per hour.

Teams developing background agents, rather than chat interfaces, emphasize throughput and number of parameters. Google has responded to these needs by introducing three distinct models: Gemini 3.6 Flash for coding and multimodal reasoning, Gemini 3.5 Flash-Lite for high-volume tasks requiring low latency, and Gemini 3.5 Flash Cyber, a restricted model for vulnerability correction.

Performance and Savings with Gemini 3.6 Flash

Google's developer documentation highlights a key figure for 3.6 Flash: a 17% reduction in output tokens compared to the previous version 3.5 Flash, according to the Artificial Analysis Index. This improvement translates into increased efficiency and reduced costs.

In synthetic testing, particularly the Datacurve DeepSWE benchmark, Google observed reductions in token usage of up to 65%. The pricing is set at $1.50 for one million input tokens and $7.50 for one million output tokens, making the model suitable for continuous reasoning loops.

On the DeepSWE benchmark, 3.6 Flash achieved a success rate of 49%, surpassing the 37% of its predecessor. Additionally, on MLE Bench, the score increased from 49.7% to 63.9%, and on Google's GDPval-AA v2 test, the model reached a score of 1421, compared to 1349 for the earlier version.

Integration by Figma, Hebbia, and Harvey

Figma has integrated 3.6 Flash into its prototyping infrastructure. According to Matt Colyer, product engineering director at Figma, this model allows developers to iterate through design iterations more quickly without sacrificing the quality of results.

In the legal field, the platform Harvey and the research tool Hebbia utilize this model to process multimodal document data. This includes ingesting financial filings, analyzing document structures, reading embedded charts, and generating preliminary reports for review.

Google has also integrated a client-side computing usage tool directly into the Gemini API and Gemini Enterprise platforms, eliminating the need for custom intermediary software. This integration has achieved a verified score by OSWorld of 83.0%, up from 78.4%. Furthermore, new protections against chemical, biological, radiological, and nuclear abuse enhance resistance to jailbreak attempts without increasing refusals for legitimate requests.

Gemini 3.5 Flash-Lite: An Economical Solution for High Volumes

The Gemini 3.5 Flash-Lite model is designed to handle documents and perform agentic searches at high volume, rather than in-depth reasoning. The Artificial Analysis Index measured this model at 350 output tokens per second, the fastest in the 3.5 series.

Pricing is set at $0.30 for one million input tokens and $2.50 for one million output tokens, which is low enough to allow engineering teams to handle simple high-volume requests while reserving higher levels of reasoning for complex tasks.

During Google's long-context test GDM-MRCR v2, Gemini 3.5 Flash-Lite recorded a success rate of 72.2%, compared to 60.1% for its predecessor. Its GDPval-AA v2 score nearly doubled, rising from 642 to 1140. The model benefits from the same native computing usage tool as 3.6 Flash.

Additionally, Google indicates that Gemini 3.5 Pro is still in testing with partners ahead of a full launch, and pre-training for the upcoming Gemini 4 architecture is already underway.

Gemini 3.5 Flash Cyber: A Specialized Model for Security

Automated vulnerability scanners now identify flaws faster than most security teams can fix them. It is in this gap that Google positions Gemini 3.5 Flash Cyber.

This model is designed to validate and correct code vulnerabilities. Google reports competitive performance on the CyberGym benchmark, although specific details have not been disclosed with the same level of precision as for consumer-targeted versions.

The distribution of 3.5 Flash Cyber is limited to governments and verified partners as part of a pilot program, a measure that Google presents as protection against the offensive use of generated code.

In Google's security agent CodeMender, multiple instances of 3.5 Flash Cyber operate in parallel, verifying each other's results before producing a single remediation report that a human reviewer validates.

Engineering teams interested in these new models can access them via the Gemini API through Google AI Studio, Android Studio, or the Gemini Enterprise Agent Platform. Consumers can also access the new models in the Gemini app, and 3.5 Flash-Lite is also being rolled out in Google Search.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.