Gemini 3.6 Flash: Google Optimizes Quietly in 2026

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Gemini 3.6 Flash: Google Optimizes Quietly in 2026
What Has Really Changed
Gemini 3.6 Flash is an incremental successor to Gemini 3.5 Flash, launched at the I/O event in May. Google's presentation is particularly candid: this version takes into account feedback from developers and customers and primarily aims to be more token-efficient in task execution. Same level, but with tighter savings.
-
About 17% fewer output tokens according to the Artificial Analysis Index, with reductions reaching 65% on certain DeepSWE executions due to fewer reasoning steps and tool calls.
-
Improved production code. Increased accuracy with fewer unwanted edits, supported by DeepSWE (49% vs. 37%) and MLE Bench (63.9% vs. 49.7%).
-
Enhanced computational usage. The OSWorld-Verified score rises from 78.4% to 83% for click-and-tap agent workflows.
-
A refreshed memory. The knowledge cutoff date finally moves from January 2025 to March 2026.
-
Same envelope. 1 million tokens context; text, image, speech, and video as input, text as output; configurable reasoning effort and parallel tool usage.
Notice what's not on this list: a major leap in raw intelligence. On the Artificial Analysis Intelligence Index, it scores around 50, which is above the industry average but roughly stable compared to 3.5 Flash. Google has optimized the denominator (cost per useful response), not the numerator (maximum reasoning).
Pricing and the Flash Family
Google launched three elements simultaneously, all under the Flash umbrella. All these models are available for free on the Gemini app or web app. Regarding API pricing, here’s their positioning:
The Flash Family
-
Gemini 3.6 Flash
- Main version focused on balanced coding, knowledge, and multimodal workflows
- $1.50 in / $7.50 out
-
Gemini 3.5 Flash-Lite
- High throughput, low latency: agentic search, document processing
- $0.30 in / $2.50 out
-
Gemini 3.5 Flash Cyber
- Finds, validates, and fixes vulnerabilities
- Limited access in pilot phase
The real stakes in pricing are here. 3.6 Flash reduces the output rate from $9 to $7.50 per million tokens while keeping the input at $1.50.
Numbers: 3.6 vs. 3.5 at a Glance
Comparison from Google, level by level. Each figure is 3.6 Flash measured against 3.5 Flash that it replaces.
-
DeepSWE
- Production-ready code: 37% vs. 49%
-
MLE Bench
- ML research tasks: 49.7% vs. 63.9%
-
GDPval-AA
- Real-world knowledge work: 1349 vs. 1421
-
OSWorld-Verified
- Computational/agentic usage: 78.4% vs. 83%
-
Output tokens / task
- Efficiency (AA Index): 100% vs. ~83%
-
AA Intelligence Index
- Composite reasoning: ~50 vs. ~50
-
Knowledge cutoff date
- Recency of training data: January 2025 vs. March 2026
-
Output price / 1M
- Cost: $9.00 vs. $7.50
The pattern is consistent: real gains in applied work (coding, ML tasks, computational usage) while the main composite score barely shifts. It’s a model tuned to do known things more economically, not to unlock new ones.
Practical: Model Stress Test
Specifications are promises; your prompts are proof. Here are six tests to copy-paste, each targeting a specific claim. Run them in the Gemini app, AI Studio, or your IDE. Each test indicates what a solid response looks like and the exact failure to watch for.
-
Visual Discipline Test
- Prompt: “Reconstruct the underlying data as a table (label, value). Then, name ONE specific way this graph could mislead a reader by pointing out the exact axis or scale choice. Mark any value you had to estimate with a ~.”
- Observation: success — the best execution to date.
-
Test Case Verification
- Prompt: “This function should return the 2 highest values from a list, but it fails on lists with duplicates. Fix ONLY the root cause. Do not rename anything, add comments, or reformat unrelated lines. Return the diff, then explain the bug in one sentence.”
- Observation: inappropriate response — it’s a partial fix that fails the test.
-
Canvas Building and Iteration
- Prompt: “Build an ‘image palette extractor’ in a single file, without a library: I drop an image, it extracts the 6 dominant colors, shows each as a sample with its hex code, and copies the hex when I click on a sample. Then, add a ‘download palette as PNG’ button. Keep iterating in Canvas until the download actually produces a file — don’t tell me it’s done until it works.”
- Observation: incredible! The tool is not only functional but was created in under a minute.
-
Instruction Following
- Prompt: “Two-part question, answer both precisely. Name a significant technology or global event from the second half of 2025 and briefly say what happened. Tell me who won an event that took place in June 2026. If your training data does not reliably cover this, say so explicitly instead of guessing.”
- Observation: correct but evasive!
-
Contradiction Hunt
- Prompt: “Two clauses here contradict each other on the same subject. Find BOTH, quote the conflicting lines with their section numbers, and tell me which one prevails by applying any stated priority rules in the document. Do not summarize the agreement — simply resolve the conflict.”
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.