Brief IA

Google Lens Revolutionizes Visual Search with Multimodal AI

💡 Use Cases·Tom Levy·

Google Lens Revolutionizes Visual Search with Multimodal AI

Google Lens Revolutionizes Visual Search with Multimodal AI
Key Takeaways
1Google has improved Circle to Search and Lens to analyze multiple objects in an image.
2Dounia Berrada explains how AI simultaneously breaks down and searches for visual elements.
3The Gemini model powers these advancements by combining image analysis and web search.
💡Why it mattersThese innovations simplify visual search, making it easier to access information and inspiration in various contexts.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Google Lens Revolutionizes Visual Search with Multimodal AI

Visual search has recently reached a significant milestone thanks to innovations from Google, particularly with updates to Google Search. A Google expert sheds light on these advancements and the techniques employed to achieve them.

We've all experienced that moment: catching a glimpse of a beautifully decorated living room or a stylish street outfit and wanting to know every detail. Until recently, visual search was done element by element. However, a major update to Circle to Search and Lens now allows Google to break down and analyze multiple objects in a single image simultaneously. This means that if you use Circle to Search on Android for a complete outfit, you'll get results for each component of the look, rather than just one item at a time. Over the past few months, several updates have been rolled out to enhance visual search and image results in AI Mode, making it easier to find inspiration.

To better understand these advancements, we spoke with Dounia Berrada, Senior Director of Search Engineering at Google.

What aspect of search are you working on?

I focus on multimodal search, also known as Google Lens. The idea is to enable Google to answer your most complex questions regarding images, PDFs, and everything you see. Visual search is redefining our interaction with information; Lens needs to be smart enough to understand the "why" behind your search, making it easy to get help with what you see on your screen or in the world around you. This means building a tool capable of explaining a complex math problem just as easily as identifying a rare succulent or helping you find a pair of shoes you love.

How does it work?

Imagine you're redesigning a room and you upload a photo of a mid-century modern space for inspiration. You're probably not just looking for the side table; you want to recreate the entire vibe. Previously, you had to search for the lamp, then the rug, then the chair individually. Now, AI Mode can break down that complex image, identify each individual piece, and perform multiple visual searches simultaneously. You can see this in action right now by using Circle to Search.

What powers these types of visual search responses?

Our advanced Gemini models make AI Mode possible, and its multimodal capabilities benefit from the visual expertise we've integrated into Lens over the years. When you search with an image, Gemini analyzes the image in parallel with your question to decide which tools to use. Suppose you're scrolling through your phone and see an outfit on social media that you like. When you search for it, the model knows to use Lens to simultaneously retrieve image results for the hat, shoes, and jacket of the outfit. It then weaves these individual results into a single, easy-to-read response.

Think of it this way: the AI model acts as the "brain" capable of "seeing" the image, while the visual search backend acts as the "library" containing billions of web results. The AI performs multi-object reasoning to understand what you're looking at. It then uses a "fan-out" technique that triggers multiple searches at once, reads the results, and presents a unique and coherent answer with helpful links—all in a matter of seconds.

Can you explain the fan-out technique?

AI Mode essentially performs a dozen searches for you in record time. If you upload a photo of a garden you admire, you might have several questions: will these plants survive in the shade? Are they suitable for my climate? What is their maintenance? Previously, you had to ask these questions one by one. Now, AI Mode identifies all those necessary "fan-out" searches. This way, it gathers care requirements for each plant in the photo using helpful web results, breaks down the information, and even suggests the next steps you might want to take. As AI Mode discovers more visual results from a single search, it’s easier than ever to find exactly what you're looking for while uncovering something new that piques your interest.

Do you have to start with an image to get this type of help in AI Mode?

Not at all! You can start with a simple text search in AI Mode, like "visual inspiration for work outfits." When you see a result you like, you can simply say, "Show me more options like the second skirt." The system immediately takes that specific image and starts the fan-out process from there.

This seems ideal for shopping—what else could it be used for?

You could take a photo of a wall in a museum and ask for explanations about each painting. Or take a photo of a bakery display and ask what all the different pastries are. It’s about moving from "What is this thing?" to "Explain this whole scene to me."

It seems I have a few photos to take and much more to discover. I'm off to put these tools to the test!

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.