⚡
Brief IA
›

EmbeddingGemma 2: Multimodal Embeddings on Device

🔬 Research·Tom Levy·

EmbeddingGemma 2: Multimodal Embeddings on Device

EmbeddingGemma 2: Multimodal Embeddings on Device
⚡
Key Takeaways
1EmbeddingGemma 2, a multimodal embedding model with 740 million parameters, is available on Hugging Face and Kaggle
2The model is touted as optimized for local execution, with a memory footprint of 191 MB (text) to 567 MB (multimodal) on Pixel 11 Pro
3Claimed top scores on MTEB Code and MAEB, with an 8K token context window and modular options
4Public demonstrations and numerous deployment tools are already available
💡Why it matters — EmbeddingGemma 2 aims to facilitate multimodal search and retrieval on-device, combining performance, modularity, and privacy.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

EmbeddingGemma 2, a multimodal embedding model with 740 million parameters under Apache 2.0 license, is now available for local deployment. The publisher claims top performance under one billion parameters and reports a memory consumption of approximately 191 MB for text-only and 567 MB for the full version on Pixel 11 Pro.

Integrations, Demos, and Tools Already Available

The model weights are accessible on Hugging Face and Kaggle, with availability announced as forthcoming in the Gemini Enterprise Agent Platform Model Garden. The LiteRT community on Hugging Face offers device-optimized variants. For development, Google AI Edge MediaPipe supports ready-to-use tasks for embedding, retrieval, and decision-making, while LiteRT aims for the integration of custom models. On the web side, execution is possible with transformers.js or WebGPU.
The model can be served with tools like transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio, and vectors can be stored with Qdrant. Unsloth provides tips for fine-tuning, along with a developer guide, documentation, and inference and fine-tuning guides. Public demonstrations are available in the Google AI Edge Gallery, including instant media search and Video Moments Finder, as well as an AI Edge Foresight application.

Local Resources and Expanded Context Window

The model is touted as optimized for strict on-device constraints. After quantization, the RAM used on a Google Pixel 11 Pro is approximately 191 MB for text weights and about 567 MB for the complete multimodal version. The context window now offers 8K tokens, a fourfold increase compared to the previous generation, allowing for the handling of up to 5.5 minutes of audio, 29 images, 58 video frames, or local assemblies of these elements.
The model's modular structure allows for the use of only 270 million parameters for text, while optional encoders of 170 million are planned for vision and 300 million for audio. For storage, the Matryoshka approach allows truncation from 768 dimensions to 512, 256, or 128, with a reported storage reduction of up to six times.

Claimed Scores on MTEB Code and MAEB

The publisher claims that the model achieves top scores among models with less than one billion parameters on benchmarks such as MTEB Code and MAEB, while matching or exceeding larger systems on text, vision, and audio tasks. For code, the reported gain is 9.92 points in MTEB Code, from 68.76 to 78.68, while maintaining the multilingual text performance of the previous generation.
In the fields of image, video, documents, and audio, the publisher presents the model as setting a new quality standard per parameter under one billion and surpassing some specialized models more than twice its size. The overall package is described as the highest-performing for on-device multimodal embeddings.

Technical Foundation, License, and Multimodal Scope

The model is based on the Gemma 4 architecture, totals 740 million parameters, and is distributed under a permissive Apache 2.0 license for commercial purposes. It shares the text tokenizer and audio encoder with Gemma 4, both of which can be executed together in a unified pipeline announced with a reduced combined memory footprint.
The embedding space natively covers text, images, audio, and video, with positioning presented as optimal for on-device inference. The first generation, launched the previous year for textual embeddings, saw significant usage according to the publisher, with over 20 million downloads, applications in on-device research, and privacy-focused RAG pipelines.

On-Device Usage, Cross-Search, and RAG

The model brings capabilities locally on edge hardware. According to the publisher, generating embeddings on-device helps preserve privacy, reduces latency, and enables offline cross-search and retrieval. When combined with generative models like Gemma 4, it is said to empower on-device RAG pipelines capable of encompassing complex multimodal data.
Described use cases include semantic similarity search in a multimedia library from text or image, pinpointing specific moments in a video from a text or audio query, and combining with Gemma 4 for local retrieval and contextual reasoning. The publisher also cites examples such as searching for a video clip from a voice memo or text queries on audio recordings, all performed by a single multimodal model.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.