NVIDIA Nemotron 3 Embed: Leader in Agentic Recovery

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
The Crucial Importance of Retrieval in Agentic Systems
In agentic systems, information retrieval is a critical step that can determine the success or failure of a process. Ineffective retrieval can lead to contextual errors, redundant queries, and resource wastage, complicating subsequent reasoning stages. To address these challenges, NVIDIA has introduced the Nemotron 3 Embed, a series of embedding models designed to optimize information retrieval.
Overview of Nemotron 3 Embed
The Nemotron 3 Embed is a collection of open and commercially available models that stand out for their ability to enhance the retrieval quality of agentic systems. These models are designed to provide developers with flexible options for deploying large-scale retrieval systems, particularly for code retrieval and agent memory management.
The collection consists of three main models, each with its own features and advantages. The flagship model, Nemotron-3-Embed-8B-BF16, is an 8 billion parameter model that ranks first in the RTEB leaderboard. It is accompanied by two 1 billion parameter variants, designed for large-scale deployment while maintaining high efficiency.
- Nemotron-3-Embed-8B-BF16: This model is recognized for its high accuracy and is ideal for enterprise applications where information retrieval is critical.
- Nemotron-3-Embed-1B-BF16: This model is optimized for production, offering a balance between cost and latency.
- Nemotron-3-Embed-1B-NVFP4: Optimized for high-throughput infrastructures, this model uses hardware acceleration technology to reduce memory footprint.
Advanced Features for Enterprises
Beyond its performance on RTEB, Nemotron 3 Embed offers production-ready features tailored to the needs of enterprises. These features include open weights and datasets, allowing teams to customize and deploy the models on their own infrastructure. The 32k context window is another notable feature, enabling retrieval over long documents and large code contexts while minimizing truncation.
The models also support multilingual and code retrieval, adapting to the needs of global enterprises. Deployment efficiency is enhanced by the use of NVIDIA NVFP4 technology, which allows for high-throughput retrieval with a reduced memory footprint.
Performance Evaluation: Quality and Efficiency
The evaluation of Nemotron 3 Embed models is based on three main axes: retrieval quality, agentic efficiency, and deployment trade-offs. The 8 billion parameter model sets a new standard for quality, while the 1 billion parameter variants offer cost-effective, high-throughput solutions.
Dominance on RTEB and Other Benchmarks
In tests on RTEB, the Nemotron-3-Embed-8B-BF16 achieved an impressive score of 78.5%, ranking first. The models were also tested on other benchmarks such as ViDoRe V3 Text and MMTEB Retrieval, where they demonstrated exceptional performance, scoring 75.5% on MMTEB Retrieval.
The Nemotron-3-Embed-1B-BF16 model managed to retain much of the retrieval quality of the 8 billion parameter model, achieving a score of 72.4% on RTEB, representing a 27% reduction in error rate compared to its 1 billion predecessor (llama-nemotron-embed-vl-1b-v2). On MMTEB Retrieval, it reached a score of 71.0%, reducing the error rate by 28%.
Importance of Enhanced Retrieval for Agents
Improved information retrieval is essential for agents, as it allows for the provision of relevant evidence more quickly, thus avoiding repeated searches and unnecessary reasoning cycles. Using a search agent powered by Nemotron 3 Ultra, tests have shown that enhanced retrieval reduces costs in agent tokens and improves overall efficiency.
Optimization for High-Throughput Deployments
For deployments requiring high throughput, the Nemotron-3-Embed-1B-NVFP4 offers an ideal solution. With native NVFP4 acceleration on NVIDIA Blackwell architectures, this model successfully combines service efficiency and retrieval quality.
Optimized Performance from Day One
To ensure optimal performance at production scale, NVIDIA has launched a NIM microservice optimized for the 1 billion parameter model. This service ensures consistent efficiency under real query loads, regardless of input sequence length.
Design of Nemotron 3 Embed Models
The Nemotron-3-Embed-8B-BF16 model is based on the Ministral-3-8B-Instruct-2512 backbone, modified to become a bidirectional encoder. Trained with contrastive pre-training, it uses text pairs from various sources to refine its retrieval capabilities.
Reduction and Optimization to 1 Billion
The 1 billion parameter model is not simply a reduced version. It has been developed by compressing a 3 billion parameter model through advanced pruning and distillation techniques, ensuring high retrieval accuracy despite its smaller size.
The table below summarizes the technical specifications and deployment targets of the Nemotron 3 Embed models:
- Nemotron-3-Embed-8B-BF16: Optimized for general GPU inference.
- Nemotron-3-Embed-1B-BF16: Designed for low latency on CPU/GPU.
- Nemotron-3-Embed-1B-NVFP4: Suitable for NVIDIA Blackwell/GB200 infrastructures.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.