Brief IA

Nvidia Nemotron 3 Nano Omni: A Revolution in Multimodal Models

🤖 Models & LLM·Tom Levy·

Nvidia Nemotron 3 Nano Omni: A Revolution in Multimodal Models

Nvidia Nemotron 3 Nano Omni: A Revolution in Multimodal Models
Key Takeaways
1Nvidia has unveiled Nemotron 3 Nano Omni, an open-source multimodal AI model capable of processing text, images, video, and audio.
2The model uses 30 billion parameters and has been trained with 717 billion tokens, incorporating data from competing models.
3Nemotron 3 Nano Omni outperforms its predecessor and competes with Qwen3-Omni across several benchmarks, with notable accuracy on OSWorld.
💡Why it mattersThis open-source model represents a significant advancement in multimodal AI, offering extensive capabilities for various commercial applications.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Nvidia recently launched the Nemotron 3 Nano Omni, an artificial intelligence model that stands out for its ability to simultaneously process text, images, video, and audio. This model is designed for agentic applications and is open for commercial use.

The Nemotron 3 Nano Omni was trained on an impressive volume of 717 billion tokens. A significant portion of this training data comes from competing models such as Qwen, gpt-oss, and DeepSeek-OCR. In addition to providing the model weights, Nvidia has also made some training data and associated pipelines public.

This open-source model, which is based on a hybrid Mamba-Transformer architecture with Mixture-of-Experts, activates about three billion parameters per query. It uses Nvidia's C-RADIOv4-H vision encoder and the Parakeet-TDT audio encoder, with a context window that can reach up to 256,000 tokens. Currently, English is the only officially supported language.

According to the technical report, the Nemotron 3 Nano Omni is primarily intended for applications such as document processing, computer usage agents, video and audio analysis, as well as voice interaction. On benchmarks like OCRBenchV2, MMLongBench-Doc, WorldSense, and VoiceBench, it outperforms its predecessor, the Nemotron Nano V2 VL, and competes with Alibaba's Qwen3-Omni. On OSWorld, a benchmark for GUI agents, its accuracy jumped from 11.1 to 47.4 points compared to the previous version. Nvidia claims that its throughput is up to nine times higher than that of Qwen3-Omni.

Influence of Competing Models on Training

The training of the Nemotron 3 Nano Omni was conducted in seven stages, with a progressively widening context window. The synthetic training data includes image captions, question-answer pairs, and reasoning traces generated from models like Qwen3-VL-30B-A3B-Instruct, Qwen3.5-122B-A10B, and OpenAI's gpt-oss-120b. Nvidia also utilized GPT-4o and Gemini 3 Flash Preview for filtering.

Using competing models for training is a common practice in the industry, although rarely as transparent. Companies like OpenAI, Anthropic, and Google have often accused Chinese labs of large-scale distillation.

The audio data includes Nvidia's Granary and SIFT-50M datasets, as well as captions from Qwen's Omni-Captioner. For reinforcement learning, a five-step pipeline has been established, covering 25 environments and various tasks such as visual anchoring and automatic speech recognition.

In addition to weights in BF16, FP8, and NVFP4, Nvidia is releasing parts of the training data, the pipelines on Megatron-Bridge, and the RL recipes on NeMo-RL. The reasoning mode is enabled by default, requiring manual deactivation for certain tasks. The model is distributed under the NVIDIA Open Model Agreement, allowing for commercial use.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.