Gemini Robotics ER 2: Advanced AI for Collaborative Robots

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Introduction of Gemini Robotics ER 2
Gemini Robotics ER 2 represents a major advancement in powering robots through video understanding, task orchestration, and multi-robot collaboration, making robots more useful in the physical world.
We are launching Gemini Robotics ER 2, a new model designed to act as a high-level brain for robots. It enables real-time spatial reasoning, multi-step task planning, and collaboration among different robots. You can access the model now via the Gemini API, Google AI Studio, or the Gemini Enterprise Agent Platform to start creating your own physical AI agents.
Advancements in Physical Agent Capabilities
Most tasks in the physical world are complex and require multiple steps to complete. Gemini Robotics ER 2 is a physical agent that orchestrates the steps for the robot, allowing it to self-correct and generalize to newer situations. To build an agentic setup, developers can declare low-level control interfaces — such as Vision-Language-Action (VLA) models or navigation APIs — as tools, and stream multimodal videos, audio, or text directly into the model.
Gemini Robotics ER 2 enhances this tool orchestration workflow. We can evaluate its performance with robots in simulation, using real robot control, and even pair it with a human remotely controlling the robot.
Temporal Intelligence for Robust Task Execution
One of the most challenging issues in robotics is knowing when a task is complete. Gemini Robotics ER 2 brings an advancement in video understanding and progress tracking to verify that complex tasks — such as screwing in a light bulb or tying a trash bag — are completed according to specifications before moving on to the next task.
Continuous Progress Classification
Progress classification refers to a robot's ability to track advancement toward the completion of a task. In our evaluations, we assign each frame of a video stream five levels of progress (0-20%, 20-40%, 40-60%, 60-80%, 80-100%). By quantifying task progress, Gemini Robotics ER 2 provides robots with real-time situational awareness, allowing them to adjust their actions on the fly or retry failed steps without restarting an entire workflow.
Gemini Robotics ER 2 achieves a 57.4% accuracy on progress classification tasks, surpassing previous generation models and competing state-of-the-art models.
Precision in Moment Detection
Moment detection measures a model's ability to identify the exact video frame where a critical event occurs (for example, when to stop pouring coffee into a cup). Gemini Robotics ER 2 achieves significant performance gains in moment detection, allowing robots to transition precisely from one task to another, verify success, and suggest corrections.
For moment detection tasks, Gemini Robotics ER 2 achieves a 91.3% accuracy and an average absolute distance of 0.96s. It closely competes with much larger model categories but offers this precision at a fraction of the computational cost and with four times the execution speed — with latency under one second being crucial for safely operating physical robots in the real world.
Multi-Robot Collaboration
No single robot is suitable for every task — a wheeled rover excels indoors, while a humanoid robot may perform better on uneven terrain. Gemini Robotics ER 2 enables collaboration among multiple robots, allowing diverse machines to communicate through shared semantic understanding to transfer and complete complex tasks.
Enhancing General Spatial Intelligence
Gemini Robotics ER 2 advances our core spatial reasoning capabilities, measured by three indicators:
-
Success/Failure Detection: Now operates on raw video streams rather than static snapshots to detect failures during execution, such as spills, slips, or misalignments.
-
General Instrument Reading: Extends beyond circular dials and sight glasses to include digital displays, linear scales, rulers, and liquid thermometers. We tested it on 10 different types of instruments.
-
Improved Spatial VQA: Enhances Visual Question Answering through Gemini's advancements in multimodal understanding.
Gemini Robotics ER 2 consistently achieves the highest accuracy across all key capabilities, with strengths including success detection (image/video), Question Answering (ERQA), and generalized instrument reading.
Advancing Safety for Embodied Intelligence
Gemini Robotics ER 2 is our safest model, achieving significant gains on Safety Instruction Following and Human Proximity benchmarks, which assess how a model adheres to physical constraints during reasoning tasks and spatial awareness to detect humans. We found that Gemini Robotics ER 2 successfully stops a humanoid robot when a person is nearby and resumes its work autonomously only when the area is clear.
In the future, our plans are to push these models toward even more complex tasks to accelerate the development of useful robots and support the robotics community.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.