July 30, 2026
Gemini Robotics ER 2 represents a major leap forward in robotics capabilities. By empowering robots with video understanding, task orchestration, and multi-robot collaboration, it enables them to better help people handle a wide range of tasks in the real world.
Steven Hansen
Senior Software Engineer
Peng Xu
Software Engineer

Accurate spatial reasoning alone is not enough for robots to assist people in everyday environments. Robots must also think quickly, keeping the pace of their decisions and reasoning in step with the real-time speed of the physical world.
That’s why today we’re introducing Gemini Robotics ER 2, our most capable embodied reasoning model for robotics to date. You can think of Gemini Robotics ER 2 as a robot’s high-level brain. It enables robots to converse with people, understand the physical world, and plan multistep tasks, while delegating motion execution to any specified low-level Vision-Language-Action (VLA) model. Gemini Robotics ER 2 can also natively call tools such as Google Search, as well as other user-defined functions. It is designed to allow robots to think about what to do next while carrying out actions.
Compared with Gemini Robotics ER 1.6, Gemini Robotics ER 2 delivers significant upgrades. By observing a continuous video stream, robots can now track their own progress, adjust when something goes wrong, and accurately determine when to move on to the next step. We’ve also introduced multi-robot collaboration, enabling robots to work together in shared spaces to complete complex workflows that a single robot could not accomplish on its own.
Gemini Robotics ER 2 is now publicly available to developers through the Gemini API and Google AI Studio, and is available in private preview on the Gemini Enterprise Agent Platform. To help developers get started quickly, we’re sharing relevant examples that demonstrate how to configure the model and use prompts to support more practical physical AI tasks.
Advancing Physical Agent Capabilities
Most tasks in the real world are complex and require multiple steps to complete. Gemini Robotics ER 2 is a physical agent that can orchestrate each step for a robot, help it correct its own mistakes, and generalize to a wider range of novel situations. To build agentic systems, developers can declare low-level control interfaces—such as Vision-Language-Action (VLA) models or navigation APIs—as tools, and stream multimodal video, audio, or text directly to the model.
Gemini Robotics ER 2 improves tool-orchestration workflows. We can evaluate its performance in simulated environments using real-world robot control, and even combine it with human operators remotely controlling robots.
Gemini Robotics ER 2 consistently outperforms ER 1.6 in tool orchestration across three control modes: real VLA, simulated VLA, and human teleoperation.

In robotics, high-level reasoning depends on execution speed. Gemini Robotics ER 2 integrates the Gemini Live API, operating through a bidirectional streaming endpoint optimized for latency-sensitive tasks. This enables smooth orchestration: Gemini Robotics ER 2 can direct action models and robot APIs to complete multistep tasks without awkward pauses caused by “stopping to think.”
To demonstrate this, we worked with our partner Boston Dynamics and its Spot robot to build a demonstration. We used Gemini Robotics ER 2 to orchestrate the Spot API, including capabilities such as navigation and arm movement, creating a robot that can fetch items for you based on interactive commands.
Gemini Robotics ER 2 powers Boston Dynamics Spot to fetch a popcorn snack in response to a natural-language instruction.
The related code and other examples are available on GitHub.
Unlocking Time Intelligence for Robust Task Completion
One of the biggest challenges in robotics is determining when a task is complete. Gemini Robotics ER 2 delivers a major leap in video understanding and progress tracking. It can verify whether complex tasks—such as tightening a light bulb or tying a garbage bag—have been completed as required before moving on to the next task.
In this update, we further advanced two foundational capabilities for understanding task progress: progress classification and precise moment localization.
Continuous Progress Classification
Progress classification refers to a robot’s ability to track how much of a task has been completed. In our evaluations, we divide each frame in a video stream into one of five progress levels: 0%–20%, 20%–40%, 40%–60%, 60%–80%, and 80%–100%. By quantifying task progress, Gemini Robotics ER 2 provides robots with real-time situational awareness, enabling them to dynamically adjust their actions or retry a step when it fails without restarting the entire workflow.
Gemini Robotics ER 2 achieved 57.4% accuracy on progress classification tasks, outperforming its predecessor and other frontier models from competing labs.

Precisely Localizing Key Moments
Key-moment localization measures a model’s ability to identify the exact video frame in which a critical event occurs, such as determining when to stop pouring coffee into a cup. Gemini Robotics ER 2 delivers significant performance improvements in key-moment localization, allowing robots to switch precisely between tasks, verify whether a task has succeeded, and propose corrective actions.
On key-moment localization tasks, Gemini Robotics ER 2 achieved 91.3% accuracy with a mean absolute distance of 0.96 seconds. Its performance approaches that of much larger model classes, while requiring only a fraction of their computational cost and running four times faster. Subsecond latency is exactly what robots need to operate safely in the real world.

Multi-Robot Collaboration
No single robot can handle every task—wheeled robots excel in indoor environments, while humanoid robots may be better suited to navigating rough terrain. Gemini Robotics 2 supports multi-robot collaboration, enabling different types of machines to communicate, hand off tasks, and complete complex operations through a shared semantic understanding. Learn how Gemini Robotics ER 2 empowers Apptronik’s [Apollo 2](https://apptron…