July 30, 2026
Models
Carolina Parada
From toes to fingertips — we are training robots to achieve intelligent whole-body control, fine dexterous manipulation, and collaborative cooperation to complete a wide range of complex tasks
For decades, we have dreamed of robots seamlessly entering our world and lending a helping hand. Today, that vision takes a major step forward.
Most robots are pre-programmed or teleoperated, capable only of executing narrow, repetitive sequences of tasks. They lack the ability to truly learn autonomously or adapt to unpredictable environments. Furthermore, transferring learned skills from one robot embodiment to another remains extremely difficult. To tackle the toughest problems at scale, robots of all shapes and sizes need to be empowered by AI models that allow them to think, act, and interact intelligently to complete tasks safely.
We previously demonstrated how Gemini’s multimodal understanding drives real-world action through Gemini Robotics. Today, we introduce Gemini Robotics 2 — the intelligence layer powering the next generation of truly adaptable robots. As it takes its first literal “steps,” this major advancement unlocks intelligent whole-body control, advanced dexterous manipulation, and multi-robot collaboration.
Gemini Robotics 2 enables robots to reason through every action, unlocking a wide range of tasks. For example, it can allow a humanoid robot to walk, squat, stretch, and manipulate objects to clean up a messy room. It can even collaborate with other robots to get work done, speeding up task progress. What’s more, this deep intelligence can run locally on-device while seamlessly adapting to entirely new robot embodiments in just a few hours.
We achieve all of this through three powerful models:
- Gemini Robotics 2: Our most advanced vision-language-action (VLA) model, which translates visual and language inputs into motor control, enabling robots to take action. This model can control an entire humanoid robot, from toes to fingertips, as well as other dual-arm robots. It also brings a new level of dexterous manipulation to hands and grippers.
- Gemini Robotics ER 2: Our most powerful embodied reasoning (ER) model. It is a vision-language model (VLM) that acts as our agent, enabling robots to communicate with humans, understand the physical world, and plan multi-step tasks spanning several minutes. We also introduce the ability for robots to work together as a team.
- Gemini Robotics On-Device 2: Our most efficient vision-language-action (VLA) model, optimized to run locally on robot hardware. The model can now quickly adapt to entirely new robot embodiments with just a few hours of data.
Our reasoning model, Gemini Robotics ER 2, is now available on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform. Our VLA and On-Device models are open to early access partners. To learn how to deploy these models to your hardware, check out our developer blog.
Humanoids in Motion: Managing Whole-Body Tasks
The world is designed for human movement; it requires us to reach, bend, and balance in tight, cluttered spaces. While our previous models controlled the upper body of humanoid robots for tabletop tasks, Gemini Robotics 2 extends physical AI to whole-body motion.
For the first time, our model can control an entire humanoid robot, translating intent into intelligent whole-body control. For instance, when controlling Apptronik’s Apollo 2 humanoid robot, we can ask it to “Put the watering can in the green bin on the bottom shelf.” Apollo processes the instruction, walks to the table, picks up the watering can, walks a few steps to the shelf, and places it precisely in the target location. While there is still room to improve the speed of our robots’ movements, this is a major step toward the skills needed to complete more complex, real-world tasks that require whole-body coordination.
Bringing Advanced Dexterous Manipulation to Hands and Grippers
To be truly useful in homes and workplaces, robots need fine manipulation capabilities. Gemini Robotics 2 unlocks new levels of physical dexterity across different end-effectors, making robots more practical than ever, whether they use hands or grippers.
The model can now control the five-fingered, 22-degree-of-freedom SharpaWave hand on the Apollo 2 robot to perform fine motor actions like tying knots and sealing zip-top bags. It can also operate standard two-finger parallel grippers on the Franka Duo platform to execute complex, dexterous tasks like tight packing. We will continue to improve precision and speed to achieve near-human levels of dexterity.
Unlocking Advanced Tasks Through Agentic Reasoning and Multi-Robot Collaboration
Most real-world tasks require multiple steps over an extended period. To address this complexity, our embodied reasoning (ER) model, Gemini Robotics ER 2, acts as the robot’s high-level “brain,” processing user instructions and communicating with humans. It observes the room, reasons through the steps needed to complete a task, coordinates with the VLA to execute actions, and tracks progress until the task is complete. This architecture enables robots to perform complex, multi-step tasks, self-correct if a step fails, and generalize to new situations and goals.
In this update, we enable robots to more reliably execute longer task sequences that last several minutes and involve hundreds of decisions. Gemini Robotics ER 2 can now understand when a task starts and ends, and precisely pinpoint the moments when key events occur, marking a leap forward in progress understanding.
Additionally, we introduce multi-robot collaboration. This allows different types of robots to communicate and work together to tackle complex workflows that