Skip to content
Main Site News Console

Gemini Robotics 2 Brings Embodied Intelligence to Robots

· DeepMind Translated
DeepMind

July 30, 2026 Models

Carolina Parada

From foot to fingertip—we are teaching robots intelligent whole-body control, fine manipulation, and teamwork to complete a broad range of complex tasks.

For decades, we have dreamed of robots seamlessly integrating into our world and truly helping people. Today, that vision is taking an important step forward.

Most robots are preprogrammed or remotely operated for narrow, repetitive workflows. They lack the ability to truly learn autonomously or adapt to unpredictable environments. In addition, transferring skills learned by one robot embodiment to another remains extremely difficult. To address the most challenging problems at scale, robots of all shapes and sizes need AI models that give them the ability to think, act, and interact intelligently so they can complete tasks safely.

With Gemini Robotics, we demonstrated how Gemini’s multimodal understanding can drive action in the real world. Today, we are introducing Gemini Robotics 2—the intelligence layer for the next generation of truly adaptive robots. As it takes a meaningful first step, this major advance unlocks intelligent whole-body control, advanced dexterous manipulation, and multi-robot collaboration.

Gemini Robotics 2 can reason about every movement a robot makes, unlocking a wide range of tasks. For example, it can enable a humanoid robot to walk, squat, stretch, and manipulate objects while tidying a cluttered room. It can even collaborate with other robots to complete tasks faster. This powerful intelligence can also run on-device and adapt seamlessly to an entirely new robot embodiment in just a few hours.

We achieved this with three powerful models:

  • Gemini Robotics 2: Our most advanced vision-language-action model (VLA), which converts visual and language inputs into motion control, enabling robots to take action. The model can control a full humanoid robot, from foot to fingertip, as well as other bimanual robots. It also brings entirely new dexterous manipulation capabilities to hands and grippers.
  • Gemini Robotics ER 2: Our most capable embodied reasoning (ER) model. This vision-language model (VLM) acts as an intelligent agent for robots, enabling them to communicate with people, understand the physical world, and plan multistep tasks lasting several minutes. We are also introducing the ability for robots to work together.
  • Gemini Robotics On-Device 2: Our most efficient vision-language-action model (VLA), optimized to run locally on robotic devices. With just a few hours of data, the model can now rapidly adapt to entirely new robot embodiments.

The Gemini Robotics ER 2 reasoning model is now available in Google AI Studio and is available in private preview on the Gemini Enterprise Agent Platform. Our VLA and On-Device models are now available to early access partners. For more information on applying these models to hardware, see our developer blog.

Getting Humanoid Robots Moving: Managing Whole-Body Tasks

The world is built around the way humans move; we need to reach, bend, and maintain our balance in tight, cluttered spaces. Earlier models could control a humanoid robot’s upper body to perform tabletop tasks, while Gemini Robotics 2 extends physical AI to full-body movement.

For the first time, our model can control a complete humanoid robot, translating intent into intelligent whole-body control. For example, when controlling Apptronik’s Apollo 2 humanoid robot, we can tell it to “put the watering can in the green box on the bottom shelf.” Apollo processes the instruction, walks to the table, picks up the watering can, walks a few steps to the shelf, and places the watering can precisely in the target location. Although our robots still have room to improve their movement speed, this is an important step toward more complex real-world tasks—tasks that require robots to coordinate their entire bodies.

Bringing Advanced Dexterous Manipulation to Hands and Grippers

To be genuinely useful at home and in the workplace, robots need fine-grained dexterous manipulation capabilities. Gemini Robotics 2 unlocks entirely new physical dexterity for different types of end effectors, making robots more useful than ever whether they use mechanical hands or grippers.

The model can now control the SharpaWave five-finger, 22-degree-of-freedom hand on the Apollo 2 robot to perform delicate actions such as tying knots and closing zip-top bags. It can also operate standard two-finger parallel grippers on the Franka Duo platform to perform complex dexterous tasks, such as tightly packing a box. We will continue improving manipulation accuracy and speed to achieve human-level dexterity.

Unlocking Advanced Tasks Through Agentic Reasoning and Multi-Robot Collaboration

Most real-world tasks require multiple steps to be completed over an extended period. To handle this complexity, our embodied reasoning (ER) model, Gemini Robotics ER 2, acts as the robot’s high-level brain, processing user instructions and communicating with people. It observes the room, reasons about the steps required to complete the task, coordinates actions with the VLA, and continuously tracks progress until the task is complete. This architecture enables robots to perform complex multistep tasks, self-correct when a step fails, and generalize to new situations and goals.

With this update, we have made robots more reliable at executing long task sequences lasting several minutes and involving hundreds of decisions. Gemini Robotics ER 2 can now understand when a task begins and ends, and accurately identify when key events occur, representing a qualitative leap in progress understanding.

We are also introducing multi-robot collaboration. Different types of robots can communicate and coordinate with one another to solve complex workflows that a single robot could not complete independently.

Enabling On-Device Models to Rapidly Adapt to Any Robot

Many robotic applications need to operate without network latency or an internet connection. Gemini Robotics On-Device 2 was specifically built for these constraints—it is our most efficient vision-language-action model (VLA), optimized to run locally on robotic devices.

The model natively supports multiple robot embodiments and inherits the advanced “action transfer” technology we introduced in Gemini Robotics 1.5. Now, with just a few hours of adaptation—typically using fewer than 200 examples—we can adapt the model to a new bimanual robot embodiment. This approach remains effective even when the new robot differs significantly in shape, sensors, and degrees of freedom. The figure below shows Dexmate, SO101, and Trossen platforms performing a variety of tasks.

Advancing Our Commitment to Safe, Responsible Robotics

Safety is foundational to our robotics research. As robots gain greater physical capabilities, we are committed to ensuring end-to-end safety and alignment. With every new release, we use a layered approach that combines traditional physical safety measures with a robust AI safety framework.

Gemini Robotics 2 is specifically designed to improve robots’ safety when handling uncertainty in the real world and collaborating with people.

We introduced ASIMOV-Agentic, a new benchmark for agentic safety orchestration and uncertainty resolution. For example, it can measure three-levelેશન