Gemini Robotics ER 2

Our embodied reasoning model: capable of understanding the physical world and complex, multi-step planning

Our embodied reasoning model understands the physical world, plans multi-step tasks, and can coordinate multiple robots in the same space to complete a shared task.

Capabilities

Gemini Robotics ER 2 is a high-level brain for robots. It allows robots to chat with humans, understand the physical world, and plan multi-step tasks. It then hands off motor execution to any given lower level vision-language-action (VLA) model. The design of Gemini Robotics ER 2 allows the robot to “think” about what comes next while simultaneously performing its actions.

Multi-robot collaboration

Solves challenges that are too complex for one robot. Multiple robots can communicate, recognize each other's unique physical strengths, and autonomously delegate tasks to complete a shared mission.

Advanced spatial logic

Considers physical spaces, precisely identifying objects and movement. Then plans effective and safe actions in response.

Success tracking

Knows when a task is complete. And knows when it needs to try again.

Tool use

Uses tools – like Google Search or Google Calendar – to better understand its environment, and use that information to instruct its reasoning.

Conversational interactions

Understands and follows everyday commands, and can explain its thinking and actions. Adjusts to new instructions and changes in its environment efficiently.



Model information

Name
Gemini Robotics ER 2
Status
Public preview: Google AI Studio, Gemini API
Private preview: Gemini Enterprise Agent Platform
Input
  • Text
  • Image
  • Video
  • Audio
Output
  • Text
Live API
Yes (Text out only)
Tool use
  • Search
  • Function calling
  • Code execution
  • Structured output
  • URL context
Availability
  • Google AI Studio
  • Gemini API
  • Gemini Enterprise Agent Platform
Model card
View model card