Skip to content
Research · Jul 30, 2026

Google DeepMind launches Gemini Robotics ER 2 for real-time robot orchestration and multi-robot collaboration

The new model enables robots to reason from video, plan multi-step tasks, self-correct, and coordinate with other robots via the Gemini API, Google AI Studio, and Enterprise Agent Platform.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • Gemini Robotics ER 2 is a new model designed to act as a high-level brain for robots, enabling real-time spatial reasoning, multi-step task planning, and collaboration between robots.

Google DeepMind introduced Gemini Robotics ER 2, a model intended to serve as a high-level controller for robots. It supports real-time spatial reasoning, multi-step task planning, and collaboration among multiple robots.

The model accepts multimodal inputs—including streaming video, audio, and text—via the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform, enabling developers to build physical AI agents.

Gemini Robotics ER 2 can orchestrate low-level control interfaces such as Vision-Language-Action (VLA) models or navigation APIs, and it can call external tools like Google Search or user-defined functions to assist with task execution.

Compared to its predecessor, ER 1.6, ER 2 improves real-time progress tracking by analyzing continuous video feeds, allowing robots to detect errors, self-correct, and determine when to advance to the next step in a task.

The model introduces multi-robot collaboration, enabling coordinated workflows across multiple robots that would be infeasible for a single robot to perform alone.

In evaluations, ER 2 outperformed ER 1.6 in tool orchestration across three control modes: real-world VLA, simulated VLA, and human teleoperation.

ER 2 integrates with the Gemini Live API using a bidirectional streaming endpoint optimized for low-latency tasks, reducing the need for stop-and-think pauses during execution.

A demonstration with Boston Dynamics’ Spot robot shows ER 2 orchestrating navigation and manipulation APIs to fetch objects on natural language command, with code examples available on GitHub.

Sources
  1. 01Google DeepMind — BlogGemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Also on Research

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.