Sponsored by Byond Boundrys Consulting - Empowering Ideas, Delivering Results
Google New Advanced

Gemini Robotics ER 2

Gemini Robotics ER 2 is Google's preview robotics reasoning model with 131K context, multimodal inputs, tool orchestration and $2/$10 pricing.

Robotics Vision-Language ModelTextImageAudioVideo Paid
In plain English

What is this model and why does it matter?

Gemini Robotics ER 2 is Google's preview embodied-reasoning model for spatial understanding, robot video analysis, multi-step tool orchestration and multi-robot coordination.

Robotics researchRobot orchestrationVideo progress understandingSpatial reasoningInstrument readingMulti-robot coordination
Model overview

Gemini Robotics ER 2: features, use cases and important details

Gemini Robotics ER 2 is Google’s embodied-reasoning model for developers building robots and physical-world agents rather than conventional text assistants.

What this model is

The model is a vision-language system specialized for understanding physical environments. Google highlights spatial reasoning, robot-video understanding, progress classification, success detection, instrument reading, multi-step tool orchestration and coordination across multiple robots. The key concept is that ER 2 does not directly replace a robot controller. It acts as a reasoning layer that interprets multimodal observations, plans steps and calls developer-defined functions or lower-level robot policies.

Technical capabilities and model behavior

The standard preview endpoint accepts text, images, video and audio with a 131,072-token input limit and 65,536-token output limit. It supports thinking, function calling, structured outputs, code execution, computer use, file search, Search/Maps grounding and URL context. This broad tool set lets a robot agent combine observations with external information and orchestrate multiple systems, although every physical action still needs engineering controls outside the language model.

How it works in real applications

Consider an industrial inspection workflow. ER 2 can watch a robot video, identify the stage of an assembly, read an instrument and decide which inspection function should run next. In a warehouse research system, it can reason over camera feeds and delegate tasks to separate robot functions. For multi-robot coordination, it can operate at the planning layer while lower-level motion planners handle collision avoidance, dynamics and actuator control.

Current status and availability

Google released Gemini Robotics ER 2 in public preview on July 30, 2026. The family includes the standard preview endpoint and a separate streaming preview optimized for low-latency robot agents. The standard endpoint supports more of Google’s general tool ecosystem and Batch API, while the streaming endpoint prioritizes live bidirectional interaction. This directory row is specifically for gemini-robotics-er-2-preview.

Pricing and deployment considerations

Standard paid pricing is $2/M multimodal input tokens and $10/M output tokens. Batch processing costs $1/M input and $5/M output. Context caching has separate pricing. Search grounding may also add query charges. In robotics, token cost is only one part of total cost: camera/video processing, robot hardware, simulation, testing and physical safety systems will usually dominate the real deployment budget.

Who should choose this model?

Choose ER 2 when the application needs semantic spatial reasoning and high-level planning across visual observations and tools. It is not a substitute for a deterministic PLC, motion planner or certified safety controller. Traditional computer-vision models may be cheaper when the task is only object detection, while a streaming robotics endpoint may be better when continuous low-latency feedback is required.

Important limitations and trade-offs

The model is still in preview, so APIs and performance can evolve. More importantly, physical-world errors can damage equipment or injure people. Developers should use constrained action spaces, independent safety layers, simulation, confirmation gates and telemetry. A model’s confident interpretation of a scene is not evidence that an action is mechanically safe.

Gemini Robotics ER 2 capabilities and use cases

In addition, its main capabilities include Spatial reasoning, Video understanding, Progress/success detection, Function calling, Agentic code execution and Multi-step tool use. For example, common use cases include Robotics agents, Industrial automation research, Physical-world reasoning, Robot monitoring and Multi-robot systems.

Who should consider Gemini Robotics ER 2?

In practice, this model may suit Robotics research, Robot orchestration, Video progress understanding, Spatial reasoning, Instrument reading and Multi-robot coordination. Also, notable strengths include Dedicated embodied reasoning model, Multimodal sensor context, Tool orchestration and Structured outputs. However, review trade-offs such as Model reasoning must not be treated as a safety controller, Preview behavior and limits can change and Robotics deployments require independent collision/safety systems before adopting it.

Gemini Robotics ER 2 pricing and access

Meanwhile, $2.00 per 1M text/image/video/audio input tokens and $10.00 per 1M output tokens. Batch pricing is $1/M input and $5/M output. Standard preview pricing is $2/M multimodal input and $10/M output; Batch halves those token rates for non-interactive jobs.

Official resources and verification

Use the official model website, official documentation, pricing or release source and additional primary source to confirm current availability, limits and pricing. Product details can change after publication, so rely on primary documentation for final decisions.

Compare with other AI models

Next, continue your research in the AI models directory, Google models and Robotics Vision-Language Model models. Compare providers, pricing, modalities and practical limitations side by side to choose the right model for your workflow.

Get started

How to use this model

  1. Create a Gemini API key.
  2. Use gemini-robotics-er-2-preview.
  3. Provide text plus robot images, video or audio context.
  4. Define functions that represent robot/tool actions with safe execution controls.
  5. Use the model to reason, plan and call tools while independently validating physical actions.
Copy and try

Example prompts

  • Analyze this robot-camera video and identify whether the assembly task has completed successfully.
  • Plan the next safe manipulation steps and call the available robot functions one at a time.
  • Read this instrument panel and explain the physical state before any action is taken.
Capabilities

What it can do

  • Spatial reasoning
  • Video understanding
  • Progress/success detection
  • Function calling
  • Agentic code execution
  • Multi-step tool use
  • Multi-robot coordination
  • Computer use
Best for

Practical use cases

  • Robotics agents
  • Industrial automation research
  • Physical-world reasoning
  • Robot monitoring
  • Multi-robot systems
Pricing

What does it cost?

$2.00 per 1M text/image/video/audio input tokens and $10.00 per 1M output tokens. Batch pricing is $1/M input and $5/M output.

Input$2.00 / 1M multimodal input tokens
Output$10.00 / 1M output tokens
Simple summaryStandard preview pricing is $2/M multimodal input and $10/M output; Batch halves those token rates for non-interactive jobs.

What stands out

  • Dedicated embodied reasoning model
  • Multimodal sensor context
  • Tool orchestration
  • Structured outputs
  • Search/Maps grounding

Things to consider

  • Preview lifecycle
  • Physical actions create higher safety requirements
  • No open weights
Limitations

Important restrictions and trade-offs

  • Model reasoning must not be treated as a safety controller
  • Preview behavior and limits can change
  • Robotics deployments require independent collision/safety systems
SimplifyAITools verdict

Our editorial take

A strategically important model for the directory because it represents the shift from chat AI toward embodied agents that reason about and coordinate physical systems.

References

Primary sources

  1. Open source 1 ↗
  2. Open source 2 ↗
  3. Open source 3 ↗