Event date · · Google DeepMind Gemini Robotics

Gemini Robotics: Bringing AI into the Physical World: Google's Gemini 2.0-Powered Robot Foundation Model

FACT STATEMENT

In March 2025, Google DeepMind released the Gemini Robotics series, including two models: Gemini Robotics (a general VLA model) and Gemini Robotics-ER (an embodied reasoning model). The former can directly control robots, performing complex manipulation tasks with robustness to object types, positions, environmental changes, and open-vocabulary instructions. The latter enhances spatial and temporal understanding, supporting object detection, pointing, trajectory prediction, grasp prediction, multi-view correspondence, and 3D bounding box prediction. Through fine-tuning, Gemini Robotics can learn new tasks (requiring only 100 demonstrations) and adapt to new robot morphologies.

What happened

Gemini Robotics is the first robot foundation model based on Gemini 2.0, extending the general capabilities of large language models to the physical world. Its VLA model can directly control robots, while the embodied reasoning model provides spatial understanding. This work demonstrates the ability to learn new tasks from few demonstrations and adapt to new robot morphologies, marking a significant advance in general-purpose robotics.

Technical significance

Gemini Robotics is based on the Gemini 2.0 multimodal model, adopting a vision-language-action (VLA) architecture. The model takes images and language instructions as input and directly outputs robot actions. Key capabilities include: 1) robustness to object types, positions, and environmental changes; 2) following open-vocabulary instructions; 3) learning long-horizon, dexterous tasks through fine-tuning; 4) learning new short-horizon tasks from 100 demonstrations; 5) adapting to new robot morphologies. Gemini Robotics-ER focuses on embodied reasoning, outputting spatial information (e.g., 3D bounding boxes, grasp points). The paper does not provide detailed benchmark comparisons but shows qualitative results on various manipulation tasks. Safety aspects discuss collision avoidance and human supervision considerations.

Industry impact

This model has transformative potential for industrial automation, warehouse logistics, home service robots, and other fields. A general-purpose robot foundation model can reduce robot programming costs and enable robots to adapt to unstructured environments. Google's ecosystem advantages (e.g., Android, Google Cloud) may accelerate deployment.

Decision value

Manufacturing and logistics companies are advised to evaluate the feasibility of Gemini Robotics in tasks such as picking and assembly. Investment opportunities lie in the robot foundation model track, especially companies collaborating with Google. On the engineering side, integrating Gemini Robotics-ER into existing robot systems can enhance perception and planning capabilities.

What to watch

Areas to watch: 1) long-term stability and safety of the model in the real world; 2) comparison with competitors like NVIDIA GR00T; 3) fine-tuning data requirements and computational costs; 4) impact of open-sourcing on the community; 5) ethical and employment implications.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.