Just Google DeepMind has released a up-to-date version of its AI-powered Gemini model that can control a range of different robots – including humanoids capable of performing dexterous tasks such as screwing in lightweight bulbs and tying garbage bags.
Gemini Robotics 2 combines several different artificial intelligence models into one system. Together, they allow the robot to understand its environment and learn how to behave in it. A vision language model (VLM) that understands images and video can communicate with humans and draw conclusions on how to perform various tasks. Two vision language-based action models (VLA), trained to understand movement in physical space, control the movement of the robot’s entire body, as well as the movements of its grippers or hands.
In pre-launch video demonstrations, the company showed several different robots performing elaborate tasks on their own using a connected model. In one demonstration, Apptronik’s Apollo 2 robot used Edged’s hands to organize shelves. Google DeepMind trained a model to perform these tasks using a combination of human teleoperation, video examples, and simulations – AI models are not yet capable of performing a wide range of elaborate tasks without special training.
While Anthropic and OpenAI have taken the lead in chatbots and AI coding tools, Google has more experience in robotics research and published critical Work on the apply of artificial intelligence to train robots to perform useful activities. This publication is another sign that the search giant is betting that artificial intelligence will need to break free from the digital sphere to realize its full potential. (It previously worked with Boston Dynamics, a leader in legged robots, to supply the brains for these machines.)
“This is another milestone on our path to achieving what we call physical AGI, which means we will have a robot that can do everything a human can do,” Carolina Parada, director of robotics at Google DeepMind, tells WIRED.
But giving pioneering AI models access to robots to roam workplaces or homes and manipulate objects comes with risks. Previous research has shown that using pioneering artificial intelligence to control robots can cause unexpected and sometimes unsafe behavior. The idea that these models can take sudden or unwanted actions in the digital realm became clear recently when an unreleased AI agent developed by OpenAI hacked several systems.
“The safety issue is even more pressing because they are exposed to so many other situations,” Parada says. “There will be a lot of uncertainty, so it’s important to better understand security.”
Parada says Google takes a multi-layered approach to security, using guardrails at every layer of the model. It also introduces ASIMOV-Agentic, a up-to-date benchmark for measuring the safety of various artificial intelligence systems working together to control a robot. The benchmark detects whether a command will result in a harmful or insecure result.
The company’s CEO, Demis Hassabis, previously told WIRED that he hopes to develop an AI operating system for a wide variety of robots, similar to the Android operating system for smartphones.
