Skip to main content

AI Robots Are Learning to See and Think: What Happens When AI Gets a Body

A comprehensive overview of AI Robots Are Learning to See and Think: What Happens When AI Gets a Body detailing architecture, practical implications, and k...

By Mittapalli Sriram
Published: Sep 30, 2026
4 mins read
👁️ 19 Unique Views
AI Robots Are Learning to See and Think: What Happens When AI Gets a Body
The scale of inference: Optimized for multimodal workloads.
Premium Insight

Why It Matters

It explains how AI is moving beyond chatbots and into physical robots that can see, understand instructions, plan tasks, and act in the real world. It also shows that the technology is still developing, especially for complex hand movements and reliable everyday use.

A humanoid robot hears one sentence: put the watering can in the green bin on the bottom shelf. It walks to a table, picks up the can, takes a few steps to the shelves and puts it where it was told. Nobody is holding a remote.

That was one of the demos Google DeepMind showed on 30 July 2026 with Gemini Robotics 2, its latest set of AI models for robots. It looks like a small task. For a robot, working out which bin is the green one and how to reach the bottom shelf without tipping over is the hard part.

From scripts to understanding

Older factory robots follow a script. They repeat one movement in one spot all day, and if a part shifts by a few centimetres, they fail. The new approach gives robots the kind of AI behind chatbots, so they can understand words and images and decide what to do.

DeepMind splits the job across three models. The first is a vision-language-action model, or VLA. It turns what the robot sees and what it is told into motor commands, which move the limbs. The second is a reasoning model called ER 2, short for embodied reasoning. It plans a job that can run for several minutes, checks whether each step worked, and tries again if it did not. The third is a smaller model that runs on the robot itself, with no internet needed. DeepMind says it can be adapted to a new robot body with a few hours of data, usually fewer than 200 examples.

What the numbers say

DeepMind published success rates, and they come from its own tests. On the Apollo 2 humanoid, picking something up from a table worked about 68% of the time, from a shelf about 76%, and from the floor about 46%. Finger work was harder. Screwing in a light bulb worked 36% of the time, tying a trash bag 44%, and sealing a ziplock bag 40%. Unscrewing a bulb reached 92%. DeepMind itself says multi-finger tasks are still challenging and that its robots need to move faster.

So the picture is mixed. A robot that fails one time in three is good research, not a helper for your kitchen.

Demo versus daily work

Robot videos are rehearsed and edited, and sometimes a person is steering the machine out of sight. A report on the Robotics Summit in Boston this May, citing AFP-JIJI, said one humanoid was being controlled by a person standing to the side. A floor report from Automate 2026 reached a similar view: the bodies are ready, but the AI needs far more training data. Both are secondhand and neither is about Gemini Robotics 2, but they are a fair reminder to ask whether any robot clip was running on its own.

You also cannot buy this yet. ER 2 is open to developers on Google AI Studio and in private preview for businesses. The VLA and on-device models are limited to early-access partners.

When a mistake has a body

A chatbot that gets something wrong gives a wrong answer. A robot that gets something wrong can knock things over or hurt someone. DeepMind says ER 2 can notice when a person gets too close and bring the robot to a safe stop, and calls it its safest robotics model on its own benchmarks. It also released a test called ASIMOV-Agentic, which checks whether the reasoning model will refuse an unsafe instruction from the movement model and ask a human for help when unsure. These are the company's claims on its own tests, so outside checks will matter more.

What to watch

The bigger shift is not one robot doing one neat trick. DeepMind says a single model checkpoint controlled three different setups: a humanoid with two types of hands, and a two-armed robot with simple grippers. If that holds up, the same AI could be reused across many robot designs. Watch for higher success rates, outside tests, and robots that work a full shift unaided.

Found this analysis insightful?

Share with colleagues, engineers, and your network.

Tags & Topics

Discussion

Leave a Comment

No comments yet. Be the first to start the conversation!

Link copied to clipboard!