For years, AI companies built separate tools for separate senses. One model could read text. Another could look at images. A third could listen to audio. That wall is coming down now. OpenAI, Google DeepMind, Anthropic, and robot makers like Figure and Apptronik are all trying to build one system that can watch a room, hear what people say in it, remember what happened last time, and then actually do something.
This matters because each skill makes the next one more useful. An AI that can only look at a photo is a neat trick and nothing more. But an AI that can see your kitchen, hear you ask for coffee, remember that you take it black, and then move a robot arm to make it, is a completely different kind of product. That is what this whole industry is building toward. But right now, each piece is at a very different stage of being ready.
Seeing and Hearing, Put Together
Google DeepMind builds two models that work as a pair. One is Gemini Robotics, which turns what a robot sees and hears into actual movement. The other, called Gemini Robotics-ER, handles the reasoning: understanding space, tracking how far along a task is, and planning several steps ahead. In one demo, a robot called Apollo 2, made by Apptronik, is told in plain words to put a watering can into a bin on the bottom shelf. It walks across the room, finds the can, and puts it away, changing its steps as it goes instead of just following a script.
The two models are at different stages, and it's worth keeping them separate. The reasoning model, Gemini Robotics-ER 2, opened up to any developer through Google's own API in July 2026, so it's now genuinely something outside programmers can test. But the model that actually controls a robot's body is a different story. Google has only given that one to a small group of partners, like Boston Dynamics, Agility Robotics, and Apptronik. You still can't buy a robot running it.
The bigger claim worth questioning is about generalization — basically, whether a skill learned on one robot can transfer to a totally different one. Google says its newest Gemini Robotics model can do exactly that, cutting the setup time for a new robot down to just a few hours instead of months. But that number comes from Google itself. No outside group has independently confirmed it. So treat it as a claim about an early research system, not a proven fact about something ready for the market.
Remembering Is the New Part
Out of everything in this article, memory is the one piece that has actually shipped to regular people, not just to lab partners. Until recently, every AI chat started from scratch, no matter how many times you'd talked to it before. That changed for Claude in March 2026, when Anthropic turned on memory for every user, including people on the free plan, so it can remember things across separate chats. ChatGPT has had some version of memory since 2024 and keeps improving it, quietly building up a picture of your job, your habits, and the projects you keep coming back to.
This doesn't mean the AI saves every word you've ever typed. Both companies boil down what they remember into short summaries, and you can look at them, edit them, or delete them. That's the trade-off: it remembers the big picture, but loses the small details. It also doesn't follow you between apps. Whatever Claude remembers about you stays in Claude, and whatever ChatGPT remembers stays in ChatGPT, unless you manually move it yourself.
The Gap Between the Demo and Real Life
Getting a robot to walk across a clean, well-lit demo room and drop something in a bin is genuinely hard to build. But it is very different from trusting that same robot alone in a messy warehouse, a busy hospital hallway, or your actual kitchen. Even Google seems to know this. It pairs every new robotics release with extra safety measures, which is itself a hint that the company doesn't think the physical, body-controlling side of these systems is ready to run on its own just yet, even while it's comfortable opening up the reasoning side to developers.
Memory has its own weak spots too. Because these systems compress what they remember instead of saving everything, they can forget one specific thing you told them weeks ago, even while still remembering the general pattern of what you like. And since vision, hearing, memory, and action are still mostly separate systems stitched together, rather than one single brain, a mistake in one part — like misreading a room — can throw off everything that happens after.
What This Could Lead To
In the near term, expect memory to keep getting better inside the chat apps we already use, since it's already a real product and cheap to keep improving. Robotics has a longer road ahead, but it's moving faster than it looks: Google opening its reasoning model to developers is a concrete, near-term sign that at least part of this stack is leaving the lab. The physical, body-controlling side is likely to widen next, but a robot that can handle messy, unpredictable, real-world settings fully on its own is still more guess than plan. And the biggest claim of all — one single AI that sees, hears, remembers, and acts, without you telling it which sense or tool to use — is a long-term goal. No lab has actually built that yet.
If it does eventually happen, it wouldn't just be a cool tech demo. It would start to change real jobs, like elder care, warehouse work, and equipment inspection, where the whole value has always come from doing exactly these things together: watching, listening, remembering, and acting. Getting there won't come down to one big breakthrough moment. It will come down to the slow, unglamorous work of closing the gap between what looks impressive in a demo video and what can actually be trusted to run on its own, every day, with nobody standing there to catch its mistakes.