A factory robot arm, a humanoid robot and a two-arm research robot can have completely different bodies, sensors and movements. Yet in 2026, AI researchers are getting closer to using a shared model across these different machines. The goal is not to make every robot move in exactly the same way, but to give different robots a common layer of intelligence that can be adapted to their physical abilities.
From Robot-Specific Programs to Foundation Models
Traditional robots are usually programmed for specific tasks and hardware. A robotic arm may be carefully configured to pick objects from one position, while a humanoid robot requires a completely different control system.
Robotics foundation models take a broader approach. They are trained using large amounts of robot demonstrations, images, language instructions and physical-action data. Instead of learning only one fixed task, the model is designed to understand instructions, recognize objects and generate actions that can be adapted to different robots.
What 2026 Models Are Showing
In July 2026, Google DeepMind introduced Gemini Robotics 2, a vision-language-action model designed to convert visual and language information into physical robot actions. Google demonstrated the same model checkpoint across different embodiments, including Apptronik's Apollo 2 with different hands and a Franka Duo robot with a Robotiq gripper.
Google also introduced Gemini Robotics On-Device 2 in July 2026. The model is designed to run locally on robots and can adapt to new robot embodiments using a few hours of adaptation and typically fewer than 200 examples, according to Google's reported results.
NVIDIA also released Isaac GR00T 1.7 in July 2026, an open vision-language-action model aimed at generalized humanoid robot skills. NVIDIA reported that the model was trained using both real and simulated robot data.
Why Skills Do Not Transfer Automatically
A skill learned by one robot cannot simply be copied as identical motor commands to another. Robots have different numbers of joints, arm lengths, grippers, cameras and movement limits.
This is known as the cross-embodiment problem. A model may understand the instruction “pick up the cup,” but the exact movement required depends on the robot performing the task.
Research published in 2026 is addressing this problem directly. An IEEE Robotics and Automation Letters paper introduced a Cross-Embodiment Interface that transfers demonstrations between different robot arms and end-effectors. In its experiments, the researchers tested transfers across 16 simulated embodiments and between different real-world robot configurations.
One Intelligence Layer, Different Robot Controllers
The emerging approach is therefore not necessarily one AI model directly controlling every motor on every robot.
A foundation model can handle higher-level understanding: what the robot should do, what objects are present and what sequence of actions is needed. A robot-specific control layer can then translate that intention into movements suitable for its joints, sensors and end-effector.
Google's Gemini Robotics ER 2, also launched in July 2026, illustrates this separation. It acts as a high-level reasoning system for understanding the physical environment and planning multi-step tasks, while lower-level models handle physical action. Google also demonstrated collaboration between different robot types.
The Goal Is General Intelligence, Not Identical Robots
The 2026 developments show that robotics is moving toward shared AI models that can be adapted across different machines. However, this does not mean a single model can immediately operate every robot without additional training or engineering.
Differences in sensors, hardware, movement, safety requirements and physical environments still matter. For now, the more realistic goal is shared robotic intelligence with robot-specific adaptation. If this approach continues to improve, developers may not need to build a completely new AI system whenever they introduce a different robot.