Skip to main content

What Does an AI Model Actually See When a Robot Looks at a Room?

A robot's view of a room is a stack of estimates, from object outlines to distances, and some of them fail in surprising places.

By Koushik Parupally
Published: Oct 05, 2026
4 mins read
👁️ 24 Unique Views
What Does an AI Model Actually See When a Robot Looks at a Room?
The scale of inference: Optimized for multimodal workloads.
Premium Insight

Why It Matters

Indian firms are developing robot-vision technology to improve object detection and depth sensing. In February 2026, e-con Systems launched its DepthVista Helix camera, while Bengaluru-based CynLr unveiled an object-perception platform. CynLr says its technology can detect transparent and reflective objects, though independent testing is limited. Meanwhile, Cognex agreed to acquire depth-camera maker RealSense for about ₹4,400 crore in September 2026.

In an August 2026 preprint, researchers at the University of Hong Kong and DiDi's Voyager Research tested a Unitree Go2 Air quadruped carrying a widely used RealSense D435 depth camera. In three lab scenes with transparent obstacles, 20 trials each, it reached its goal without a collision or human help in only 10 percent of trials on average (5, 15 and 10 percent by scene). The camera often failed to register the glass at all. A robot doesn't see a room as you do. It sees numbers, and some of them are wrong.

From pixels to objects

A camera gives a robot a grid of colour values, nothing more. Object detection software compares patterns in that grid with patterns learned from many labelled images, then draws a labelled box around each match. Segmentation goes further and assigns every pixel to an object. That matters because a gripper needs an outline, not a rectangle.

Meta's SAM 3, released in November 2025, is background rather than news. According to Meta, it uses text or example-image prompts to identify, segment and track matching objects in images and video. It is a research model, not a robot product.

Adding distance

A flat image doesn't say how far away anything is. Robots get depth in three main ways:

  • Stereo cameras compare two views, as eyes do.
  • Time-of-flight sensors time how long emitted light takes to return.
  • AI models estimate depth from a single image.

Software then combines depth with an object's outline to locate it in 3D, and a motion planner turns that into a path for the arm or wheels.

e-con Systems, which issues releases from California and Chennai, launched a time-of-flight camera, DepthVista Helix, on 17 February 2026. It claims under 1 percent deviation over 0.2–2 m and 0.5–6 m ranges, plus filtering that suppresses unreliable pixels. Those figures are the company's own.

Reasoning about space

The newer step is reasoning. Google DeepMind's Gemini Robotics-ER models take images or video, point at objects, count them and plan steps. In April 2026, DeepMind showed ER 1.6 improving on its predecessor, which miscounted tools and hallucinated a wheelbarrow in one test image. On 30 July 2026 it released ER 2 to developers through the Gemini API. It acts as a planning layer and hands motion to a separate action model.

DeepMind's own numbers are revealing. ER 2 scored 91.3 percent accuracy at locating a key moment in a video, such as when to stop pouring coffee. But it classified a robot's task progress into one of five bands with only 57.4 percent accuracy. These are company benchmarks, not independent audits. Demonstrations such as a Spot robot fetching a snack are not factory deployments.

Where the picture breaks

Glass and mirrors are the classic trap. Light passes through or bounces away, so sensors return missing or background distances, and a solid door can look like free space. The preprint behind the opening test, OptiGeo (University of Hong Kong and DiDi's Voyager Research, August 2026), trains a compact model to repair such errors. Across three lab scenes, navigation success rose from 10.0 to 78.3 percent. The authors say it can still fail on large glass doors and windows, and it needs extra calibration for true metric scale. It is a research preprint, not a product.

A robot's view of a room is a reconstruction, built from sensors and models that each have blind spots. When a company says its robot "sees," the useful follow-up is: through which sensors, in what conditions, and measured how? For India, e-con's work is a reminder that robot "eyes" are a component market, not only a finished-robot one.

Found this analysis insightful?

Share with colleagues, engineers, and your network.

Link copied to clipboard!