Most AI today learned from text and pictures on the internet. That source is drying up. It also never taught any AI what happens when you drop a glass, or turn a corner too fast. So a new race has started. Google DeepMind, Meta, NVIDIA, and Fei-Fei Li's company World Labs are all building AI that learns how the real world works, not just what it looks like in a photo.
Here's why that matters. The internet has tons of pictures of a ball. But it has almost nothing that shows what happens one second after you throw it. A robot, a self-driving car, or a warehouse arm needs to predict that next second to work safely. That's a much harder kind of learning.
Three Different Ideas, Not One
People use the term "world model" loosely. But it really means three different things. Mixing them up leads to big, misleading claims.
The first idea is building 3D scenes. World Labs makes a tool called Marble. You give it a text prompt or a photo, and it builds a 3D scene you can walk through. Nothing in it moves unless you move it. This one is real, and you can use it today. World Labs launched Marble in late 2025. It opened its tool to outside developers in January 2026. Then in February 2026, it raised $1 billion, backed by Autodesk, NVIDIA, AMD, and others. Some reports said this valued the company at $5 billion. World Labs itself never confirmed that number.
The second idea is video simulation. A model predicts what a scene will look like one moment later, frame by frame. Google DeepMind's Genie 3 and NVIDIA's Cosmos both work this way. Cosmos isn't really one product. It's more like a toolkit: data and ready-made models other companies build on, to train robots or self-driving cars. NVIDIA calls its newest version, Cosmos 3, a step toward "World Action Models," one model that imagines a scene and decides what to do in it. That's NVIDIA's own claim, and nobody outside NVIDIA has proven it works at scale.
The third idea is closest to how scientists think real brains work. That's Meta's V-JEPA. Instead of generating a picture, it predicts a simplified, compressed version of what happens next. V-JEPA doesn't need to paint a full, realistic video just to know a cup is about to fall off a table, which saves a lot of computing power. Meta built a robot version called V-JEPA 2-AC. It tested it on robot arms in two labs that had never given it training data before. That's a real test of generalizing, not a demo picked to look good.
What's Working vs. What's Just a Demo
It helps to separate what's actually being used from what's just shown off. In February 2026, Waymo said it built its own tool on top of Genie 3, made for driving. The goal is to create rare, dangerous situations, a flooded street, an elephant on the road, so the self-driving system can train on them without facing them for real. This is a real tool, used inside a company already running self-driving cars for paying customers. But here's the catch: Waymo shared no test results or outside checks. So how much safer it actually makes things is still just Waymo's own claim.
Compare that to most Genie 3 demos you've seen. They show a few minutes of an explorable environment, built from one prompt. Impressive, but still a research demo. You can't just license Genie 3 and plug it into your product the way Waymo built on top of it.
What's a Claim, and What's Actually Checked
NVIDIA's "World Action Model" language, and its claims about how accurate Cosmos 3 is, come straight from NVIDIA's own reports. Some grading was even done by NVIDIA's own AI, built to judge if a video looks physically real. That's not an outside check. Meta's robot tests come closer to one, since the test labs' data was deliberately kept out of training. World Labs hasn't published outside tests proving Marble's 3D scenes hold up at bigger scale. Its momentum so far is measured in funding and user growth, not test scores.
What Could Happen Next
In the near future, expect these world models to mostly be used as training tools. They'll likely keep showing up for self-driving cars and warehouse robots, not as something regular people buy and use directly. That's already happening at Waymo and inside NVIDIA's robot partners.
Further out, the big open question is which of these three ideas, if any, leads to AI that understands the real world in a general way. Yann LeCun, Meta's former top AI scientist, left to start his own company chasing this idea. He believes predicting in a simplified form, the V-JEPA way, is the right path, and that generating full video is a costly detour. People at NVIDIA and DeepMind are betting the opposite. Nobody has proven either side right. What's real today is smaller than the pitch: AI that's learned enough physics to train other machines safely, not AI that understands the physical world the way a person does.