"World model" has become one of those terms that gets applied to three genuinely different things, which makes it hard to tell what's actually shipping versus what's a research demo with a good video attached. So let's split it apart properly. A world model, broadly, is a system that learns to predict how an environment changes over time: given a current state and an action (or just given the past), what happens next. That definition covers Meta's JEPA-style predictive models, the simulators autonomous-driving companies use to test policies, and the "playable video" systems like Genie that generate interactive game-like worlds frame by frame. They share a lineage but solve different problems with different levels of maturity, and conflating them oversells the least mature one.
Here's where each actually stands as of late 2026.
Robot control and manipulation: early production, real gap over baselines
This is the most concrete, benchmarked use of world models right now, largely because Meta has been unusually open about publishing results. V-JEPA 2, trained on more than a million hours of public video, is a video-based world model used for two related things: understanding and predicting what happens next in video, and (via a variant called V-JEPA 2-AC) planning robot actions.
The mechanism is worth understanding because it's genuinely different from how most robot-learning systems work. Instead of training a policy to imitate labeled demonstrations end to end, V-JEPA 2-AC uses its learned world model for model-predictive control: it defines an "energy" score as the distance between a predicted future state and a goal state, then searches over candidate action sequences (using a method called the cross-entropy method) to find the sequence that gets closest to the goal. Critically, the underlying world model was trained on passive video: it never needed action labels for the bulk of its training. Only a small amount of robot interaction data (Meta reported 62 hours) was needed to fine-tune it for actual control.
The results that matter: in published benchmarks, V-JEPA 2-based control reportedly achieved around 80% success on tasks like picking up and relocating an object in a new environment, compared to roughly 15% for Octo, a comparable video-language-action baseline model, on the same unfamiliar setup. That's a real, published gap (not a demo reel claim) and it's specifically about generalization to environments and objects the model wasn't trained on, which is the hard part of robotics.
What this isn't yet: a production robot stack you can buy. It's a research model with strong benchmark numbers, released for the community to build on, not an off-the-shelf manipulation system running in warehouses today. Several robotics labs are reportedly exploring JEPA-style world models as a planning backbone rather than end-to-end imitation learning, but broad production deployment of this specific architecture is still ahead of it, not behind it. Treat any claim of large-scale commercial robot deployment running on V-JEPA-style world models as unverified until a specific company states it.
Autonomous driving: world models as simulators, not (yet) as the driving policy
Driving is where "world model" gets used loosely, so it's worth being precise about what's actually deployed. The clearest, most verifiable 2026 use case is simulation and evaluation, not the live decision-making policy in the car.
Waymo has published its own "Waymo World Model", a large generative model used to create hyper-realistic driving simulations at scale, letting their driving software encounter rare and dangerous scenarios in simulation millions of times before ever seeing them on a real road. That's a world model used as a training and validation environment, which is a mature, already-in-production use of the concept: it just isn't the same as the world model directly steering the car.
Wayve, a UK autonomous driving company, has taken this further with its GAIA series of models, and GAIA-3 specifically is described as extending world modeling from pure visual synthesis toward genuine autonomy evaluation, recreating realistic traffic dynamics, including rare safety-critical events, for testing driving policies against. Wayve has been explicit that this generative-simulation approach is central to how it validates and iterates on its driving stack, and the company has been pushing toward licensing this kind of software to automakers rather than only running its own fleet.
Where things get less verifiable: whether any production driving policy (Tesla's FSD, Waymo's Driver, Wayve's stack, or others) is directly using a JEPA-style predictive world model as its live planning core, as opposed to using world models for simulation and offline training. Public information supports the simulation/evaluation use case clearly. It does not clearly support claims that a JEPA-architecture world model is the real-time decision engine behind any commercially deployed robotaxi or FSD system today, that's a meaningfully different claim, and treat it skeptically if you see it asserted without a named source. The honest summary: world models are already load-bearing infrastructure for how self-driving systems get trained and tested; they are not confirmed to be the thing making the driving decision in the car, in production, yet.
Interactive and generative video: fast-moving, but distinguish two different techniques
This is the category most likely to get miscategorized, because "world model" and "video generation" get used almost interchangeably in coverage, and they're related but not the same thing.
Most of what you've seen under names like Sora 2 are diffusion-based video generation models: trained to generate plausible video (or extend/edit existing video) by iteratively denoising, conditioned on a text or image prompt. They're generative in the classic sense (producing pixels) and extremely good at photorealism and stylistic range, but they're not necessarily learning a controllable, interactive, action-conditioned model of an environment you can steer moment to moment.
The other lineage (Google DeepMind's Genie series is the clearest public example) is explicitly built as an interactive, action-conditioned world model: the system learns a latent space of possible "actions" from watching large amounts of unlabeled video, then lets a user or agent control a generated environment frame by frame in something closer to real time. Genie 3, per DeepMind's own description, generates navigable environments at 720p and 24 fps with minutes of coherent, playable consistency: you can move around inside it, not just watch it play out. Other systems in this space (Oasis, GameNGen, Matrix-Game, and several open-weight variants) target similar real-time, playable generation with varying degrees of physical consistency and how long the "world" stays coherent before it drifts or breaks.
The precise distinction: diffusion video generation optimizes for the best-looking output given a prompt; interactive world models optimize for a consistent, controllable environment that responds correctly to input over time, even if the visual fidelity is lower. JEPA-style architectures are conceptually closer to the second category's goals (predicting consistent, controllable state) though most current interactive video systems, including Genie, are still built on generative/diffusion-adjacent techniques rather than JEPA's embedding-prediction approach specifically. This space is genuinely research-stage, the "minutes of coherence" framing is itself an admission that long-horizon consistency is still the open problem, not a solved one. Don't read demo videos of these systems as evidence of production readiness; they're closer to compelling research previews than shipped products.
The maturity ranking, plainly stated
If you need the one-paragraph version: robot manipulation is the most benchmarked and closest to practical use, with real published numbers showing meaningful generalization gains over baselines, but it's still research-grade rather than a bought-and-deployed product. Driving world models are genuinely in production, but specifically as simulation and evaluation infrastructure, the leap to a world model as the live driving policy is not something you should assume has happened just because the term gets used near driving companies. Interactive video generation is the least mature of the three in terms of real-world deployment, moving fast on research benchmarks (frame rate, resolution, coherence duration) but still firmly in the demo-and-paper stage rather than anything resembling a shipped consumer or enterprise product.
None of this requires you to change how you prompt or build with today's LLM-based tools, see our models page and glossary if you need current terminology, but if you're tracking where the frontier is moving, robotics is the place to watch first. It's the use case with the clearest evidence that predicting an abstract world state, rather than generating raw pixels or tokens, produces something measurably better than the alternative.
Sources: Meta AI (V-JEPA 2, Waymo) The Waymo World Model, Wayve (GAIA-3, Genie (world model)) Wikipedia.



