Everything covered so far in this track, whether it's an LLM or a reasoning model, shares one trait: it predicts the next token. Feed it text, it predicts more text, one piece at a time. That's an incredibly powerful design, but it's not the only way to build a predictive model, and it's arguably not a good fit for understanding things that aren't sequences of discrete symbols, like physical motion in the real world.
JEPA is a different answer to the same underlying question: how do you get a model to predict what happens next?
The core idea
JEPA stands for Joint Embedding Predictive Architecture, a design proposed by Yann LeCun. Instead of predicting the next token like an LLM, or generating pixels like a diffusion model, JEPA predicts the embedding, the internal representation, of masked or future content.
Here's the training setup in plain terms: take an image or a video, hide part of it (a patch of the image, or a future segment of the video), and ask the model to predict what the hidden part's representation looks like in embedding space. Not to reconstruct the actual pixels. Not to generate a plausible-looking image. Just to predict the abstract representation of what's missing, based on the context around it.
That distinction matters more than it sounds like it should. A model that has to reconstruct exact pixels ends up spending capacity on details that don't matter, like the precise texture of a carpet or the exact shade of a shadow. A model that only has to predict a representation can skip all of that and focus on the higher-level structure: what's likely to be there, how it relates to what's around it, what's about to happen. It's a bit like the difference between memorizing a photograph and understanding a scene.
Why this connects to "world models"
The phrase "world model" describes a system that has learned something like a causal, predictive understanding of how the physical world behaves, not just how words tend to follow other words. LeCun's argument, which drove the JEPA line of research, is that LLMs trained purely on text don't build this kind of understanding. They're extremely good at manipulating language, but language is a lossy, compressed description of the world, not the world itself. A model that has only ever seen text has never seen an object fall, a hand grip a cup, or a car turn a corner.
JEPA architectures are trained directly on visual and video data with this gap in mind. By learning to predict representations of masked video segments, a JEPA-style model builds up something closer to an intuitive physics: what a falling object does next, what a scene looks like a few seconds later, what actions are even possible from a given state. That's the "world model" part: not a chatbot that talks about the world, but a system whose internal predictions are shaped by how the world actually behaves.
V-JEPA 2, the flagship example
Meta's V-JEPA 2 is the clearest concrete example of this approach in action. It's trained on video, and it's been shown to do three things worth knowing about:
- Video understanding: reasoning about what's happening in a video clip.
- Video prediction: predicting representations of what happens next in a sequence, without generating actual future frames.
- Zero-shot robot planning after fine-tuning: because the model has learned something like intuitive physics from video, it can be adapted to help a robot plan physical actions, like reaching for or moving an object, without task-specific training on that exact robot behavior.
That third capability is the one that gets attention, because it's a concrete demonstration of the "world model" argument actually paying off: a model trained on passive video observation transferring to active physical planning.
Why this is having a moment
The person most associated with this research direction, Yann LeCun, left Meta in 2026 to found AMI Labs (Advanced Machine Intelligence Labs), raising roughly $1.03 billion specifically to pursue world models as a research bet distinct from scaling up LLMs further. That's a meaningful signal: a researcher with LeCun's track record betting a nine-figure raise on the idea that predicting-the-next-token is not the path to machine intelligence that understands the physical world, and that predicting representations is.
You don't need to take a side in that debate to get the practical point of this lesson: JEPA-style world models are a genuinely different architecture from anything else in this track, built for a different job (physical and visual understanding) using a different training signal (representation prediction instead of token or pixel prediction).
Where to go deeper
This lesson is deliberately compact. The mechanics of JEPA, how it compares to diffusion and autoregressive models, how to actually prompt or interact with a world model, and what it means for robotics specifically are each substantial topics on their own. If you want the full picture, start with what is JEPA, the pillar post for our full JEPA cluster, which links out to deeper pieces on world models versus LLMs, using V-JEPA 2 hands-on, prompting world models, and the AMI Labs story.
Next: architecture is only half the picture. The other half is who controls the weights once a model like this ships. See open vs closed model weights.