Every LLM you've used predicts the next token. Every image generator you've used predicts the next pixel (or the next denoising step). JEPA does neither. It predicts the next embedding: a compressed representation of what's missing, not the raw thing itself. That one design choice is the entire pitch, and it's worth understanding because it's the architecture behind the most well-funded new AI lab of 2026.
JEPA stands for Joint Embedding Predictive Architecture. Yann LeCun (Turing Award winner, formerly Meta's chief AI scientist) proposed it in a 2022 position paper called "A Path Towards Autonomous Machine Intelligence," and has spent the years since building it out: I-JEPA for images, V-JEPA and V-JEPA 2 for video, and now an entire company, AMI Labs, built around scaling the idea. If you've only encountered JEPA as a buzzword in the wake of LeCun leaving Meta, this is the deep dive that explains what it actually does.
The problem JEPA is trying to solve
Generative models (the GPTs, the Stable Diffusions, the Sora-likes) all share a training objective: reconstruct the input, or some corrupted version of it, at the level of raw pixels or tokens. An LLM is graded on whether it predicted the exact next word. A diffusion model is graded on whether it predicted the exact noise that was added to an image.
LeCun's argument is that this is wasteful. Reconstructing every pixel of a video frame means the model has to spend capacity on things that don't matter for understanding or prediction: the exact texture of grass blowing in the wind, the precise dither pattern in a low-light shot, the wallpaper in the background of a room. None of that helps you predict what happens next or plan an action. It's detail for detail's sake, and generative models burn enormous compute getting it right anyway, because that's literally what the loss function rewards.
JEPA's bet: don't reconstruct, predict — but predict in a space where irrelevant detail has already been thrown away. Compress first, then predict the compressed version of what's missing.
How JEPA actually works
A JEPA setup has a few essential pieces, and they're consistent across the image, video, and (eventually) other modalities LeCun's team has built:
- A context encoder takes the visible part of the input (part of an image, the first few seconds of a video) and compresses it into an embedding: a vector representation that keeps what's semantically relevant and discards the rest.
- A target encoder does the same thing to the missing part (the masked image region, the future video segment), producing a target embedding. This is what the model is trying to predict.
- A predictor takes the context embedding (plus, in some variants, a description of what's being asked, such as which region, how far into the future) and tries to predict the target embedding directly.
- The training signal is the distance between the predicted embedding and the actual target embedding, not pixel-level reconstruction error.
This is the "joint embedding" part of the name: both the context and the target get embedded into the same latent space, and the model learns to navigate between them. It's "predictive" because the whole point is forecasting one embedding from another: a missing image patch, a future video state, the outcome of an action.
The self-supervised objective is simple to state: mask or withhold part of the input, and train the model to predict the embedding of what you withheld, using only the embedding of what's still visible. No labels needed. No reconstruction of pixels needed. Just embedding-to-embedding prediction, at scale, over huge amounts of unlabeled video and images.
This matters because chain-of-thought prompting and other techniques you'd use with an LLM are fundamentally about getting a token-predicting model to externalize its reasoning in language. JEPA doesn't produce language at all by default: it produces embeddings. Prompting it, in the traditional sense, isn't really the interaction model. That's a genuinely different kind of system, which is why it's showing up in the glossary as its own term rather than a variant of an LLM.
I-JEPA and V-JEPA: the model lineage
I-JEPA (Image JEPA) was the first working implementation, released by Meta AI as open source. It masks blocks of an image and predicts their embeddings from a single visible context block. It performed well on standard vision benchmarks while being notably more compute-efficient than contrastive or generative vision pretraining approaches, because it never has to reconstruct actual pixels.
V-JEPA 2 extended the idea to video, and this is where it starts looking less like "a clever vision pretraining trick" and more like a world model. V-JEPA 2 was pretrained on over a million hours of public video plus a million images, using the same masking objective. But now the "missing" piece is a future segment of video rather than a masked patch. Meta reported strong results on motion-understanding benchmarks (77.3% top-1 on Something-Something v2) and action-anticipation tasks (39.7 recall-at-5 on Epic-Kitchens-100), alongside new physical-reasoning benchmarks (IntPhys 2, Minimal Video Pairs, and CausalVQA) designed specifically to test whether a model actually understands physical plausibility rather than just pattern-matching pixels.
The more interesting result is V-JEPA 2-AC, an action-conditioned variant. Post-trained on fewer than 62 hours of unlabeled robot video, it enables zero-shot planning: a robot arm it has never seen before, in a lab it has never seen before, using nothing but a monocular RGB camera, can perform reaching, grasping, and pick-and-place tasks: no calibration, no task-specific fine-tuning, no reward function. The model predicts future embeddings conditioned on candidate actions and picks the action sequence whose predicted outcome embedding is closest to the goal. That's planning by prediction in latent space, not by trial-and-error reinforcement learning in pixel space.
In March 2026, Meta released V-JEPA 2.1, which focuses on learning temporally consistent dense features (representations that hold up frame to frame rather than just at a coarse, global level). It reportedly improved real-robot grasping success rate by around 20% over V-JEPA 2-AC, largely by making the underlying features more spatially and temporally grounded.
Why this became a company, not just a research thread
In early 2026, LeCun left Meta after roughly a decade there and founded Advanced Machine Intelligence Labs (AMI Labs) to build world models around the JEPA architecture full-time. The company closed a seed round of roughly $1.03 billion, reportedly Europe's largest seed round on record, specifically to scale this approach rather than continue it as one research program inside a larger company.
The stated bet is pointed and a little combative toward the rest of the industry: LLMs, no matter how large, are trained on text, and text is a lossy, secondhand description of the physical world. A system that has only ever seen language will always be missing something a system trained directly on video and sensor data would pick up implicitly: how objects move, what's physically plausible, what an action actually does to a scene. AMI Labs is targeting robotics, industrial applications, and healthcare specifically because those are domains where "the model doesn't actually understand physics" is a real, practical blocker rather than an academic complaint.
Whether that bet pays off is a genuinely open question — you're reading this in September 2026, and the company is six months old. But the funding size alone is a signal that a meaningful chunk of the field now thinks token prediction isn't the only path forward, and that latent-space prediction is worth betting a billion dollars on.
How this fits with everything else you know about AI models
If your mental model of "how AI works" is built entirely around LLMs (chatting with a reasoning model, comparing GPT, Claude, and Gemini on the model comparison page), JEPA is a useful reset. It's not a bigger or better chatbot. It's a different answer to a different question: not "what word comes next," but "what state comes next, and what does that state mean for what I should do."
The two other posts in this series go deeper on the comparisons that matter most in practice: world models vs. LLMs unpacks the latent-space-vs-token-space distinction in more detail, and JEPA vs. diffusion vs. autoregressive models puts all three prediction strategies side by side so you can see exactly where each one spends its compute and what each one is actually good at.
The short version
JEPA predicts embeddings, not pixels or tokens. It's trained by masking part of an input (an image region, a future video clip) and learning to predict that missing part's compressed representation from what's still visible, with no reconstruction step and no labels required. I-JEPA proved it works for images; V-JEPA and V-JEPA 2 proved it works for video and, with action-conditioning, for zero-shot robot planning. And as of 2026, it's the architectural bet behind a billion-dollar startup betting that understanding the physical world matters more than predicting the next word.



