Nearly every major AI model you've used (text, image, or video) is doing some version of the same thing: predicting something it hasn't seen yet, given something it has. The differences that actually matter aren't in that high-level description. They're in exactly what gets predicted, in what representation, and in what order. Three architectures dominate right now, and they answer those questions in genuinely different ways: autoregressive models, diffusion models, and JEPA.
The three objectives, side by side
Autoregressive models (the architecture behind essentially every current LLM: GPT, Claude, Gemini) predict the next discrete token in a sequence, one at a time, each prediction conditioned on everything generated so far. The training objective is straightforward: cross-entropy loss against the actual next token. Generation is inherently sequential (token 500 depends on token 499, which depends on token 498, all the way back) because each step's input includes every previous step's output. Training is parallelizable (teacher forcing lets you compute the loss for every position in a sequence at once), but inference is not: you generate one token, then the next, then the next, and that sequential dependency is why long generations take longer and why speeding up LLM inference is its own subfield.
Diffusion models (the architecture behind most image and video generators, and, notably, image-generation heavyweights like the products referenced across this site's model comparison) predict noise, not content, directly. Training corrupts an image with random noise added in steps, and the model learns to predict the noise that was added at each step. Generation runs the process backward: start from pure random noise and repeatedly subtract the model's predicted noise, denoising over many steps until a coherent image emerges. Unlike autoregressive generation, this happens in parallel across the whole image at each step, with no left-to-right ordering, which is part of why diffusion models are so good at maintaining global coherence across an entire scene rather than a strict linear read order.
JEPA (Joint Embedding Predictive Architecture: see the full deep dive if you haven't read it yet) predicts neither tokens nor pixels. It predicts the embedding of missing information (a compressed latent representation of a masked image region or a future video segment) from the embedding of what's visible. There's no reconstruction step at all: the model is never graded on how close its output is to the actual pixels or tokens, only on how close its predicted embedding is to the target embedding. It throws away whatever detail doesn't matter for prediction before it even starts predicting.
What each one is actually optimizing for
This is the part that explains why the three produce such different capabilities, and it's worth being precise rather than hand-wavy about it.
Autoregressive training optimizes for getting the exact next symbol right, symbol by symbol, in a fixed left-to-right order. That's an excellent fit for language, where meaning really is built up sequentially and the vocabulary is a manageable discrete set (tens of thousands of tokens). It's a much worse fit for anything continuous and high-dimensional, like raw pixels, which is why nobody builds image generators by predicting pixels one at a time in raster-scan order; it technically works (early PixelRNN/PixelCNN-style models did exactly this) but it's slow and doesn't capture global structure well.
Diffusion training optimizes for recovering a whole coherent sample from corrupted input, without committing to any fixed generation order. Because the model works on the entire image at every step, it naturally enforces global consistency: objects don't get generated in isolation with no awareness of the rest of the scene. The cost is that the model is trained to reconstruct at pixel-level fidelity, which means a huge fraction of its capacity goes toward getting fine detail right, whether or not that detail matters for anything downstream. A diffusion model spends real training signal getting individual blades of grass or skin pores correct, because pixel-level accuracy is literally the loss function.
JEPA training optimizes for predicting what matters and only what matters, because the encoder that produces the target embedding has already compressed away low-level detail before prediction even happens. There's no pixel-level or token-level reconstruction loss anywhere in the objective: the model is never rewarded for getting fine detail right, because fine detail was never part of what it's asked to predict. This is the efficiency argument LeCun makes for JEPA: generative approaches (both autoregressive and diffusion) waste capacity reconstructing things a good world model shouldn't need to reconstruct at all, if the goal is understanding and prediction rather than photorealistic output.
That last point is the crux of the comparison. If your goal is producing content (an image, a video, a piece of text a human will read), you need a model that ultimately outputs something in the original space (pixels, tokens), so autoregressive or diffusion approaches are the natural fit, and current results bear that out: diffusion dominates image/video generation, autoregressive dominates text generation. If your goal is understanding or predicting state (what will this video show next, will this action succeed, is this physically plausible), you don't need pixel-perfect output at all, you need an accurate compressed representation of what's going to happen, and that's exactly what JEPA is built to produce cheaply.
A concrete comparison: predicting a falling glass
It's easier to see the difference with one example run through all three architectures.
- An autoregressive video model (predicting frames as sequences of tokens) would generate the falling-glass video frame by frame, token by token within each frame, conditioning each new token on everything generated before it — accumulating error the way long LLM generations can drift off-topic, except here it can drift into physically implausible motion.
- A diffusion video model would denoise the entire video (or large chunks of it) toward a plausible falling-glass sequence, enforcing frame-to-frame and within-frame coherence throughout the denoising process, and outputting an actual watchable video of a glass falling and shattering.
- A JEPA model, given video of the glass starting to tip, would predict the embedding of what the scene looks like a second later: not a video, not pixels, just a compressed representation that (ideally) encodes "the glass is now on the floor, broken, in roughly this configuration." It never has to render a single pixel to make that prediction, which is exactly why V-JEPA 2's action-conditioned variant can do zero-shot robot planning cheaply: it's evaluating candidate actions by predicting their embedded outcomes, not by rendering full video for each candidate and inspecting it.
That third case is the one with no real analog in the other two architectures, and it's why JEPA gets talked about as a distinct thing from "yet another generative model" rather than a video-generation competitor to diffusion.
Where the categories blur
None of this is perfectly clean in practice. There's active research on diffusion-based language models that apply a diffusion-style objective to text instead of autoregressive next-token prediction: early comparisons suggest diffusion language models can do better on tasks requiring global consistency and long-range coherence, while autoregressive models stay competitive on short, sequential reasoning. And "diffusion forcing" approaches blend next-token-style sequential generation with full-sequence diffusion within a single model. The three-way split (autoregressive / diffusion / JEPA) is the clearest way to think about the current landscape, but treat it as three points on a spectrum of "how much do you commit to sequential order, and how much do you commit to reconstructing raw output" rather than three hermetically sealed categories.
Which one you're actually using
If you're prompting Claude, GPT, or Gemini, anything you'd typically reach for using chain-of-thought or effort-setting techniques for reasoning models, that's autoregressive, full stop. If you're generating an image or video from a text prompt, that's almost certainly diffusion. If you're reading about robots learning to grasp unfamiliar objects with no task-specific training, or about world models predicting physical outcomes without rendering anything, that's JEPA — and as of 2026, it's the architecture an entire billion-dollar company is betting will matter more than the other two for anything that has to operate in the physical world rather than just talk about it.
For the deepest look at how JEPA itself works (the encoder/predictor setup, I-JEPA and V-JEPA 2 specifically, and why LeCun left Meta to build a company around it), see what is JEPA. The glossary also has quick definitions for JEPA, world model, and latent space if you just need the short version.



