"World model" got thrown around loosely for a couple of years: people used it for anything that seemed to have some internal representation of how things work, including LLMs themselves. That's gotten more precise in 2026, mostly because Yann LeCun left Meta to found a billion-dollar company explicitly built around the distinction. It's worth being exact about what separates the two, because they're not competing on the same axis. They're doing genuinely different jobs.
The one-sentence version
An LLM predicts the next token in a sequence of text. A world model predicts the next state (represented as an embedding, not as text), given the current state and, often, a candidate action. Same general shape (predict what comes next from what came before), completely different substrate for what "next" means.
That sounds like a small distinction. It isn't. It determines what each kind of system is actually good at.
What LLMs are doing under the hood
An LLM (GPT, Claude, Gemini, whatever's current when you're reading this) is trained autoregressively: given a sequence of tokens, predict the next one, then feed that back in and predict the one after, and so on. The training signal is cross-entropy loss against the actual next token in the training data. Everything the model "knows" about the world, it knows because that knowledge was compressed into text somewhere in its training corpus and it learned the statistical patterns of how that text is structured.
This is an enormously general and enormously useful trick, precisely because so much of human knowledge is already in text form. But it has a structural limitation: an LLM's model of "the world" is a model of what people wrote about the world, filtered through language. It's never seen an object fall. It's read millions of descriptions of objects falling, in physics textbooks and news articles and fiction, and it's very good at producing plausible continuations of those descriptions. That's not nothing, but it's a secondhand, lossy representation of physical reality, mediated entirely through the bottleneck of language.
If you've worked with chain-of-thought prompting or reasoning models, you've seen this play out directly: the model's "reasoning" is itself a sequence of tokens it generates and then conditions on, which is powerful for problems that are naturally expressible in language or symbols, and much weaker for problems that are fundamentally about physical dynamics: how a fluid moves, what happens when a robot arm's gripper is 2mm off target, whether a stack of boxes is actually stable.
What world models are doing under the hood
A world model, in the specific sense LeCun and AMI Labs mean it, built on the JEPA architecture, is trained differently from the ground up. Instead of predicting the next token in a text sequence, it predicts the next embedding in a latent space that represents visual or physical state. Feed it video, mask out a future segment, and train it to predict that segment's embedding from the embedding of what came before. No text involved at any point. No tokens. The training data is video and images (direct observation of the physical world), not descriptions of it.
The practical unit of "state" in a world model like V-JEPA 2 is a compressed vector representation of a scene: what's semantically relevant about it (object positions, motion, physical relationships) with the irrelevant pixel-level detail thrown away. The model predicts how that compressed state evolves, optionally conditioned on a candidate action, which is what makes action-conditioned variants like V-JEPA 2-AC usable for planning. You don't ask a world model "what happens next?" as a text question. You give it a current state and a candidate action, and it predicts the resulting state directly in latent space, which you can then use to evaluate whether that action gets you closer to a goal.
That's the core mechanical difference: an LLM's "next" is the next token in a linguistic sequence. A world model's "next" is the next state in a physical or visual sequence, represented as an embedding.
Why this isn't "world models are better"
It's tempting to read all this as "world models fix what's wrong with LLMs," and that's not the right takeaway. They're built for different jobs.
LLMs are extraordinary at anything that's naturally expressed in or reducible to language and symbols: writing, code, structured reasoning, synthesizing information from disparate text sources, holding a conversation, following complex multi-step instructions. Most of the AI agents built today (coding assistants, research agents, customer support bots) are fundamentally LLMs orchestrating tool calls, and that's the right tool for that job because the interface to the world in those cases is language and structured data (APIs, function calls, files).
World models are built for tasks where the thing you're predicting is physical state, not language: what a video shows next, whether an action will succeed, how a robot arm should move to grasp an object it's never seen before. V-JEPA 2-AC's headline result (zero-shot pick-and-place on unseen robot arms, using under 62 hours of robot video for post-training, with no reward function or task-specific fine-tuning) isn't a task an LLM is well-suited to at all. There's no clean way to express "predict the visual outcome of this exact gripper motion" as a next-token prediction problem over text.
So the honest framing is: LLMs operate over the space of things humans have written down. World models operate over the space of things that can be observed and acted on. Where those spaces overlap (physical common sense embedded in text), LLMs get some of it for free but shallowly. Where they don't overlap (precise physical dynamics, robot control, video-native reasoning), you need something trained directly on the modality, which is what world models are for.
Is this actually a rivalry?
Not really, or at least not yet in practice — it's more that 2026 is the year "world model" stopped being a vague aspiration and became a specific, fundable architecture with real benchmark results behind it. LeCun's departure from Meta to found AMI Labs, and the roughly $1.03 billion seed round behind it, is a bet that this distinction matters enough commercially to build an entire company around, specifically targeting robotics, industrial automation, and healthcare, domains where physical-world understanding is the actual bottleneck, not language fluency.
Whether the two approaches eventually merge (an LLM that reasons in language but plans using an embedded world model for physical tasks, for instance) is an open architectural question nobody's settled yet. For now, treat them as separate tools solving separate problems: reach for an LLM (see the current landscape on /models) when the task lives in language, and understand that world models exist as a distinct, increasingly well-funded category when the task lives in physical or visual state.
If you want the full mechanical comparison, including how JEPA's embedding-prediction objective differs specifically from diffusion's denoising objective, not just from autoregressive token prediction, the three-way comparison lays out all three prediction strategies side by side. And if you haven't read it yet, what is JEPA is the deep dive on the architecture itself, including how V-JEPA 2 and V-JEPA 2.1 actually work under the hood.



