I kept catching myself writing "prompt V-JEPA to predict..." while drafting other posts, and every time I had to stop and correct it, because it's wrong in a way that matters. World models like V-JEPA don't have a text input. There's no chat box, no system message, no instruction you type in English that changes what they do. If you're used to prompt engineering for LLMs (wording things carefully, giving examples, setting a role), none of that transfers here, and pretending it does will send you looking for an API that doesn't exist.
So what does "steering" a world model actually look like? And if you want language-level control over what it does, how do people actually wire that up? That's what this post is about.
Why there's no text box
A world model like V-JEPA 2 is trained on video, not on video-caption pairs the way an image generator is. It learns by masking parts of a video clip and predicting the missing embedding from the visible context: pure self-supervision on pixels, no language anywhere in the training loop. There's nothing in its weights that maps the string "make the arm move left" to anything. It was never taught that mapping.
Contrast that with an LLM, where every token you type is literally the input format the model was trained on. A world model's native input format is a visual observation (a frame or clip) and, for the action-conditioned variants, a proposed action. Its native output is a predicted future embedding: what the world (or the sensor's view of it) will look like, in representation space, after that action happens.
So "prompting" a world model, in the loose sense of "giving it something that steers its output," means one of:
- A goal image or goal video clip: "get to a state that looks like this."
- A candidate action sequence: "if I do these motor commands, what happens?"
- Some combination of both, scored against each other.
None of that is text. It's pixels and numbers.
Goal-conditioning: showing, not telling
The most common pattern in the world-model literature is goal-conditioning: you hand the model a target observation (a photo or short clip of the end state you want) and the planner's job is to find the action sequence that gets the predicted future embedding as close as possible to the goal embedding.
Concretely, for a system like V-JEPA 2-AC, the loop looks like this:
- Encode the current observation into an embedding.
- Encode the goal image into an embedding (same encoder, so they're comparable).
- Sample a batch of candidate action sequences.
- For each candidate, roll the action-conditioned predictor forward step by step, producing a predicted embedding trajectory.
- Score each candidate by distance between its final predicted embedding and the goal embedding: this scoring function is often called an "energy" in the JEPA literature, since lower distance means lower energy, i.e. a better match.
- Execute the first action of the best-scoring sequence, re-observe the real world, and repeat. That's the receding-horizon part of model-predictive control (MPC): you never commit to the whole plan blind, because your predictions drift from reality the further out you project.
This is exactly the "energy landscape" framing in Meta's own V-JEPA 2-AC example notebook, and it matches the general shape of MPC-style planning used across the goal-conditioned world model literature: sampling-based action optimization (methods like the cross-entropy method are common) scored against a learned dynamics model, executing one step, then replanning. The "prompt," if you insist on the word, is the goal image plus the cost function. Nobody's typing sentences into this loop.
Where language re-enters: the LLM-as-planner pattern
Here's the part that actually matters for anyone coming from the prompting world. You clearly want to be able to say "clean up the table" and have a robot figure out what that means physically. That request is language, and somewhere it has to get turned into pixels and actions. The way current systems bridge that gap is hierarchical, not monolithic: no one is training a single model that takes English in one end and motor torques out the other.
The pattern, consistent across the hierarchical-planning and LLM-agent literature, splits the problem into two layers with a genuine division of labor:
- High-level layer (LLM): takes the natural-language task, does the reasoning and decomposition ("clean up the table" becomes "pick up cup, move to sink, pick up plate, move to sink, wipe surface") and outputs a sequence of subgoals. This is squarely a language and reasoning problem: it needs world knowledge about what "cleaning up" typically involves, common sense about object affordances, and the ability to sequence and revise steps. It's the same skill set covered in prompting reasoning models: decomposition, step-by-step planning, checking its own intermediate reasoning.
- Low-level layer (world model): takes each subgoal, converts it into a target visual state (a goal image representing "cup is in sink"), and handles the physical prediction and control problem: what sequence of continuous motor actions gets from the current observation to that goal state. This is a perception-and-dynamics problem, not a reasoning problem, and it operates in embedding space at a much higher frequency (real-time control loops) than the LLM's occasional subgoal updates.
The division isn't arbitrary: it's playing to what each architecture is actually good at. LLMs are strong at symbolic decomposition and weak at continuous physical prediction (ask one to predict exactly where a dropped ball lands and watch it guess). World models are the reverse: strong at predicting how a scene evolves under an action, with no notion of "task" or "why." Gluing them together (LLM sets the subgoal, world model figures out the physics to get there, LLM checks progress and issues the next subgoal) is how you get both halves of the competence without training one model to do everything.
This is the same high-level/low-level split you'd recognize from classical robotics (symbolic planner plus continuous controller), just with an LLM standing in for the symbolic planner and a learned world model standing in for hand-coded dynamics. The research literature on this frames it explicitly as a "translation problem": abstract intentions have to get translated into something a continuous controller can execute, and that gap is exactly where most of the current research effort in this pairing is going, because it's genuinely unsolved in general, not a wiring detail.
What this means practically, today
If you're building with these pieces right now, the honest state of things is:
- The LLM side is mature. You can already build a subgoal-decomposition layer using any current model: this is a straightforward application of the same planning and tool-use patterns covered in the AI agents track. The LLM doesn't need to know anything about world models; it just needs to output subgoals in a format (text description, or a reference image it selects/generates) that the next layer can consume.
- The world-model side is much less packaged. As covered in our V-JEPA 2 walkthrough, Meta's action-conditioned checkpoint ships with an example notebook, not a general planning library. You'd be implementing the MPC loop yourself, and hooking it to your own robot's action space is nontrivial engineering, not a config change.
- The glue between the two layers (turning an LLM's text subgoal into a goal image the world model can condition on) is itself an open research problem. Some approaches use a separate image generator to render the goal; others retrieve a matching image from a library of known states. There's no standard off-the-shelf component for this yet.
If you want to experiment today, the realistic path is: use an LLM for the subgoal decomposition (fully doable now), and treat the world-model planning loop as a research project you're implementing from a paper, not a library you're importing. That's a very different timeline than most LLM tooling, and it's worth setting expectations accordingly before you scope a project around it.
The takeaway
"Prompting" is a language-model concept. World models take goals and actions, not words, and the way you steer one is by showing it what you want (a goal observation) and letting a search process (not a sentence) find the path there. If you want language-level control over a physical task, you're not prompting the world model at all — you're prompting an LLM that plans in language, and handing its output to a completely different system that plans in pixels. Keeping those two loops mentally separate will save you from a lot of wrong assumptions about what either half can do on its own.



