Most "how to use" posts about a new model assume there's a clean, stable API waiting for you. V-JEPA 2 mostly has one, for video understanding. For the robotics half, what you get is closer to a research checkpoint with an example notebook than a product. I'm going to be specific about which is which, because conflating them is how you end up debugging an import error for two hours before realizing the feature you wanted was never packaged for general use.
V-JEPA 2 is Meta FAIR's self-supervised video model: the second generation of the Joint Embedding Predictive Architecture (JEPA) applied to video. It was trained on internet-scale video with no labels, and Meta's paper ("V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning") reports state-of-the-art results on motion understanding and action anticipation benchmarks, plus an action-conditioned variant for robot manipulation. Here's what you can actually run today.
What V-JEPA 2 predicts — and what it doesn't
The "JEPA" in the name matters for how you use it. A JEPA model doesn't predict raw pixels for masked or future frames, and it doesn't predict discrete tokens the way an autoregressive video model would. It predicts embeddings: the encoder's own internal representation of the missing content. Given a masked view of a video clip, the predictor tries to output the embedding the encoder would have produced for the missing parts, working entirely in representation space.
That distinction is why V-JEPA 2 is described as good at "understanding" and "prediction" rather than as a video generator. You won't get a rendered next-frame image out of it. You get a vector that encodes what's semantically happening: useful for classification, retrieval, and, per Meta's paper, downstream planning, but not for making a video.
Setup
There are two ways to get V-JEPA 2 running: the standalone facebookresearch/vjepa2 GitHub repo (via torch.hub), or the Hugging Face transformers integration. For feature extraction and classification, transformers is the path of least resistance, so that's what I'll walk through.
pip install -U git+https://github.com/huggingface/transformers
pip install torchcodec numpy
torchcodec is what the official example uses for video decoding, worth noting if you're on macOS, since Meta's own repo README flags that its default decoder (decord) doesn't support macOS and recommends eva-decord or decord2 as substitutes if you go the torch.hub route instead.
The published checkpoints, from Meta's Hugging Face collection, include:
facebook/vjepa2-vitl-fpc64-256(ViT-L, ~300M params)facebook/vjepa2-vith-fpc64-256(ViT-H, ~600M params)facebook/vjepa2-vitg-fpc64-384(ViT-g, ~1B params, higher resolution)- Fine-tuned variants for classification, like
facebook/vjepa2-vitl-fpc16-256-ssv2(trained on Something-Something V2)
V-JEPA 2.1, a March 2026 update, adds checkpoints (ViT-B through ViT-G) with more temporally consistent dense features, useful if you're doing anything frame-by-frame, like tracking or dense video QA, where V-JEPA 2's features could flicker slightly between adjacent frames.
Getting embeddings from video
This is the core use case, and it's the one with an actual documented API. Here's the pattern straight from the Hugging Face model docs for vjepa2:
import numpy as np
from torchcodec.decoders import VideoDecoder
from transformers import AutoModel, AutoVideoProcessor
processor = AutoVideoProcessor.from_pretrained("facebook/vjepa2-vitl-fpc64-256")
model = AutoModel.from_pretrained(
"facebook/vjepa2-vitl-fpc64-256",
device_map="auto",
attn_implementation="sdpa",
)
video_url = "https://huggingface.co/datasets/nateraw/kinetics-mini/resolve/main/val/archery/-Qz25rXdMjE_000014_000024.mp4"
vr = VideoDecoder(video_url)
frame_idx = np.arange(0, 64) # 64 frames — matches this checkpoint's fpc64
video = vr.get_frames_at(indices=frame_idx).data # T x C x H x W
inputs = processor(video, return_tensors="pt").to(model.device)
outputs = model(**inputs)
# encoder embeddings — same as model.get_vision_features()
encoder_outputs = outputs.last_hidden_state
# predictor output (for masked/future prediction tasks)
predictor_outputs = outputs.predictor_output.last_hidden_state
last_hidden_state comes back shaped (batch_size, sequence_length, hidden_size): for vjepa2-vitl-fpc64-256, hidden_size is 1024. sequence_length is the number of spatiotemporal patches: the model tubelets frames (2 frames per tubelet by default) and patches each tubelet spatially (16×16 patches by default), so a 64-frame, 256×256 clip produces a few thousand patch tokens, each a 1024-dim embedding. If you just want a single vector per video for retrieval or clustering, mean-pool over the sequence dimension.
That embedding is the actual product here. It's not text, it's not a caption: it's a dense representation you'd feed into a downstream classifier, a similarity search index, or (per Meta's research) a planning module.
Video classification
For classification, use the fine-tuned checkpoints with AutoModelForVideoClassification instead of the bare encoder:
import numpy as np
import torch
from torchcodec.decoders import VideoDecoder
from transformers import AutoModelForVideoClassification, AutoVideoProcessor
hf_repo = "facebook/vjepa2-vitl-fpc16-256-ssv2"
model = AutoModelForVideoClassification.from_pretrained(hf_repo, device_map="auto")
processor = AutoVideoProcessor.from_pretrained(hf_repo)
video_url = "https://huggingface.co/datasets/nateraw/kinetics-mini/resolve/main/val/bowling/-WH-lxmGJVY_000005_000015.mp4"
vr = VideoDecoder(video_url)
frame_idx = np.arange(0, model.config.frames_per_clip, 8)
video = vr.get_frames_at(indices=frame_idx).data
inputs = processor(video, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits
top5 = logits.topk(5).indices[0]
probs = torch.softmax(logits, dim=-1).topk(5).values[0]
for idx, prob in zip(top5, probs):
print(model.config.id2label[idx.item()], float(prob))
Note the frame sampling changes with the model config: fpc16 checkpoints expect 16 frames, sampled with a stride, not the 64 used for the base encoder. Get this wrong and you won't get an error, you'll get silently degraded predictions, because the model will just process whatever frame count you hand it without complaint in most transformers pipelines. Always read model.config.frames_per_clip before you sample.
The -ssv2 suffix means fine-tuned on Something-Something V2, a benchmark specifically about fine-grained motion ("pushing something from left to right" vs. "pulling something from right to left") rather than object identity. That's the axis V-JEPA's embeddings are supposed to be strong on: it was trained to predict what happens, not just what's in frame.
Action-conditioned planning — what's real vs. what's a demo
This is where I want to be careful, because it's the part most likely to get overhyped. Meta's paper describes V-JEPA 2-AC: a version of the predictor post-trained on a relatively small amount of robot trajectory data (they report using data from a lab's existing robot arm datasets, not internet video) so that it can predict how the world's embedding changes conditioned on a specific robot action, not just conditioned on time passing.
The repo does publish an actual checkpoint for this: a single ViT-g/16-based vjepa2_ac_vit_giant predictor, loadable via:
import torch
encoder, ac_predictor = torch.hub.load(
'facebookresearch/vjepa2', 'vjepa2_ac_vit_giant'
)
But what's shipped alongside it is an example notebook (energy_landscape_example.ipynb) that computes an "energy landscape" (essentially, how well different candidate actions match a predicted goal state) using robot trajectory data collected in Meta's own lab setup. It is not a general-purpose "give it a goal image, get back a robot control sequence" API you can point at your own robot out of the box. There's no wrapped planning loop, no off-the-shelf environment interface, no documented way to bring an arbitrary robot arm's action space and have it just work. You'd be reimplementing the planning loop (typically something like model-predictive control: sample candidate action sequences, roll them through the predictor, score against a goal embedding, pick the best) using their predictor as one component.
So: if you want video embeddings or video classification, V-JEPA 2 is genuinely a "clone repo, pip install, run inference" experience today. If you want action-conditioned robot planning, what you can actually do right now is reproduce Meta's benchmark results and study their notebook's structure, not deploy it against your own hardware without substantial additional engineering. That's a fair place for a research release to be nine months post-launch; it's just not the same as a product API, and it's worth knowing which one you're signing up for before you start.
Where this fits
V-JEPA 2 isn't something you prompt: there's no text interface to steer it, which is a genuinely different interaction model from working with an LLM. If you're coming from prompting reasoning models or building agents that reason in language, the mental model doesn't transfer directly; goals here are images, video clips, or target embeddings, and outputs are predictions in that same embedding space. For the LLM side of things, our models page tracks current model capabilities, and the AI agents track covers how language-based planning systems work if you want the contrast.
If you're building something that needs both (an LLM to decide what to do and a world model to predict what happens physically if it does), that pairing is a distinct architecture worth understanding on its own, which is what I get into next.



