Reasoning Prompts
V-JEPA Experiment Planner
A prompt for a chat model (Claude or GPT-6) that helps you plan a V-JEPA 2 project: checkpoint choice, sampling strategy, fine-tuning approach, and evaluation.
Prompt
I'm planning a project using Meta's V-JEPA 2 (a self-supervised video world model). Help me design the approach. MY USE CASE: [DESCRIBE THE TASK, e.g. "classify defect types in manufacturing line video" or "detect unusual behavior in warehouse security footage"] WHAT I HAVE: - Video data: [DESCRIBE — how much, what resolution, what format, is it labeled?] - Compute available: [e.g. "single A100", "cloud, budget-flexible", "CPU only for inference"] - Latency requirement: [real-time / near-real-time / batch/offline is fine] Based on this, recommend: 1. Which V-JEPA 2 checkpoint to start from (ViT-L/H/g, and whether I need the base encoder or a fine-tuned classification variant) and why, given my compute and accuracy needs 2. A frame sampling strategy (frame count, stride) appropriate for my task and checkpoint 3. Whether I need to fine-tune at all, or whether the pretrained embeddings plus a lightweight classifier head on top would work for my use case 4. An evaluation plan: what metric fits my task, how to build a labeled validation set if I don't have one, and a sanity-check baseline to compare against 5. The biggest risk in this plan and how I'd catch it early Be specific about tradeoffs, not just recommendations. If two checkpoint choices are close, say what would tip the decision either way.
How to use
This is not a prompt for V-JEPA 2 itself. V-JEPA doesn't take text prompts; it takes video and returns embeddings, not a text response. This prompt is for a chat model like Claude or GPT-6 to help you plan a project that uses V-JEPA 2 as a component. Paste it in when you're at the "which checkpoint, what approach" stage, before you write any code. See the V-JEPA 2 hands-on guide for the actual implementation once you have a plan.
Variables
[DESCRIBE THE TASK]— be as specific as possible; "classify X" plans very differently from "detect anomalies in Y"[DESCRIBE — how much, what resolution...]— your data constraints shape the whole recommendation[compute available]— determines whether ViT-g is realistic or you should stay with ViT-L[latency requirement]— real-time changes the frame-count and model-size recommendation significantly
Tips
- Give the model your labeled-data situation honestly. "I have 50 labeled clips" changes the recommendation (lightweight head, heavy reliance on pretrained embeddings) very differently than "I have 5,000 labeled clips" (real fine-tuning becomes viable).
- If your use case sounds like it needs action-conditioned planning (predicting outcomes of specific actions, not just classifying what's happening), say so explicitly. As of today that's checkpoint-and-research-notebook stage rather than a packaged API, and the model should tell you that rather than implying a polished path exists. See how to prompt a world model for what's actually available there.
- Ask a follow-up: "what would make you change this recommendation?" A good answer surfaces the assumptions baked into the first plan.