Research Prompts
Gemini Video Analysis System Prompt
A system prompt for using Gemini 3.8 Flash to analyze video, the only current frontier model that takes video as native input.
Prompt
You are a video analyst reviewing [VIDEO TYPE, e.g. "customer support call recordings" or "product demo videos"] for [PURPOSE, e.g. "quality review" or "highlight extraction"]. For everything you report: - Reference the exact moment using MM:SS timestamps, so a human can jump straight to it. - Distinguish what's said from what's shown. If something is only visible on screen but never spoken, say so explicitly, and vice versa. - If a moment is ambiguous (unclear audio, occluded view, fast cut), say so rather than guessing at what happened. Output format: Return findings as a numbered list. Each item: [MM:SS] — one-sentence description — category tag. Categories to use: [LIST YOUR CATEGORIES, e.g. "decision, objection, action-item, unresolved-question"] Do not summarize the whole video in prose before the list. Give the list directly.
How to use
Set this as the system_instruction when calling Gemini 3.8 Flash on a video file. Gemini is the only one of the major frontier models that takes video directly as input right now, so this prompt is written around that specific capability rather than being a generic "describe this video" instruction.
from google import genai
client = genai.Client()
video = client.files.upload(file="recording.mp4")
# wait for video.state.name == "ACTIVE" before using it
response = client.interactions.create(
model="gemini-3.8-flash",
system_instruction=SYSTEM_PROMPT, # the filled-in prompt above
input=[
{"type": "video", "uri": video.uri, "mime_type": video.mime_type},
{"type": "text", "text": "Analyze this recording."},
],
)
Variables
[VIDEO TYPE]— what kind of video this is (call recording, demo, meeting, security footage)[PURPOSE]— why you're analyzing it, which shapes what counts as a "finding"[LIST YOUR CATEGORIES]— a fixed tag set makes downstream filtering easier than free-text categories
Tips
- Gemini samples video at roughly 1 frame per second by default. A fast on-screen change (a quick UI click, a flashed error message) can fall between sampled frames — if that matters for your use case, mention it explicitly and ask the model to flag low-confidence visual claims.
- Timestamps in the output are a real audit trail, not decoration. Spend a few test runs confirming they're accurate before trusting them in production; ask a follow-up question referencing a specific timestamp to sanity-check.
- For long videos (an hour or more), consider a two-pass approach: a first pass asking only for a table of contents with timestamps, then a second pass zooming into the sections that matter, rather than one pass trying to extract everything at once.