"Multimodal" gets used as if it's a single checkbox a model either has or doesn't. It isn't. In practice it's a per-vendor list of accepted input types, and that list is different for every major model family right now. If you're building something that needs to process video or audio, the model you pick determines whether that's even possible, not just how well it's done.
Here's where things stand as of this writing, across the three big frontier families.
What each family actually accepts
GPT-6 (the Astra, Sol, and Luna model line) takes text and images. That's it, natively, in the chat models. If you need video transcription or audio understanding from OpenAI, you're routing through separate realtime or transcription models, not the same model that's writing your code or answering your questions. It's a second API call, a second integration, a second thing that can fail.
Claude, across the current lineup (Fable 5.1, Opus 5.5, Opus 5, Sonnet 5, Haiku 4.5), takes text, images, and PDFs. The PDF support is worth calling out separately from images, because it isn't just "convert to an image and read it": Claude can process a PDF's text layer and its visual layout together, which matters for documents where the structure carries meaning, like tables, forms, and multi-column layouts. No video or audio input on any current Claude model.
Gemini (3.8 Flash and 3.5 Flash-Lite) is the outlier, and it's the only one of the three families that takes video and audio natively alongside text, images, and PDFs. Both models also carry a 1,048,576 input token limit, large enough to matter for video specifically. A lengthy video eats tokens fast once it's been converted to frames plus an audio track, and Gemini's context window is built to absorb that.
If your project needs to understand a video file or transcribe and reason over spoken audio in the same call that also handles text, Gemini is currently your only native option among the big three. That's not a preference, it's a hard constraint that should shape your model choice before you write a single prompt. Check model comparison if you're deciding between families for a specific project, and see the Gemini model guide for more on its multimodal specifics.
Why this matters more than people expect
Teams often prototype with whatever model they're already using for text, then discover mid-build that it can't touch their actual input data. If your product ingests security camera footage, sales call recordings, or user-submitted videos, that's a Gemini-shaped problem from day one, not a "we'll add multimodal later" afterthought. Retrofitting a video pipeline onto a text-only architecture is a bigger rewrite than picking the right model up front.
The inverse is also true: if your workload is PDFs (contracts, invoices, research papers), Claude's PDF handling is purpose-built for that, and you don't need a video capability you'll never use.
Prompting multimodal inputs well
Once you've got the right model for your input type, the second mistake is prompting it the way you'd prompt a search engine: "analyze this image" or "summarize this video." That's not wrong, exactly, it's just underspecified, and underspecified prompts on multimodal inputs produce vague, generic output far more often than they do on text alone.
The fix is the same instinct that makes text prompting better: say exactly what you want extracted, not just that you want something.
Weak: "What's in this image?" Better: "List every product visible in this shelf photo, with approximate quantity and whether the price tag is legible. Flag any items that look out of stock."
Weak: "Summarize this PDF." Better: "Extract the payment terms, termination clause, and any auto-renewal language from this contract. Quote the exact clause text, don't paraphrase."
Weak: "Analyze this video." Better: "Identify each distinct scene change with a timestamp, and describe what action is happening in each scene in one sentence."
The pattern underneath all three: name the specific fields or facts you need, specify the output format, and tell the model what to do with ambiguous or missing information rather than leaving it to guess. This is the same specificity discipline from earlier in this course. It just applies to pixels and waveforms instead of only to words.
A few things that trip people up specifically with multimodal inputs:
- Image resolution and cropping matter. A model reading small text in a screenshot does better if you tell it where to look ("read the text in the top-right corner") than if you make it scan the whole image for something small.
- PDFs with scanned, non-text-layer pages are effectively images to the model, even though the file extension says PDF. If extraction quality is poor, that's often why; ask the model to note when it's uncertain due to scan quality rather than silently guessing.
- Video and audio token cost adds up fast. A ten-minute video isn't a rounding error the way a short text prompt is. Budget for it, and consider whether you actually need the full video or just key frames plus a transcript.
- Don't mix unrelated asks in one multimodal prompt. "Describe this image and also write me a poem about it and also translate the poem" spreads the model's attention. One clear extraction task per call is more reliable than three tasks bundled together.
The takeaway
Multimodal capability is a moving target and it's vendor-specific, not a settled industry standard. Before you commit to a model for a project involving anything beyond text and static images, confirm what that specific model family actually accepts. Don't assume "multimodal" means "handles everything." And once you've picked the right model, treat multimodal prompting with the same specificity you'd apply to text: say precisely what to extract, in what format, and what to do when the input is ambiguous.
Next: the next lesson looks at a fundamentally different kind of architecture, one that doesn't predict tokens or pixels at all. See world models and JEPA.