Once you understand that models come in families and tiers (covered in how model families differ), the next question is which one to actually reach for on a given task. Most people answer this backwards: they pick a favorite vendor, then try to make every task fit that model. A better approach works through four questions in order, each one narrowing the field before the next.
This lesson gives you the framework and the checkpoints to practice it. For the deeper practical guide, with specific benchmarks, real-world tradeoffs, and vendor-by-vendor recommendations, see which AI model should I use. Read that when you want the long answer. This lesson is for building the instinct so you don't need to look it up every time.
Question 1: what does the input look like?
Before anything else, check what you're actually feeding the model. Plain text opens up every vendor. The moment your input includes video or audio, your options narrow fast.
Right now, Gemini 3.8 Flash is the only current fast-tier model across the three major vendors that natively takes video, audio, image, and PDF input alongside text in a single call. If your task is "summarize this video call" or "transcribe and analyze this audio clip," that's not a preference, it's a hard constraint that picks the vendor for you before you've even thought about task type or budget.
If your input is text, code, or standard images, all three vendors are in play and you move to question 2.
Question 2: what kind of task is it?
With input type settled, look at what the model needs to do. A few rough categories:
- Extraction and classification: pulling structured data out of unstructured text, sorting items into categories. Low reasoning demand, high volume tolerance.
- Writing and drafting: emails, content, documentation. Needs good instruction-following and tone control more than deep reasoning.
- Coding: ranges from autocomplete-style suggestions to multi-file refactors. The upper end of this category benefits heavily from models tuned for agentic work, like Claude Opus 5.5.
- Analysis and research: synthesizing multiple sources, working through ambiguous requirements, multi-step reasoning. This is where flagship-tier models earn their price.
- Chat and conversation: general assistant behavior where latency matters as much as depth.
Task type tells you which tier you're likely to need before you've looked at a single price tag. Extraction and simple writing usually fit the fast tier. Coding and analysis usually need balanced or above.
Question 3: what's your budget and volume?
Now bring in cost. This isn't just "which model is cheapest," it's "what does this cost at the volume I'm actually running." A model that's twice as expensive per call but handles the task correctly on the first try is often cheaper in practice than a cheap model that needs three retries and a fallback.
Rough math worth doing before you commit: estimate calls per day, multiply by average input and output tokens, multiply by the per-token rate. Do this for the fast, balanced, and flagship tiers of your chosen vendor and look at the delta. If the flagship tier costs 10x more but doesn't measurably improve your task's success rate, that's a signal to drop down a tier, not a reason to accept the bill. The pricing reference has current rates if you want to run these numbers for real.
Question 4: are there specific capability requirements?
Last, check for hard requirements that override everything above. These come up less often but rule out models fast when they apply:
- Guaranteed structured output: some models offer strict JSON schema enforcement, others only offer best-effort formatting. If your pipeline breaks on malformed JSON, this matters more than raw capability.
- Context window size: processing a 500-page document or an entire codebase needs a model with enough room. GPT-6 Astra and Gemini 3.8 Flash both currently offer roughly 1.05M token windows; Claude Haiku 4.5 tops out at 200K.
- Reasoning controls: if you need fine-grained control over how much the model "thinks" before answering (covered in the next lesson), check what each vendor's effort parameter actually supports at that tier.
- Latency guarantees: real-time or interactive features have hard response-time budgets that rule out slower flagship-tier reasoning regardless of quality.
If none of these apply, you likely already had your answer after question 3.
Working through an example
Say you're building a feature that reads customer support tickets (text only), classifies urgency, and drafts a suggested reply.
Input type: plain text, so all three vendors qualify. Task type: this splits into two sub-tasks, classification (fast-tier territory) and drafting (balanced-tier territory). You could run both on a balanced model for simplicity, or split them: fast tier for classification since it runs on every ticket, balanced tier for drafting since it only runs on tickets that need a reply. Budget: if you're processing thousands of tickets a day, the split approach saves real money, since most tickets never reach the draft step. Capability requirements: check if you need structured output for the classification step (you probably do, since something downstream needs to read "urgent" vs. "normal" reliably).
That's the framework in miniature: input narrows the vendor, task narrows the tier, budget confirms the tier or pushes you to split the workload, and capability requirements catch anything that would otherwise force a change.
Checkpoint
Before moving on, try applying this to a task you actually have in mind: what's the input type, what task category does it fall into, roughly what would it cost at your expected volume on the fast vs. balanced tier, and does it have any hard capability requirements? If you can answer all four without hesitating, you've got the framework down.
Next: Reasoning models covers the fourth question in more depth, specifically how "thinking" settings work and how they change the way you should prompt a model.