Every few months someone asks me "which AI model should I use" like there's one right answer sitting in a spreadsheet somewhere. There isn't. There are three or four vendors, each with a fast/cheap tier, a balanced tier, and a flagship tier, plus a couple of specialized options that only make sense for specific inputs. The names change constantly: by the time you read this, some of the models I mention will have successors. What doesn't change is the set of questions that actually determine the right pick.
This is that framework. I'm deliberately not leading with a spec table, the site keeps one current at /models, with a full interactive comparison at /models/model-comparison. What I want to give you here is a way to reason it out yourself, so it still works in a year when the names on that table are different.
Start with what you're feeding the model
Before task type or budget, input type narrows the field hardest, because it's the one axis where models are flatly incapable rather than just worse.
Text only. Every major model family handles this: OpenAI's GPT-6 line, Anthropic's Claude line, Google's Gemini line. This is the unconstrained case; move to the next section.
Images. Also broadly supported across current-generation models from all three vendors. Vision quality varies more by task (dense document OCR vs. general scene description vs. chart reading) than by vendor, so if image understanding is central to your product, test your actual images rather than trusting a benchmark.
Native video or audio input. This is where the field narrows sharply. As of this check, Google's Gemini family is the only major lineup with native video and audio input built into the same models you'd use for text: you feed in a video file directly, not a transcript or a sampled frame sequence. If your workload is genuinely video-native, analyzing footage, understanding a screen recording, processing a call recording without a separate transcription step: that capability gap, not price or reasoning quality, is what should decide your vendor. Check /models for which specific models currently support this, since it changes as vendors ship new versions.
Huge context. If you need to hand the model an entire codebase, a stack of long documents, or a full book at once, context window size becomes the filter. Current flagship and mid-tier models across vendors mostly cluster around the 1 million token mark, which is enormous, but max output tends to be much smaller (often 128K or less), so "huge input, huge output" isn't a given even when the context window is. If you need both a large input and a large generated output in one call, verify both numbers, not just the headline context window.
Then narrow by task type
Once input type has filtered your options, task type tells you which tier within that filtered set to reach for.
Chat and conversational assistants. Balanced mid-tier models are usually the right default, fast enough to feel responsive, capable enough to handle most conversational turns without heavy reasoning. Save the flagship tier for conversations that need genuine multi-step reasoning mid-chat.
Classification, extraction, and formatting. This is fast-tier territory almost without exception. These are narrow, well-specified tasks (pull a field out of a document, sort a ticket into a category, reformat JSON) and a flagship model's extra reasoning capacity is wasted spend here. If you're paying flagship prices for a task like this, it's very likely the single easiest cost cut available to you; see AI model pricing in 2026 for the cost math on this.
Coding. This is where the flagship-vs-mid-tier line gets genuinely blurry, and it's worth testing rather than assuming. Some vendors now ship a coding-optimized tier that sits below their absolute flagship on price but matches or beats it on coding benchmarks specifically: Anthropic's Opus 5.5, for instance, launched positioned exactly there relative to Opus 5. Don't assume "most expensive = best for code." Check current model pages for which tier a vendor recommends for coding and agentic work specifically, and validate against your own repo if the choice matters.
Agentic workflows (multi-step tool use, autonomous task completion). These tasks compound errors across steps, so this is genuinely flagship territory, the cost of a wrong intermediate decision (a bad tool call, a misread instruction) is much higher than the token cost difference between tiers. This is also where reasoning effort settings matter most: set effort too low on an agent and it makes shallow, greedy decisions that cost more to unwind than the tokens it saved. See prompting reasoning models for how effort settings interact with agentic task quality, and the prompt library for agent-oriented prompt patterns to start from.
Classification-at-scale or LLM-as-judge. If you're running a model to grade or route thousands of other model outputs, fast-tier is almost always the right call unless the judgments are genuinely subtle, and even then, a mid-tier model with a well-specified rubric usually beats a flagship model with a vague one. promptfoo for LLM testing covers building the eval harness this kind of judging setup needs.
Then apply your budget and volume constraint
Capability and task type get you to a shortlist. Budget and volume decide which model on that shortlist you actually ship with.
Low volume, latency doesn't matter much. Use the most capable model that fits your task type, at low volume the cost difference between tiers is often a rounding error, and you're better off not thinking about it further.
High volume, cost-sensitive. This is where the fast tier earns its keep, and where the levers in AI model pricing in 2026 (caching, batching, right-sized reasoning effort) matter more than which vendor you picked. A well-cached, well-batched request to a mid-tier model frequently beats a naive, uncached request to a fast-tier model on total cost.
Mixed workload: some requests need the flagship, most don't. Don't pick one model for the whole product. Route by request: cheap/fast model for the easy majority, capable model for the hard minority. This is the single highest-leverage cost decision most teams haven't made yet. The full framework for building this is in LLM routing: how to choose the right model for each task.
Latency-critical, user is watching a cursor blink. Fast-tier models return short outputs in well under a second; flagship models with reasoning turned up can take several seconds on hard problems. If your UX can't absorb that wait, that's a hard constraint that overrides most other considerations: test perceived latency with real users, not just raw token throughput.
Then check for a specific capability requirement
A few requirements don't fit neatly into input type, task type, or budget, but they're common enough to call out on their own.
Guaranteed structured output. If you need the model to reliably return valid JSON matching a schema (not "usually," but guaranteed) check whether your chosen model supports a strict structured-output mode rather than relying on prompt instructions alone. Every major vendor now offers some form of schema-constrained generation; the exact mechanism (strict tool schemas, a dedicated structured-output config) differs, so check the current docs for your vendor rather than assuming the pattern from another provider carries over.
Extended reasoning on genuinely hard problems. If your task is math, complex logic, or multi-step planning where getting it right matters more than getting it fast, look for a model with a high-effort reasoning mode and actually turn the effort setting up: the default is often tuned for a broad workload, not your hardest case.
Regulatory or data-residency constraints. If you're required to keep inference in a specific cloud or region, your vendor choice may be constrained before any of the above criteria apply. Check what's available on the cloud platform you're already committed to before evaluating models on merit alone.
Putting it together
The order that actually works, in practice: filter by input type first (it's a hard capability gate, not a preference), narrow by task type second (it tells you which tier to reach for), apply budget and volume third (it tells you whether to route rather than pick one model), then check for anything capability-specific you might have missed. Most "which model should I use" questions collapse to a clear answer once you've walked through those four steps in order: the hard part isn't the decision, it's remembering that the decision has more than one axis.
For the current model roster this framework applies to, see /models; for per-vendor prompting specifics once you've picked one, the guides for Claude, Gemini, and OpenAI's GPT line are a good next stop. And if cost is the axis you're least confident reasoning about, AI model pricing in 2026 is the companion piece to this one, it covers the reasoning-token and caching mechanics that make "cheap" and "expensive" models cost more or less than their sticker price suggests.



