Open any provider's pricing page and you'll see three or four models with different names but a lot of shared DNA. GPT-6 Astra, Sol, and Luna. Claude Fable 5.1, Opus 5.5, Opus 5, Sonnet 5, and Haiku 4.5. Gemini 3.8 Flash and 3.5 Flash-Lite. These aren't unrelated products, they're tiers of the same family, built on the same underlying architecture and training approach, then scaled up or down for different jobs.
Understanding what a family is (and isn't) saves you from the most common mistake people make when picking a model: assuming "newest and biggest" is always the right call.
What makes a family a family
A model family shares a training recipe, a data pipeline, and usually an architecture generation. What changes across tiers is scale: parameter count, training compute, sometimes context window, which trades off against cost and latency.
Take OpenAI's current lineup. GPT-6 Astra is the flagship: $10 per million input tokens, $50 per million output tokens, a 1.05M token context window. GPT-6 Sol is the balanced tier at $2/$10. GPT-6 Luna is the fast tier at $0.10/$0.50. Same family, same general capabilities profile, but Astra will out-reason Sol on a genuinely hard problem while Luna will answer a simple classification question in a fraction of the time and cost.
Anthropic's family is wider. Claude Fable 5.1 sits at the top ($10/$50, 1M context) as the most capable model overall. Claude Opus 5.5 ($4/$20) is tuned specifically for coding and agentic work. It's not "weaker" than Fable, it's specialized. Claude Opus 5 ($5/$25) is the previous flagship generation, still in service. Claude Sonnet 5 ($2/$10) is the balanced workhorse most people reach for by default. Claude Haiku 4.5 ($1/$5, 200K context) is the fast tier, with a smaller context window but built for throughput.
Google's Gemini family right now is narrower: Gemini 3.8 Flash ($0.75/$3.75 as a promotional rate through December 2026, 1.05M context, and the only one of the three vendors' current fast tiers that natively takes video, audio, image, and PDF input alongside text) and Gemini 3.5 Flash-Lite ($0.30/$2.50) below it.
Flagship, balanced, and fast: what actually changes
Three tiers show up across every vendor, even though the names differ:
Flagship (Astra, Fable 5.1) is the model you reach for when a task genuinely needs maximum reasoning depth: ambiguous requirements, multi-step analysis, code review where subtlety matters. You pay 5 to 10 times the balanced tier's rate for this.
Balanced (Sol, Sonnet 5) is the default for most production work. Strong enough for the majority of coding, writing, and analysis tasks, priced to run at volume without a second thought.
Fast (Luna, Haiku 4.5, Flash-Lite) is built for high-throughput, low-latency work: classification, extraction, routing decisions, anything where you're making thousands of calls and the per-call task is simple. The capability gap between fast and flagship tiers is real, but it matters less than you'd think for well-scoped tasks.
Notice that Anthropic breaks this three-tier pattern slightly with Opus 5.5, a model that's priced below Fable 5.1 but isn't really a "weaker flagship." It's a task-specialized model tuned for coding and long agentic loops. Not every family maps cleanly to flagship, balanced, and fast tiers, so check what a tier is actually optimized for, not just where it sits in the price list.
Why bigger isn't automatically better
The instinct to always reach for the flagship model is understandable (it tests best on every benchmark), but it's usually the wrong default for three reasons.
Cost compounds. A $10/$50 model running at the volume a $2/$10 model would handle isn't a rounding error, it's a 5x line item. If you're building something that makes hundreds of calls a day, that difference shows up on the invoice fast. The pricing reference has the full current rate card if you want exact numbers before committing to an architecture.
Latency compounds too. Flagship models think longer and generate more thoroughly, which is exactly what you want for a hard problem and exactly what you don't want for an autocomplete-style feature where users expect a response in under a second.
And past a certain point, extra capability doesn't change the output. A fast-tier model correctly extracting three fields from a structured document doesn't get more correct by upgrading to the flagship, because it was already right. You're paying for headroom you're not using.
The practical rule: match the tier to the hardest part of the task, not to the task's importance. A high-stakes but simple task, like routing customer emails to the right queue, doesn't need a flagship model. A low-stakes but genuinely hard task, like debugging a subtle race condition, might.
Where to check current specs
Model names, prices, and context windows change often enough that hardcoding them into your head is a bad idea. The models registry tracks current models across all three vendors, with dedicated pages for Claude, Gemini, and OpenAI's GPT line, plus a side-by-side comparison if you're deciding between vendors rather than just tiers.
Next: Choosing a model walks through the decision framework for picking a specific model once you know what your task actually needs.