Which Claude model should you use, and when the answer matters
Four models, a confusing pricing page and a decision most people on a subscription never actually make: what the tiers are for, what they cost relative to each other, and when the choice matters.
Somebody in your organisation has already spent an afternoon deciding which Claude model to use, and it changed nothing about their week. For most of the people I work with, the answer is that you should not have to think about it, because the app has already chosen for you. There are about three situations in which the choice changes what you get back. The rest of the time it is decoration, and telling the two apart takes two minutes.
The four you are likely to meet
As things stand, four models are generally available to anyone. A fifth exists on invitation only, and older versions of each tier stay listed on the pricing page long after the newer ones arrive. The wording in the middle column below is Anthropic's own positioning rather than my opinion of it.
| Model | Best for | Context window |
|---|---|---|
| Claude Fable 5 | Highest capability, and long-running agent work | 1M tokens |
| Claude Opus 5 | Complex agentic coding and enterprise work | 1M tokens |
| Claude Sonnet 5 | The best balance of speed and intelligence | 1M tokens |
| Claude Haiku 4.5 | Fastest responses, near-frontier intelligence | 200k tokens |
If the names in front of you do not match the names here, the official model overview is the list Anthropic keeps current.
The practical differences between the top three are smaller than the naming suggests. All four take text and images, all four work across languages, and all four will draft a decent set of meeting notes. What you buy as you move up the range is depth on hard problems and stamina on long ones, not a different sort of tool.
The cost question, in relative terms
Printing prices would be a disservice, because they change and articles do not. The shape is stable enough to be useful: the top tier costs about ten times as much as the entry tier for the same quantity of text, in both directions, on what you send and on what comes back. The tiers in between sit closer to the bottom of that range than the top.
Two facts matter more than the ratio. Per-token pricing applies to use through the API, meaning software your organisation builds or buys. If you are working in the Claude app on a subscription, you are not being billed per unit of text and will probably never see a token price in your life. And most subscribers never pick a model at all, because the app selects a sensible default and is right about it for ordinary work.
When the choice genuinely matters
At volume through the API, the gap between tiers becomes a budget line, and choosing well is an engineering decision, not a preference. If you want an answer faster and can accept less depth, stepping down a tier is a fair trade. If a piece of work keeps coming back shallow, stepping up sometimes helps.
Sometimes. In the sessions I run, the prompt is at fault far more often than the model, and changing model is a comfortable way of avoiding that conclusion. It costs nothing to try and it feels like progress, which is exactly why people keep doing it instead of rewriting the request.
If the naming itself is what bothers you, the tier and the generation number are doing two different jobs. If your real question is how much material you can put in front of it at once, read what a context window holds next.
Where your attention is better spent
Model choice is close to the least important decision you will make about Claude. An afternoon spent comparing tiers buys you less than ten minutes spent on how you ask, and the second of those is the one nobody puts in the diary.
The table above will be out of date before your prompting habits are.