
Which Claude model should you use? transcript
Claude · @claude
Words
580
Runtime
3:45
Speaking pace
155wpm
Reading time
2min
155 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Picking the cheapest Claude model at the lowest effort looks like the obvious way to save money or usage on your plan. But with the launch of Claude Fable 5.1, a more intelligent model can cost less to complete a task. Three things affect what a task costs: model, effort, and cache. In general, the more capable the model, the more it costs per token. But just like a college student can solve a maths problem in fewer
78 words, the words spoken in the first 30 seconds at 155 words per minute.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 36 |
| Average words per sentence | 16.1 |
| Longest sentence | 30 words |
| Questions asked | 1 |
| Sentences containing a number | 8 |
Most used terms
- model17
- effort16
- task10
- work10
- cost6
- fable6
- costs5
- routine5
- tasks5
- token5
- claude4
- ended4
Filler phrases
6 in total: like 4 · kind of 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, published by the channel, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Transcript
Picking the cheapest Claude model at the lowest effort looks like the obvious way to save money or usage on your plan. But with the launch of Claude Fable 5.1, a more intelligent model can cost less to complete a task. Three things affect what a task costs: model, effort, and cache. In general, the more capable the model, the more it costs per token. But just like a college student can solve a maths problem in fewer steps than a fifth-grader, a more intelligent model often finishes a task with less work.
Effort is how much reasoning the model does before it answers. More effort costs more, but often gets a better result. Lastly, caching. The model doesn't remember anything between requests, so every time you send a message, it has to reprocess the whole conversation. Caching lets it reuse what it already processed at about a tenth of the cost. So the price per token is only a part of what a task costs. The model and the effort decide how many tokens a task takes, and caching helps reduce the costs.
Let's take the Claude Fable 5.1 and 5 models as an example. 5.1 has the same price per token as 5. But where usage is billed by token, its cache reads cost 75% less. Overall, that reduces the cost by around 25% for typical work and up to around 45% for long agentic tasks. On CursorBench, a third-party coding benchmark, 5.1 at medium effort produced a similar result to 5 at max effort for about a fifth of the cost. With that in mind, the model and effort you choose depend on the kind of work that you do.
Broadly speaking, there are two kinds of tasks: unconstrained, open-ended work like deep research, complex analysis, and long-running agent work. And constrained routine work like summarizing a document, drafting an email, and pulling data from a report. So what should you use? If your work involves more open-ended tasks, use Fable 5.1 and adjust the effort, starting with medium and raising it as needed. If your work involves more routine tasks, stay on Opus or Sonnet and reach for Fable as needed.
On routine tasks, a cheaper model can get a similar result, so you don't pay for the capability the task never uses. Similarly, subagents are constrained to a specific task, so a cheaper model is often enough. Now, if you're an admin managing Claude for your organization, you can configure the available models and effort levels for everyone or per role. There are three controls you can reach for. Model entitlements decide which models users can pick from.
Effort caps set the highest effort users can select for a model. Defaults set which model and effort a new conversation starts on. For a role where long open-ended work is most of the job, you can make Fable 5.1 the default and consider setting an effort cap. For roles where work is mostly routine, keep Opus or Sonnet as the default model and entitle Fable, so users can switch to it when a task call for it. Alright, so when you're deciding what model and effort to use, consider the cost per task, not the price per token.
Choose the model based on the kind of task, whether it's long and open-ended or routine. And for more capable models, start with a lower effort and increase if needed. To learn more, read the resources linked below.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Use this transcript
Three free tools that work on the material around a video like this one. No signup, no login.
Hook Analyzer
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Policy Pre-Flight
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Channel Skill Generator
Read this channel's public videos and transcripts, and download a writing brief for it.