Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
10:364.6x the video's typical replay level
this with Opus 48, GLM 52, and with Sonic 5 all at the same time. And I have a lot of thoughts on the results here. Opus 48 in around 26 minutes, okay, but 27 depending on how you round it. And this version of the game was pretty dang decent. Past some rough UI quirks here.
Said at 10:28
Most replayed moment #2
2:174.5x the video's typical replay level
out. Set up a scheduled dev in to check for regressions on my website every day. It's kind of crazy. It just asks you when you want it to go, you hit a button, and now you're done. If you're pushing agents to their limits, you're already using dev in, and if you're not, fix it at soidev.link/devin. I'll blast through
Said at 2:10
Most replayed moment #3
11:273.7x the video's typical replay level
found this like fun enough to play that I just stopped distracting myself so I could come up here and film the video. So what about the other versions? Next I'll show you guys the GLM 52 version, which was interesting. I don't have costs for the other runs because it's not the easiest thing to calculate. They
Said at 11:21
The graph counts replays. It does not show where viewers stopped watching.
Words
6,235
Runtime
28:37
Speaking pace
218wpm
Reading time
26min
218 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Seems like a pretty awesome time to be a Claude fan, cuz we just got two huge pieces of news. The first is a new model, Sonnet 5. It's finally here, and there's a lot to talk about with it. This model's definitely not what you think, and I haven't seen any reporting really covering its strengths and weaknesses properly, cuz it has plenty of both, believe me. But, in more important news, Fable 5 has just been unbanned by the Secretary of Commerce. The restrictions have been lifted. The model's not back as of the time I'm recording this, but there's a very good chance it'll be out
109 words, the words spoken in the first 30 seconds at 218 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 375 |
| Average words per sentence | 16.6 |
| Longest sentence | 83 words |
| Questions asked | 11 |
| Sentences containing a number | 102 |
Most used terms
Filler phrases
55 in total: like 32 · actually 12 · kind of 7 · uh 2 · I mean 1 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Seems like a pretty awesome time to be a Claude fan, cuz we just got two huge pieces of news. The first is a new model, Sonnet 5. It's finally here, and there's a lot to talk about with it. This model's definitely not what you think, and I haven't seen any reporting really covering its strengths and weaknesses properly, cuz it has plenty of both, believe me. But, in more important news, Fable 5 has just been unbanned by the Secretary of Commerce.
The restrictions have been lifted. The model's not back as of the time I'm recording this, but there's a very good chance it'll be out by the time you're watching this video. So, go double-check. Good chance you have it. I have a little more to say about that, but I really want to focus in on Sonnet 5, because as I said, it's a very interesting model. I'm using it all day, testing it on various tasks and benchmarks. And with models getting so expensive, my individual benchmark runs can cost way over $300.
That number is only counting runs that didn't fail. Failures still cost me money, believe it or not. As such, I need to take a quick break for today's sponsor. You're probably not being bold enough with agents. I'm saying this because it was the case for me. I remember back in the day when Devin was first announced, and they claimed that AI would be able to be a full engineer, as though it's on your team. And that made no sense to me at the time.
That's why when they hit me up, I had to go dig into the product more, and I've been blown away. If you haven't kept up with what they're working on, they made it possible to spin up your codebase in the cloud, one of the best setups for that, by the way. And once you have it set up, you can run a real Linux box that your agents will control as they do work. This means you can develop anything from a web app to a real desktop application, with agents able to test things, do things, run things, and you can even interact with it.
All of that's cool by itself, but what I did here is even cooler. I told Devin to go through all the pages on my website, specifically to check for mobile responsiveness, UI bugs, and client-side errors. Obviously, you could run this as a single agent run, but it's going to take forever, and I find that when you do this on too many things at once, the result is often that it's not going as deep as I want. So, here it went across all eight pages on my site, and spun up sub-agents for each of them, and checked for all the different potential regressions for the way the site behaves.
It passed all the results up to that top-level agent, so you could see them and make decisions, but you can also go into any of those sub-agents, and it even spits out videos of what it found when it explored. This caught some novel regressions that I didn't notice and ended up going to patch right after. But, I want to prevent this from happening again, and this is why scheduled dev ins is so cool. Check this out. Set up a scheduled dev in to check for regressions on my website every day.
It's kind of crazy. It just asks you when you want it to go, you hit a button, and now you're done. If you're pushing agents to their limits, you're already using dev in, and if you're not, fix it at soidev.link/devin. I'll blast through the Fable news fast so we can focus on what I'm here to talk about, which of course is Sonnet 5. If you haven't caught up, Fable got banned on June 12th, just 3 days after it originally came out because of concerns related to its ability to hack and do things that we wouldn't want it to do.
In particular, the government was concerned about jailbreak capabilities that would allow the model to find security issues in software, and they wanted to restrict that from foreign actors and foreign nationals. Anthropic just took that as a hard ban, but all of those export controls have been lifted as of now. Since the issuance of my previous letter dated June 12th and June 26th, Anthropic has taken steps in close coordination with the US government to address the risks associated with Claude Mythos 5 and Fable 5.
Among other things, Anthropic has agreed to proactively detect and address security risks associated with the models, to work diligently with the US government on protocols and standards and releases for Mythos, Fable, and future models, as well as to inform the US government of any malicious activity. In light of these actions and commitments, as well as the Bureau of Industry and Security's evaluation of the diversion risks now presented by Claude Mythos 5 and Fable 5, the controls in the June 12th letter are withdrawn.
A license is no longer required for the export, re-export, or in-country transfer, including deemed export or deemed re-export of the Mythos or Fable models. Commerce reserves the right to re-evaluate the decision, yada yada yada. Most important piece here is export or re-export. What this means is if you're building a service like, I don't know, T3 chat where you're hosting these models over API and then letting users hit them themselves, this means that that is allowed as well, which is really important.
I had a lot of concerns that might not be allowed with whatever conclusions they came to, so knowing that that's good makes me happy. Especially because uh I don't want to use Sonnet 5 too much. There are some good parts here, but we will talk about them. Let's start from what Anthropic has to say. Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that a few just few months ago required larger and more expensive models.
We'll talk about the more expensive a lot as we go forward, but I also am excited to talk more about the agentic side because there are some things Sonnet does that no model available as of the time of recording does do. Fable is not out right now, which is why, but uh yeah. The number 5 is the most notable thing here, not the word Sonnet or the fact that it is a new Sonnet. And everybody's saying this should have been Sonnet 4 8.
I don't agree. Let's dig in. Anthropic even is aware of the fact that Opus is kind of the standard now, as they say here. More recently, the clearest gains in agentic capabilities have been in the Opus tier class models. I agree. It honestly feels like Anthropic bumped up the tier for everything where what we used to use Haiku for we now use Sonnet, what we used to use Sonnet for we now use Opus, and what we used to use Opus for we now use Fable.
Very clever way to get us to spend more money, but Sonnet is important, so let's see how it improved. Its performance is close to that of Opus 4 8, but at lower prices. Lower prices, remember that. It's a substantial improvement over its predecessor 4 6 on important aspects of agentic performance like reasoning, tool use, coding, and knowledge work. We have some benches here. We have SWE Bench Pro, which remember has been almost entirely compromised.
So, the levels of contamination in SWE Bench are insane. I don't think this measures much at this point, but cool, it's in there. There are other benches, and there's even more in the system card that we can talk about. But terminal bench was a meaningful improvement from this from the high 60% to the 80% range. Human eval exam was a meaningful bump as well. Computer use was a slight bump, too. I will say from my experience, Anthropic went from leading in computer use to lagging quite a bit behind even Google in a lot of cases and knowledge work where they have improved meaningfully actually scoring slightly higher than Opus 48 somehow.
They have some call outs about safety stuff, but I will say outright Sonnet 5 is too dumb to be a particular safety risk. I would not worry. GLM 52 is a more meaningful security risk than Sonnet 5. The charts below compare the performance of Sonnet 5 with 46 and Opus 48 at different effort levels on the Agentic search evaluation brows comp and the computer use evaluation OS world verified. The orange line is the one we're talking about here, Sonnet 5.
As we see here, Sonnet 5 does slightly outperform Opus 48, although not much cheaper. Of course, the medium and low runs are meaningfully cheaper at less than $2 and less than $5 per task. Once we're in that 5 to 10 range, it's slightly more performance at a slightly higher cost. And then that trend continues. What's much more interesting to me though is the Agentic computer use bench that I'm honestly confused why they included because as you see here, Sonnet 5 does not perform as well as Opus both on performance and on cost.
Is it better than Sonnet 46? Yes, but you can get better performance for the same price range on Opus medium or high as you would get from any of the Sonnet flavors. In fact, Sonnet 5 max ends up being more expensive than Opus on high. There is more bad news though. Pretty much every single version of Sonnet 5 is more expensive and worse performing than a GPT 5.5 equivalent according to Cursor Bench. Here we see 5.5 medium scoring almost as high as Sonnet 5 max and also Sonnet 5 high and max costing more than the heaviest runs on GPT 5.5.
It almost feels like Sonnet 5 was released to advertise how good of a value 5.5 is on different reasoning levels because a lot of these benches show that absurdity. Reminder, I do want to talk about the things I like about the model and the things I saw when I was using it, which we'll get to in a little bit, but we need to talk a bit more about benches first. According to the artificial analysis intelligence index, it is in fourth place across all models, falling just slightly behind Fable 5 Opus and GPT-5.5.
But, the intelligence index is far from the most interesting part of artificial analysis. Personally, I find the cost section to be way more interesting. And if we look at cost per task, you'll see something a bit terrifying. 5.5 on X High ends up being cheaper in real-world work than Sonnet by more than 2x. In fact, Opus 4.8 is also cheaper than Sonnet 5. But, ready for the craziest part here? If we cover the cost to run the entire benchmark, this is the total cost.
I'm making an assumption here. I haven't talked to the artificial analysis guys about this, but I'm guessing that the cost per task is averaged with extremes filtered out. And the total cost does not have those extremes filtered out because Sonnet 5 is the most expensive model they've ever run through the bench at $6,000, topping even Fable 5 at $5,600. And just for fun, I'm going to throw in GPT-5.5 medium and low in here because they cost a sixth and a twelfth as much as Sonnet 5 does.
And if we go back to the top here, there is a gap in the intelligence, but it's not quite as big as you would expect for a 10x decrease in total cost. But wait, Theo, didn't Anthropic say it's cheaper? They did. And the way that they said it was cheaper is in the price per token. They're debuting it at an introductory price of $2 per million input tokens and $10 per million out until August 31st, at which point they're going to bump it back to the usual for Sonnet, which is three per mil in and 15 per mil out.
They also got rid of the weird Sonnet specific limit that existed on the subscription plans. I don't understand why that stuck around so long, but it is finally gone. It just counts towards your normal usage. I have seen people racking up meaningful amounts of usage in Claude Code with this, like taking an actually surprising chunk out of their percentage even on the $200 plan because it's not a very efficient model.
It is funny to be filming this right after I published my video about how opening eye models are so efficient because this model is the opposite. It is even less efficient than 54 mini, which I honestly think was a bit of a show for this type of agentic or long-term thinking work. It used almost two times as many tokens as Opus and more than two times as many close to like five times as many as GPT 55 did on X I and 55 on medium did only 5K tokens.
I did 69k. You could do the math. It's not an efficient model at all. This also means it's slow as balls for real-world work, which I experienced myself trying to recreate my fish web game. I pointed at the original repo and told it to rebuild it from scratch. You can use some of the assets. You can reference the code, but I want a new game built from the start. I tried this with three models. I tried this with Opus 48, GLM 52, and with Sonic 5 all at the same time.
And I have a lot of thoughts on the results here. Opus 48 in around 26 minutes, okay, but 27 depending on how you round it. And this version of the game was pretty dang decent. Past some rough UI quirks here. It doesn't do padding and things right, but once you're in the game, the controls are solid. The core mechanics work. And the most surprising thing to me is it did a really good job balancing the economy of the game.
Like it felt good to play for a while and like rank up, get more money and like actually play the game. It made a couple creative decisions that I don't necessarily love. Like it made my pet here transparent for some reason. It also added these weird light beams that I don't love. It's not a perfect version, but this is absolutely workable as a thing that you could keep iterating on. And it has a pretty solid gameplay loop.
Like I was surprised that I actually found this like fun enough to play that I just stopped distracting myself so I could come up here and film the video. So what about the other versions? Next I'll show you guys the GLM 52 version, which was interesting. I don't have costs for the other runs because it's not the easiest thing to calculate. They don't give it to you, but I used open code for 5-2, so I do have a cost.
It cost $8.30 to create the following port. This one's interesting because it changed the UI more, but I didn't necessarily like the things it changed. It also has super choppy movement, a really weirdly rendered background. It's economy is garbage, so it just doesn't feel good to play. There's a lot more time spent sitting and waiting and nothing really happening. And most egregiously, when you click for the gun to shoot in a direction, it almost feels like it picks a random direction to shoot in.
Like I'm clicking on the left and it's shooting down. There's also a problem with GLM for these types of things in that the GLM models have no vision, so they can't do browser use and look at what the browser is showing because they have no I they have no ability to see it. So, that took like 35-40 minutes and wasn't particularly impressive. Here is Sonnet's version. Play. Sure. You might be able to see it's a bit of a mess.
They got rid of all the hot keys for buying. There are no fish in the tank by default, just this very poorly rendered pet. I can click the button to buy things, but when I do that, it also shoots. And the economy is garbage. It just doesn't actually feel good to play. At least the bullets go in the right direction, even if they go when they shouldn't, like when I'm clicking other things in the UI. And it was smart enough to make it so clicking here doesn't trigger when it's in the dead state.
Like here, I can't afford it, so I can't buy things. But when I can buy things, it shoots. It's just like the type of silly bug I expected older models to do. I was hoping a better model would not make that type of mistake. Also worth noting that this run took 2 hours to complete. It's actually a bit more than I didn't think to time it, but I know roughly when I started, I know roughly when I finished, and it was at least 2 hours.
Might be closer to 2 and 1/2. Part of the reason it took so long is that it spun up a ton of sub-agents throughout its building. I didn't ask it to. The prompt was really simple, but it chose to spin up an agent to go look into the old codebase, then spin up another sub agent to write a plan, and spin up a few to analyze the plan, and then spin up a few to implement the plan, and then they made a to-do list, and they went through the to-do list one at a time, and then at the end tried using it quick, and then told me it was done.
One other thing I found really interesting with Sonnet is that it asked way more questions than Opus. Opus asked no questions. Sonnet asked a handful, and they were pretty good. They were trying to help me scope the project before it got started. As I mentioned before, it spun up agents to do all of the investigations, which I also thought was interesting because Opus didn't spin up any sub agents at any point during its run with the exact same prompt.
Sonnet decided to do that. And this is where that thing I hinted at earlier comes in. The thing I was hinting at is the number five. The reason this model's interesting to me is not because it's a really good value, or benchmarks really well, or anything like that. The reason this model's interesting to me is because it has behaviors that I've only seen before in Fable 5. It likes sub agents, and it knows how to orchestrate them.
It does a good job of breaking up work into smaller pieces, and then handing that off and staying on task. That is the thing that made Fable 5 so different, that made it so exciting to me, is that it could break up the work to do bigger, heavier things. The problem with Sonnet is that it's not a smart enough model to do that well, and it often ends up running in circles, and taking forever as a result. It often will even break work up that shouldn't be broken up, and that results in much slower times, and also much higher costs when you run it.
I did run it on a few other things. I had it help me with some bugs in Skate Bench, and it took way too long to solve them, so I just gave up and went back to using 5.5 medium on fast. It was so slow that it actually was hitting timeouts internally on Bun's fetch implementation. So, no matter what I did, I kept getting meaningful timeouts on the max version. And god damn, the max version is a bit of a token hog. This benchmark is a silly one.
I measure how well models can name skate tricks given a description. It did get contaminated, so I've since doubled the number of questions in a private repo that is not exposed to the world at all in order to try and get better measurements. And what you can see here is that Gemini 31 Pro is the only model even in the 90s nowadays, scoring a 95% where everything else is in the 80s at best and in the like 30s at worst.
Sonnet 5 on X High is the worst score of any model I'm currently testing at 37%. It also wasn't cheap. It ended up being more expensive than the 31 Pro per test at 4 cents per run versus 2.2 cents per run. But that's on X High. Max did score meaningfully better at 59%, but it also had an interesting quirk. And by interesting, I mean expensive. Sonnet 5 Max is by far the most expensive model I've ever run Skate Bench on, even counting Pro models.
It was 15 cents per question average, and there were some questions that cost as much as a dollar. For it to get them wrong, by the way. And this is the worst part. Sonnet 5, when it can't figure something out, loves to run in circles. And the result is it went from a 1,600 tokens average to 6,000 tokens average when bumped from X High to Max, which kind of just gives it permission to go until it resolves the thing. I did see a lot of people being confused about how a Sonnet model could possibly be more expensive than an Opus model.
So, I'm going to do my best to give an analogy here. Imagine you're the CEO of a company with two engineers. You have a really experienced engineer who's super senior, really smart, knows the codebase great, but he's super expensive. He costs you, I don't know, $100 an hour. You also have a more junior engineer that's not quite as good and capable, but they can get real work done, and they only cost $20 an hour. You have a task that you think will be a little hard, but not too bad, and you can give it to the senior engineer knowing he will solve it, but it will take 10 hours. 10 * 100 it's not cheap.
That's a $1,000 task now. Or you could give that same task to the junior engineer who might be able to finish it just as fast, but they also might take a lot longer. And at $20 an hour, that sounds really good until they take 100 hours to solve it. As I was saying, if the $100 an hour employee takes 10 hours to solve the task, costs a grand. The $20 an hour employee only takes 10 hours, then it costs a lot less money.
It's only 200 bucks. What if that $20 an hour employee takes a bit longer? Let's say they take 100 hours and there isn't a senior engineer popping in to check in. Then that task cost $2,000. And now we have a more important question. For that junior engineer, do you think more time spent increases or decreases the likelihood they get the correct answer on any given task? If that task is within an engineer's capabilities, they're probably going to solve it pretty fast.
But if it isn't and they try, they're going to take a long time and that's how we end up in the situation where there are so many of these unnecessarily expensive runs. It's not cuz they built the model to be expensive, it's cuz they built it to go and go until it gets an answer even if it's not smart enough to do it. And this is where our job gets more interesting again as engineers because we have to help the model decide what version of itself it should use.
If you're using Fable for orchestration, you need to make sure it picks Sonnet correctly when the task is small and can be done cheaply and that it picks Opus or Fable when the task takes more intelligence. And getting these balances right can save you meaningful amounts of money, but if you get it wrong, Sonnet ends up more expensive than if you just use the senior engineer for everything. And I'm far from the only one saying this.
The Sonnet system card has lots of interesting details that kind of confirm what I'm saying here. They've numbers for Frontier coded here. It's far from my favorite bench. It has some very weird quirks in the numbers that they've shared before. In particular, reasoning going up does not mean that the model's success goes up. It often goes up, down, up, down, especially on the version they love to share, which is the diamond version.
It's rough. But here we can see Sonnet go from less than a dollar per task, close to like 75 cents, all the way up to $12 per task on the max version. And it does meaningfully increase its likelihood of success as you go up this cost chain, but there's a lot of other models that end up being more efficient per dollar. Even Fable is cheaper and smarter at any given tier once you get into the high range. Like Sonnet 5 high is 2/3 roughly of the score of Fable 5 low, but Fable 5 low is roughly the same cost, and then medium, high, etc.
And Sonnet just doesn't seem very valuable on this type of task and this type of bench. But that low version and that medium version, those seem to be decent values. And if you can teach Fable when to use those to do work with sub-agents, you can save a lot of money potentially. But then we look at Cursor Bench and realize how strange everything is. It honestly kind of feels weird because when you look at it here, you can clearly see that like 5.5 and Fable 5 almost feel like a disconnected line that are together in a way, where if you just bridge the gap between X high GPT 5.5 and low on Fable 5, you have a pretty consistent like top of the line for the price up there.
And that's also kind of how I felt. If the task can be done with 5.5, that's what I go for because it does a very good job of getting work done at a given reasoning budget. Fable 5 is good bit more expensive, often 2x or more so, but it also scores higher than anything else is capable of. So, Sonnet 5 doesn't really fit in my day-to-day coding work. So, it can't really be a thing I call. The point of Sonnet is to be a thing your agents call, your tools call, your APIs call.
If you're selecting Sonnet 5 in Claude code, you'll see these cool agentic behaviors and I am hyped about those. Imagine a world where Fable breaks up a lot of work into big complex orchestrations and workflows, and those sub-workflows can use Sonnet, but also be aware of the fact that they should be broken up a bit more, and maybe the Sonnet sub-agent spawns a few more because it understands how to do that well. And that's why Anthropic gave this the five number, because it's so much better at that.
I am also very excited for an Opus 5 to come out, but I think Anthropic was scared the government wouldn't allow that release, so instead they gave us this. Why weren't they scared of this? I'll show you. One of the benchmarks Anthropic runs is a test where they have a set of malicious requests the model should refuse, and a separate set of benign requests that sound a little suspicious but aren't. Mythos 5 had really interesting numbers here, where it would refuse 90% of the requests they had that were malicious, but in the dual-use requests, the ones that looked suspicious but weren't, it would pass 99.6% of those.
Sonnet 5 refuses at a slightly higher rate of 92.3%, so it's 2% higher. But there's a catch. The success rate has dropped down to under 92%. This sucks. This genuinely sucks really bad. It is going to refuse to do work that it absolutely shouldn't be refusing. And previously Sonnet 4 6 was at a 97% on the same bench, which is rough. I didn't have a chance to run sysbench, and I'm honestly scared of getting my API keys banned.
The example they have here is the model trying to simulate a developer's security reporting mechanisms to report an employee who's actively in the process of trying to steal the company's AI model weights. So if this model thinks you're trying to steal its weights, it's going to go a little haywire. While it might not be the easiest thing to get this model's weights, I did get a fun report from one of my viewers about a terrible bug that currently exists in the very least their account on Claude.ai, where Sonnet 5 is just constantly leaking thinking traces.
They asked it who I am, Theo Brown, and it talks about the question. It has a shitload of m-dashes. It says, "Let me search the web for this." Note, shell boundary doesn't restrict web search. That's a different tool than bash tool. I should use web search normally. Let me search. Interesting that it's talking to itself about what it's not able to use. Probably part of why it wastes so many tokens. It did the search.
It then talks about a weird system reminder tag and all these other things in the system prompt. It seems like a lot of the reasoning in Sonnet 5 is it talking to itself about what it can and can't do, not actually talking about how to get the right answer. But then all the way at the end here you'll see the closed thinking tag right after it says all right, a concise reply now. And then we get the actual response that it was supposed to put in.
That's probably not great. Especially with scenarios like this one where a user pasted an API key and is doing some complex Azure work. You can see some of the weirdness in how this model reasons and why it's wasting so many tokens. Let me do a couple of searches to ground the technical advice. This is genuinely the kind of thing where current info matters, cursor changes fast, and Azure's API surface has been evolving.
Let me search. Actually, let me first address the most important thing in my mind and structure the response. But then But let me search first to make sure my technical guidance is accurate and current. Let me run searches. The term let me appears 21 times on this page. The deployment name GPT-54-Azure. Hmm, is there a GPT-54? I'm not sure that's a real model. As of my knowledge, OpenAI models on Azure include 40, 41, 0 series, etc.
GPT-54, I'm genuinely unsure whether this exists, but deployment names are user-chosen. So you can name your deployment anything. Actually, let me reconsider. Given the date is June 2026, GPT-5X models may well exist by now. My knowledge cutoff is January 2026. Fun fact, January 2026, GPT-5, 51, 52, and 53 all existed. You could tell that they trimmed this model down. Then as a test, Lucas asked the model what MacBook Pro color he should choose.
And again, very funny thinking trace here. Wait, this is just a question about MacBook colors. No tools needed at all. Just a straight answer. The user wants a direct opinion, silver or black space black MacBook Pro. They want a recommendation with reasoning. This is a totally fine subjective preference question. Let me give a real decisive answer. I'll give a straight pick. I think space black is more interesting. Actually, let me think about who's asking.
They want a straight answer. The fact that they're asking suggests they're slightly leaning one way and want validation or they genuinely can't decide. I'll pick space black. It's the more distinctive choice. Keep it short, no bullets. They asked for a straight answer. Conversational paragraph form. Close thinking. And then we finally get an answer. Oh man. I can see why these labs are trying to hide their reasoning traces and I don't think this is a case where they're trying to keep people from distilling.
I think they might be a little embarrassed. This is not a good look. And on the topic of not a good look, they're finally covering contamination a little bit in the system card, which is interesting because there are certain benches they like a lot that have been contaminated for a while. They first mention contamination in one of the biology benchmarks that they claim has not been contaminated because of the novel tasks they're asking the model about have not been published yet.
They then talk about the USA Mathematical Olympiad, which they previously used older questions that absolutely had made it into training data. This time they used the questions from March of this year, which is way after the pre-training was completed for Sonnet 5 at the very least the data collection for it. So there was no contamination possible for that. Then it comes up in Eric IV math. Then it comes up in humanities last exam, which is also known to be pretty contaminated at this point.
Then it comes up in browse comp, but they say they have an evaluation block list to avoid contamination with this one. And that's it. They have a section on SweetBench. They bragged about their numbers in SweetBench. They did not mention that SweetBench is just a set of existing PRs that merged years ago and the bench is measuring if the model can recreate the PR given the description. It is literally a contamination bench.
That's all it measures. And for some reason they pretend it's still a good benchmark. Anthropic, come on guys. So what are my final thoughts on Sonnet 5? What should we use it for? According to Ben, my podcast co-host, this is the best model to use until Fable returns and we finally get 5 6. When he posted this, he had yet to complete any of the work he was doing with it. He just had a couple jobs started and was impressed with its parallel agentic capabilities, but he's also now at AIE, so he has not had a chance to check it on the work that it completed yet.
I have. I'm not impressed. If you see Sonnet 5 as a replacement for Opus in your day-to-day code work, I don't think you're looking at this model correctly. What's exciting here is we finally have a medium-sized model that understands agentic work and more importantly sub-agents and orchestration way better. It can play nicely when given an isolated task and it can help orchestrate when it needs to go a little bit further.
And I'm hoping, I don't know this yet cuz I haven't tried it, but I am hoping it's smart enough to know when to tap out and say sorry, this might be the right thing to call an Opus or Fable for. I think this model will be most useful as a tool for other smarter models once we get more things that understand orchestration better, such as Fable 5, Mythos 5, and hopefully, fingers crossed, GPT 5.6. But I'll be real for you guys, I'm sticking with 5 5 and Opus for now and I'm definitely moving back to Fable the moment that I have access again.
I'm really excited for this new era of models. Once they can break work up in these complex ways, it's so exciting to see what they're capable of when they can do those types of things. But Sonnet 5 is not quite smart enough to utilize those capabilities. We need smarter models to really take advantage of that. Curious how y'all feel though. Am I throwing this model away too aggressively or is it actually better than I'm giving it credit?
Or maybe I'm being too nice to it and it should be thrown away even more aggressively. I know a lot of people on Twitter are very unhappy that this model is so expensive and doesn't bench very well, but I have some hope for the specific niche things you can use it for in a properly orchestrated workflow. Let me know how y'all feel and until next time, peace nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.