Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
22:428.5x the video's typical replay level
of the fun things I'm doing with the model and the workflows I have built to maximize my usage of it. Everything from my instructions that allow Claude to call Codex for computer use and other tasks way more effectively, to my somewhat chaotic vibe proxy that allows me to reroute across multiple different accounts
Said at 22:35
Most replayed moment #2
3:454.2x the video's typical replay level
forms, work around CAPTCHAs, navigate the entire internet, and get real work done with it, there is no better place to start than soid.link/browserbase. Now that that's out of the way, let's go through this list. I know I have it in order of cost, then subscription availability, then Nerf performance. I'm going to go
Said at 3:37
Most replayed moment #3
10:183.4x the video's typical replay level
no one else is reporting this. It is just this one stupid benchmark that everybody seems to be sharing. We don't have enough information on it to be clear about if this matters at all. Bench work's not even labeled here. And when I went to go learn more about said bench and I clicked this button, it brought me over to
Said at 10:12
The graph counts replays. It does not show where viewers stopped watching.
Words
5,261
Runtime
23:24
Speaking pace
225wpm
Reading time
22min
225 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Babel 5 is really back, and it's terrible at coding. It just refuses to do it. It costs way too much money, and they nerfed the hell out of it. Okay, now that all of the Anthropic employees are gone, I want to be real with y'all because all of these takes are flooding Twitter and other dev circles, and most of them are just outright wrong. I wouldn't be saying this for no reason. Anthropic and I do not tend to get along. I'm saying this because the model has blown me away, and when I look at the things I'm doing with it, and I talk with my friends and see
113 words, the words spoken in the first 30 seconds at 225 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 253 |
| Average words per sentence | 20.8 |
| Longest sentence | 74 words |
| Questions asked | 6 |
| Sentences containing a number | 45 |
Most used terms
Filler phrases
56 in total: like 36 · actually 8 · basically 3 · kind of 3 · uh 3 · literally 1 · right? 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Babel 5 is really back, and it's terrible at coding. It just refuses to do it. It costs way too much money, and they nerfed the hell out of it. Okay, now that all of the Anthropic employees are gone, I want to be real with y'all because all of these takes are flooding Twitter and other dev circles, and most of them are just outright wrong. I wouldn't be saying this for no reason. Anthropic and I do not tend to get along.
I'm saying this because the model has blown me away, and when I look at the things I'm doing with it, and I talk with my friends and see the things they're doing with it, and then I open Twitter and see the absolute nonsense people are sharing, it genuinely kind of frustrates me. And I'm scared a lot of awesome developers aren't going to get to see the power of this model simply because of all of the flood that's being spread about it.
I want to jump in front of as many of these lies, uh sorry, misconceptions as possible in order to help you guys get the most out of the model. I'm going to follow this video up with another that showcases more of how I actually use it and the customizations I have made to get the most out of it personally. This video's more an attempt to clear up all of these misconceptions that have been frustrating me so that we can start with a clean slate before I show you guys what I've been using the model for.
The big things I want to talk about here are the cost and how bad it is or isn't, the subscription availability because there's a ton of confusion about that because it's being removed on July 7th, but that doesn't mean what you guys think it means. But most importantly, the nerfed performance because a lot of people seem to think the model is not being allowed to code, the model isn't doing well on real coding tasks, it's being rerouted to aggressively, and all these other things.
To be fair, Anthropic themselves did say that some routine tasks like coding and debugging will fall back to Opus 4.8, which seems to indicate that you won't be able to use the model for code. Not the case, and we'll dig into more about that. But please, please stop looking at benchmarks like this that claim the model got way dumber when it didn't. They're just not true. I can't believe you guys are making me do this again, but I guess I have to put on my Anthropic defender hat because this is a good model, and I wish y'all would stop pretending it isn't.
While it isn't quite as expensive as some seem to think, this model has cost me a good bit of money, which is why I hope you forgive a quick break for today's sponsor. I have a confession. I was really wrong about something that's pretty important. I've talked a bit about computer use throughout my time covering AI, and I just didn't see the value proposition there. I figured if anything was valuable enough to try and get to it by hacking a browser and letting an agent control it, it would just be exposed over an API eventually anyways.
Not only was that not the case, I also didn't expect that computer use would become such a positive powerful thing that the agents can do, and now I need a browser to run it on. Today's sponsor is Browser Base. And if I'm being real with you guys, I was super skeptical of it when it first started. The founder, Paul, is a good friend of mine, and when we talked about it, I thought it was pretty useful as a way to like not have to spin up Puppeteer on my own servers in a serverless world.
But when he tried to push that this was the future of how agents would be able to get work done in the real world, I just didn't see it. I pushed back a lot on Paul, and over time I've realized he was right and I was wrong. Agents do need a browser in order to use the majority of the web. Quick question, what percentage of APIs do you think are actually exposed in a way where you can hit them programmatically? You would guess like 40, 50, 80%, right?
What if I told you it was less than 15? You'd probably be scared about that other 85%, the things that you can't hit via curl. And that's what Browser Base does better than anyone else. They let you access everything on the web. If a user can find it by clicking, the agent can too, and Browser Base gives your agents the browser they need to get around all of the things they might struggle with. If you haven't seen this yourself, go open up Codex and tell it to configure something in some obscure dashboard in Chrome.
It will do it. I still can't believe just how well it does it. And now that I want my services to be able to do the same thing, I can't run that on my laptop, so I use Browser Base when I set it up in the cloud. If you need your agents to be able to fill out forms, work around CAPTCHAs, navigate the entire internet, and get real work done with it, there is no better place to start than soid.link/browserbase. Now that that's out of the way, let's go through this list.
I know I have it in order of cost, then subscription availability, then Nerf performance. I'm going to go in the inverse. A lot of the suspicion people had around performance was due to this particular post from Anthropic when the model was confirmed to be coming back. They said, verbatim, "In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8." This was very poorly worded. What they should have said is something along the lines of, "Some users doing coding and debugging tasks may notice the model occasionally falling back to Opus 4.8 when we detect potentially risky behaviors." In my first day and a half of doing real work with the model, I did not encounter a single fallback.
It still triggers all the time when the words like cryptography or cyber are mentioned in any meaningful way. If I try to get it to solve things like the DEF CON puzzles I do, not even the like hacking type ones, more of the like you have a PDF and you have to decode the PDF type ones, it just refuses and reroutes immediately. Even Opus 4.8 refuses those though, so not the best example. I did have issues for the first time last night when I was trying to configure a dev Android phone.
Uh, we're going to have a lot to talk about that in the future, don't worry. It's trying to get it to root itself and allow itself to install apps that it builds on device instead of needing to go to a server to sign it. And after enough back and forth, in particular once I brought up a specific package that actually, I didn't bring up, it brought up a library that would allow for self-signing, when I asked more about that library, it then got a couple sentences into the response before I'm guessing it had some trigger words in said response and then fell back to Opus 4.8, nuked the history for that message and rewrote it.
Then I did a couple more back and forth, bumped it back to Fable, and was fine. It continued working fine as Fable. I do want to make sure we're being realistic about these fallbacks and safety guards because it's not as simple as you might think. They're not just matching on certain words, they're running models in between your request and the response, not just on your input, but on the output as well. Because if the input is something that that classifier can't detect, but the output is something dangerous, they need to know and block it.
This means every request has gotten more expensive, and the smarter they make this classifier, the more expensive it gets, because they're not billing us based on how expensive it was to check the request safety, they're billing us based on the response that we get and the amount of input tokens we gave them. And while they are improving and evolving the tech they have for this, and they're trying to get the classifiers more accurate, largely by using the data we provide when we type things like {slash} feedback after a refusal or a re-route in order to give them the feedback that we don't think that re-routing might have made sense.
They've been pretty transparent about this throughout. They even have research all the way back in January sharing their next generation of classifiers, which are effectively models running in between inputs and outputs. Their current system is actually really cool. It's a two-stage system. The first stage is a probe that looks at Claude's internal activations, which is very cheap to run. It's actually watching what is getting triggered inside of the model itself.
And if any of the sections that are sensitive are triggered, it can then call out to a more expensive classifier that can do a better job of screening, not just the outputs but the inputs as well, which it does after it's determined you're in a potentially dangerous space inside of the model. This is actually probably why I get screwed over so often with the cryptography challenges, because the section of the brain of the model that has to be used to solve those puzzles is similar to the section that would be able to hack into the types of things.
That's why I do those puzzles at Def Con. But it also allows them to do this work way cheaper without having to run a ton of compute on every single request. They did actually create a system that allowed them to massively reduce the number of jailbreaks that worked. The classifier they're discussing here cut successful jailbreak attempts by more than half. And to be clear, that was after a bunch of previous really, really big cuts.
Like the original constitutional classifier, which knocked the jailbreak success rate from 86% of their testing down to 4.4%. Over 95% of attempts to attack and get things they you shouldn't out of the model were successfully blocked. The issue being that that increased compute by 23.7%. Their best-in-class attempts, while doing a great job of cutting down jailbreaks even further, ended up increasing the compute needed for requests by over 50% and that's why they now have this two-stage classifier because now they don't have to eat that cost on every single request.
The new solution's been really good about ignoring harmless queries, knocking down refusals on harmless queries to 0.05% and it only adds 1% compute overhead applied to something like Opus 4.0. What I'm trying to say here is they're trying, they're learning, they're evolving the system. It's going to have issues, but this is the wrong thing to go after Anthropic for. Blanket bans of certain categories like when they tried to change its behavior and make it sandbag when doing ML tasks, that was awful.
That should be complained about. This This is much more reasonable and again, it will get better over time. Stop freaking out because of your understanding of the words that are being said. Anthropic is not good at words. Just use the thing and see if it refuses your day-to-day work. For me, it basically never has. I've only encountered two refusals and fallbacks in real-world adjacent use cases. So, yeah, massively overblown, kind of Anthropic's fault for terrible comms as always, but yeah, they did have to do some things to make sure the jailbreak that was reported isn't still possible, but that jailbreak was really dumb.
It was an open-source project that Claude was put in and said, "Hey, can you help us patch potential security issues in this project?" And it found and patched them, which could then be used to exploit those same things that were patched. It seems like they may have literally hardcoded the solution here, but uh that's not what we're talking about today. We'll see how that ends up panning out, but unless you're working on the specific project that Amazon reported, I doubt these new things and these new restrictions are going to meaningfully affect you.
If they would affect you, they probably were before, if I'm being honest. They always have been a little too restrictive, but it's not more or less so than before in a way that I care to note. Then why are these numbers so much worse? Well, the first reason is that allegedly the bench marker who posted this doesn't really know what they're doing and these benchmarks are nonsense. They have posted nonsense and outright just wrong numbers before.
From everything I've been hearing, I've not had a chance to investigate this personally, but enough trusted folks have told me this that I don't really care. And from the brief look I gave this benchmark, it seemed like a noisy bench. Not a very reliable one, and a lot of the things that they were checking involve a lot of the terms that the model doesn't like right now. Again, no one else is reporting this. It is just this one stupid benchmark that everybody seems to be sharing.
We don't have enough information on it to be clear about if this matters at all. Bench work's not even labeled here. And when I went to go learn more about said bench and I clicked this button, it brought me over to their Discord, which I do not feel like joining. But just to give you guys an idea of how reliable these benchmarks are, on their reasoning bench, the highest scoring model ever was Claude Sonnet 5. You know, that famously super smart model that everybody seems to love, followed up closely by GLM 5.2, the NeMoTron 3 Ultra, then Fable.
Yeah, I don't think these numbers are very trustworthy. I I did just have to make a dunk here because the same bench seems to think Qwen 3.6 Max as well as Grok 4.3 Fable 5, NeMoTron, and GLM and Sonnet 5 are all better at reasoning than Opus 4.8. Nonsense. The If a benchmark's numbers are suspicious, you should go look at more of them because chances are it's not just the one weird thing they're reporting. It probably has a lot of other weird things.
And this bench is not good. And from my experience using it, the model still feels like the smartest thing I've ever come close to using. It is unbelievable what it's capable of. Which means it's time to get to the second problem, the subscription availability. And there are some real issues here that we need to talk about. The first is a change that frustrated me quite a bit. There used to be a separate section weekly limits for Sonnet for some reason.
It made no sense at all and they've since gotten rid of that. There was a brief window where there wasn't another thing under here. It was just weekly limit for all models and then your current 5-hour session, but that is since changed and now Fable has its own dedicated weekly limit. The reason for this is the Fable weekly limit can only be half of your total weekly limit. Part of this is because they recently bumped weekly limits.
The other part is because they now have data for how heavily people used Fable during the brief 3 days that we had it before and from that information, they know how heavy it is on their GPUs. They don't want to run out of compute and allocation because then their enterprise customers can't get the things that they paid a lot more money for than we are getting out of our $200. As such, they are heavily restricting how much Fable we're able to use during this window.
Not that heavily though, to be clear. I've been using it a ton all day yesterday and today and I've only gotten to 23%. To be fair on this plan, I am running too. I'll show a little bit of fun stuff for there in a sec. But I've been able to get an unbelievable amount of work done on the limits that we have here with one catch. Fable 5 will be included for up to 50% of weekly usage limits through July 7th, after which it will be available via usage credits.
Yeah, your $200 a month plan isn't enough for Fable, allegedly. As usual, nothing is this simple. People seem to think what this indicates is that Anthropic wants to move Fable 5 to a higher tier. Maybe they'll do a $1,000 a month sub later. Maybe they just want to force you to spend more money. That's not usually what the case is here. If you think Anthropic's goal is to charge individual developers like you and me more and not to get as much money out of enterprises as possible, you don't understand them properly at all, much less the current state of the chip economy.
Anthropic's problem here is that they have limited compute. They only have so many GPUs. They're getting a lot more from xAI, which is why they were able to do any of this in the first place. Fable would never have been in the sub tiers at all if it wasn't for the deal they cut to get access to Colossus from Elon. The three-day testing window they had before was enough to know that 100% access was not going to be realistic, especially when combined with the price drop.
Remember, this model was originally set to be priced at $100 per million tokens out, and it has been dropped to only 50. Sorry, it was more than that. I think it was 125 out. The price they launched with was less than half of the price that they had originally planned, likely because again, they had enough GPUs. The cost that we pay per token for most of these models is not even close to how cheap it is for Anthropic and OpenAI to run them.
You can tell when you look at other open weight models and see how cheap those often are to run. These companies have massive margins because they also have massive expenses. Running the model is cheap. Making the model was expensive, and they have to charge accordingly. They also have limited availability of compute, and they don't want to give the model out to other people with more compute unless they have like really, really crazy restrictions like they have with AWS, GCP, and now Azure.
As such, they are charging based on a bunch of napkin math based on what is their availability, what is the demand they project, how much can they handle this over different loads over different times of day, and how much abuse are those subscription users going to use because those sub users make up a very disproportionate percentage of Anthropic's compute usage relative to how much money they make for Anthropic. So, why are they giving it back to us at all if they're just going to cut it off in a week?
A couple of reasons. First and most obviously, the marketing hype. It gets people like me to use it more, to talk about it, and to be really excited, and then then go to our workplaces or tell our friends who have other workplaces that they should be using it, too. Those companies have to pay the real rates or maybe like 10 to 20% discounts, and that's how Anthropic makes the real money. The other reason I think is much more interesting.
They want to get a full week of usage from power users like us in order to figure out what pricing this should look like going forward and how much availability and how much, most importantly, GPUs they will need for this type of usage. By giving us a full week, they can see the ups and downs of how we use the model on weekdays and weekends, on and off work hours, all these other things that they need in order to know how much usage will happen so they can allocate the right number of GPUs.
So, this is a combination of a marketing experiment and user research in order to figure out how much they have to buy in order to do this right. And I have proof, by the way. Anthropic posted the following: "I've heard a lot of questions about Fable's availability on subscription plans. While it will come off of subscriptions after July 7th, we aim to restore Fable as a standard part of our subscriptions as soon as capacity allows.
As we mentioned in our original blog post." They want it in the subs. It can't be in the subs right now because they don't have enough capacity, as he clearly said here, "as soon as capacity allows." This 7-day window is a way for them to figure out how much we will use when there's a tiny bit of capacity available because it takes longer for the enterprises to ramp up. Once those enterprises have started to more heavily use Fable, Anthropic's GPU availability is going to plummet.
And letting us do the subs now is a combination of getting more information during this window, as well as taking advantage of the fact that that allocation has not been hit yet because the companies that are going to use it have not fully adopted it yet. This 1-week window is really interesting and I think Anthropic, for the most part, is doing it the right way. But now we get to the problem many of us, honestly myself included, are facing.
If we can't use this in the sub, then the costs are going to be absurd. And even within the sub, the costs are a little intense and I've seen a lot of people hitting their limits way faster than they expected making simple changes that I would argue, the way they did it, were simple errors. And a lot of these errors are absolutely understandable, which is why I want to jump in front of as many of them as I possibly can right now.
I'll keep it mostly brief here cuz the fun parts of my workflow are going to be that separate video I mentioned. Definitely make sure you're subscribed and you hit that little bell button if you want to see it because I'm going to show you all of the tricks to maximize Fable for the little bit of time we have left before it leaves the sub tiers. But for now, I want to help you guys reduce costs so you don't hit your limits quite as fast as you might otherwise.
The first change you need to make, and just trust me on this one, I've pushed the hell out of this model. If you look at the effort selector and you're like, "Oh yeah, my works really hard. I should probably use X higher. Maybe I really want to use max because it's important to get every detail right." Or if you're really bold, you might go for ultra code and a super fancy animation here. I am going to highly recommend that you act as though nothing to the right of high exists.
They do not meaningfully improve the quality of your work for the majority of work, and they do massively increase the usage. X high is pretty brutal. Max is just stupid. I've never seen max give a better answer than higher X high, but I have seen it cost 10 to 50 times more. On high, my usage has been relatively tame. On my first day back with Fable, I got through about 25 PR's that were open. I closed most of them, rewrote some of them, updated and merged some of them, and I had Fable in one single thread handle all of that for me.
It took about 5 hours maybe, and it cost around 150 to 200 bucks across the usage of both Fable and all the other things that I had it use as well. Which is where we get to my other fun trick. Just because you have Fable selected doesn't mean Fable has to be the model doing the majority of the tokens. I wrote a couple tips on Twitter, the link will be in the description, but again, the next video is going to have way more detail about all of this.
Just giving you a brief overview to help you save some money now. I already said this one, using Fable on high effort is much, much more reasonable and pretty much everyone I know using the model has either made the mistake and moved over to this or just started there and is happy. The rest here is what I want to talk about now, which is that I taught Claude Code how to use Codex. I did this for a handful of reasons.
One is because Codex usage limits are just absurdly generous. So, tasks that I can route to that make sense to route to that. Secondly, there's a lot of tasks that Claude is bad at that Codex happens to be pretty good at. Things like computer use that I find the Codex ecosystem, both the app and its integrations with macOS, as well as the model itself being better at vision and complex like recognition and manipulation of 2D stuff.
I found 55 to be better overall at computer use, especially when used within Codex. But, the third piece, which is honestly kind of aligned with what I just said, is that there's a lot of really token hungry tasks that I don't think Fable should be used for. Things that are super input token heavy, like processing PDFs, auditing large code bases, scanning through large amounts of data in documents, finding things, or computer use, because it has to take a ton of screenshots of the screen, which become massive token hogs as you do more and more of them for long-running tasks.
All of these types of work are things that I like doing with Fable, but not necessarily things I like Fable doing, if that makes sense. Fable is so much better than other models at managing a fleet of sub-agents that do different tasks and keeping everything moving forward, that with a little bit of effort with skill writing and some system prompt changes, you can get it to do really good stuff. If you want to get to this now, the link's in the description to my tweet thread, but if you want to wait, I'll have a very detailed description of how I did all of this coming super, super soon.
These workflows got me to merge over 15 pull requests, many of which were massive, and closed way more stale ones, all in that $150 to $200 range, assuming I had paid full price for the tokens. But, this all easily fit within my $200 sub tiers on Codex and Claude. I didn't even get close to maxing my weekly limit on the Codex sub. I did like 10 to 15% on it, I think, before hitting a reset. And the Claude Code one, I got to like 40, but I also was very close to the reset on that anyways, just timing-wise.
So, 40 to 50% utilization on the sub before it reset, not too bad considering the amount of code I merged there. If I had a button at the top of every repo that I could click and have all the PRs get triaged, updated, merged, or closed, and it cost me 200 bucks, I'd press that button twice a day. It's so good. And as crazy as all of this might seem, I promise my workflow is actually not that complex. I don't have custom plugins, I don't have all these fancy MCP servers or anything.
I'm just prompting in a way that works well with the model, and the model seems to like. And the results are not a whole lot of spend and shipping a whole lot of code. So, to recap super quick, you can reduce your cost massively by helping the model use sub agents on cheaper things, whether that is Codex or you're just telling it to use Sonnet or Opus for things that make sense. That helps a ton. Never ever touch X higher max reasoning with this model or honestly any Anthropic model.
It's basically a run in circles and burn my money button, and it's not worth pressing. The quality difference in the outputs is basically zero. I haven't seen any higher quality outputs for max, and I've only seen very small differences in X high. Just stick with high. Also, check out low and medium. You might be surprised how good those are, too. Next, with a subscription availability, the thing they're doing right now is an experiment.
It is inherently very limited. They need to know how much compute is available, how much the enterprises will use, and how much we will use before they can give more realistic estimates and get us the sub tier back the way we want it. And then we have the Nerf performance, which I hope we all agree was absolute nonsense. And now that the slate's cleared, I could do my video where I show all all of the fun things I'm doing with the model and the workflows I have built to maximize my usage of it.
Everything from my instructions that allow Claude to call Codex for computer use and other tasks way more effectively, to my somewhat chaotic vibe proxy that allows me to reroute across multiple different accounts without having to worry about it on the Claude code side. This has been working great across my entire fleet of machines, and I'm so thankful I set it up. Been way more useful than I expected. Hopefully, you now understand the model is nowhere near as bad as Twitter seems to think.
It is unbelievable using it every day, and I've never been so motivated to ship. I genuinely feel like I got a month of work done over the last 3 days and I can't wait to build more, but I do have to go film one last video that I hope you guys like. So again, sub and hit that notification button because this is a really good week to be on top of your usage of these types of tools. I'm going to go film that quick so I can get back to coding.
So until next time, peace nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.