Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Theo - t3․gg's most watched videos.
Most replayed moment at 2:36
5.6x that video's typical replay level
Well, you have nothing to worry about cuz the first million users are free. Get yourself enterprise ready at soidiv.link/workos. I'm very excited to read into what Linear is cooking here. You know they're cooking something different cuz this is the only not dark mode page I've ever seen Linear ship. They actually took
Said at 2:29
Most replayed moment at 18:41
4.1x that video's typical replay level
about it, that will save you so much trouble because you won't go waste all of the time building the good code when you're building the wrong thing. Good code has two important characteristics. It solves the right problem and it doesn't suck to read. People get way too focused on the second part and they often, if not
Said at 18:34
The graph counts replays. It does not show where viewers stopped watching.
Words
6,329
Runtime
30:58
Speaking pace
204wpm
Reading time
26min
204 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
They're back. Okay, let me catch you guys up quick. Last week, Anthropic dropped a new model, Opus 5.5, and it was unbelievably good. It was so unbelievably good that OpenAI rushed out two model drops, GPT6 Soul and GPT6 Luna. You might have noticed I didn't do a video on those models. There's a reason. They weren't that good. I was not particularly impressed with either of them and didn't really have much to say. But there was one other model I happened to get early access to that is now available for all. That model is called GPT 6.1 Soul.
102 words, the words spoken in the first 30 seconds at 204 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 391 |
| Average words per sentence | 16.2 |
| Longest sentence | 65 words |
| Questions asked | 7 |
| Sentences containing a number | 108 |
Most used terms
Filler phrases
63 in total: like 39 · actually 12 · uh 7 · kind of 2 · literally 2 · I mean 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
They're back. Okay, let me catch you guys up quick. Last week, Anthropic dropped a new model, Opus 5.5, and it was unbelievably good. It was so unbelievably good that OpenAI rushed out two model drops, GPT6 Soul and GPT6 Luna. You might have noticed I didn't do a video on those models. There's a reason. They weren't that good. I was not particularly impressed with either of them and didn't really have much to say. But there was one other model I happened to get early access to that is now available for all.
That model is called GPT 6.1 Soul. And this model made it very, very hard to film that Sonnet 5.5 video because I knew something was coming. Something surprisingly cheap, surprisingly capable, and most surprisingly, better than Astra. So yeah, I got a lot to say about this one. I'm filming this at 2 in the morning, right after filming my Sonnet 5.5 videos. So, pardon me for stumbling over a few words here and there. I'm doing my best to get this out as reasonably quickly as possible because I want to have some coverage and I'll be real.
It is also quite fun to cover these things before they are out so you are getting my true honest take and not the distilled version of what everyone else is saying. I'm sure this model is going to cause some pretty crazy waves. So, uh it will be nice to have my take out initially separately first. As always, I feel obligated to remind you I do have early access, but I'm not being paid in any way, shape, or form. OpenAI has no influence over what and how I say things, just when.
They've politely asked me to wait until the model is out to talk about it, which makes a lot of sense. But I have to wait for one other thing first. Today's sponsor. In order to build good software with agents, you need to get feedback. And let's be real, they're getting a lot of that feedback from our CI. That's why we've all been seeing our CI bills skyrocket. And also why we've been getting more and more frustrated with GitHub actions.
Today's sponsor is depot and they're here to solve all of this and more. Not only can they make your CI up to 10 times faster, as well as your Docker builds up to 40 times faster, especially when they're downloading cache, they're also cheaper and they give better feedback for your agents. All this is possible due to depot metal. They're running their own bare metal with AMD epic processors that are way faster than what you get from traditional CI providers like of course GitHub actions.
If you want it to be a drop in replacement, it absolutely can be, but APIs are so much better that you should probably use those instead. They enable parallelization and most importantly resilience when GitHub inevitably goes down randomly for no good reason. We've had our releases get blocked because we weren't using depot. And I'm so thankful that I've been moving more and more stuff over. For example, when Ben moved pick thing over to bun, we immediately had some CI failures.
Normally, this would be obscure piles of text that our agents pars through for us. But when we use depot, it becomes way easier to see. They'll even analyze the failures and give suggestions which make it much simpler to get this feedback back to our agents. This is especially useful when you tell your agents that you can use depot because they'll no longer have to push changes and wait for that to trigger a build. They can just run the CLI to trigger the exact same CI that you'd be triggering through GitHub instead.
No longer do you have to file PRs with broken code just to get feedback to your agents. They could just run a tool instead. Your agents will also get way better breakdowns of what is taking so long in your actual CI runs so that you can figure out how to improve them and make them faster and more reliable. You and your agents deserve faster Docker, faster build times, faster CI, better results, and ideally a cheaper price.
Get all of that and more at swive.link/devo. Let's talk about this model a bit because it is not quite what I expected and it's probably not what you guys expected either, especially when you consider that GPT6 soul just came out like a week ago. It'll be around a 1 week gap from 6.0 soul to 6.1 soul. I also want to disclose the numbers I'm currently showing on my screen are unlikely to be exactly accurate because I am running Terminal Bench for myself.
In the first two times I ran it, I screwed things up. The third one seems to be doing much better. I didn't run it on medium initially, so there's a miss there. But low, high, XH high, and max, although the max run is incomplete, so I'm currently back filling scores from X high for the ones that Max either got wrong or in a previous run or didn't do yet because it takes like eight plus hours and some of these tasks. I this bench is nuts.
It's I'm more skeptical of benchmarks than ever now that I've been running a lot more of them myself in order to get the coverage I want to give here. For what it is worth, Terminal Bench 4 is state-of-the-art score here. As is deep SWE, although this one's weirder because it goes down on X high and max and stays even on low and high. But those even low and high scores are scoring around what Astra did on high. The difference being it's doing it for comically cheaper.
Switch over to the log scale, you'll see what I mean. This model on low costs 21 versus Astra on low costing $1.46 and Opus 5.5 on max getting the same score for $14.65. While I will gladly admit that Deep Su is far from a perfect measure of how good a model is at day-to-day code work, the fact that 61 Soul is scoring the same as Opus and is also 73x cheaper is at least worth noticing. Here's where I'm going to say some of the things that I probably shouldn't.
Considering that GBT6 Soul came out last week on Tuesday and this model's coming out this week on Tuesday, I think it's reasonable to infer that 6.1 Soul was not meant to be 6.1 Soul. There are things I'm not supposed to say, and I'm definitely walking a thin line here by sharing it. So, uh, I hope this proves I'm not paid off by Open AI because I'm about to give you guys info that I Yeah, just let let me get through this.
First and foremost, 6.1 is significantly smarter than GPT6, thereby indicating this isn't just a one bump. There is something fundamentally different here. Next point is that it has meaningfully slower tokens per second. That tends to indicate the model is bigger. Hard to know for sure. Seems like this model might be different. Most importantly, we have Tibo's tweet. What am I referring to there? Well, right before I started filming, Tibo dropped quite a wall of text.
The thing I want to emphasize here is first off that the pro $200 subscription is back. But more importantly, but more importantly is this sentence. They are changing how they calculate the usage in the sub. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old pro $200 plan. Why in the world would they do this? especially right now where there's allegedly an internal code red because Opus 5.5 is so unbelievably good and has made the $200 quad code sub such an unbelievable value.
The only reason in the world Tibo would post this right now is uh I don't know, maybe a new model is coming where their margins aren't as good. So, the ability to subsidize has gone down because remember you can get $8 to $9,000 of usage in a month on the $200 Cloud Code plan and you can get over 12 grand on the $200 codeex plan. I did actually run a lot of numbers before this and the amount you could get on Astra did go down slightly closer to like eight grand or so.
Hard to know for sure because they differ for everyone everywhere and it's hard to log all of this stuff, but from my math roughly 9 grand a month of usage. And that's where the price for this model comes in. This model is $2 per million input tokens and $10 per million out. This makes it way cheaper than 5.6 Soul was at launch. Half the price of 5.6 Soul after discounts and the same price as GPT6 Soul. 1/5 the price of Astra.
However, this is not the whole story cuz cash reads matter. And the cash read price for this model is going to be 10 cents per mill in. That's a big deal. OpenAI has not changed cash read price as far as I know ever before. It's always been exactly 10% of the normal read price. That makes it a 90% discount and now it's a 95% discount. That means they cut the cache read cost in half, massively reducing the cost for real world agentic use, which to be clear is what we're using these for most of the time.
So this makes the model absurdly cheap for doing realworld code work. That also means that they are almost certainly cutting into their margins. Historically, these margins are rumored to be as high as 95%. Like for every $10 you spend, they only have to spend 50. And as crazy as that sounds, it makes a lot of sense. especially when you consider how expensive it is to make and train these models. But that also gives them wiggle room to change things around a bit, which appears to be what's happening here.
That also means that if they were to keep subsidizing the same level that they were on the subscriptions that your electricity cost for your sub would be more than you're paying. So, I get why they have to change this. They've kind of just left the details out there for us to uh reverse engineer. So, uh take this as you will. 6.1 coming so fast seems to indicate it is not just a new snapshot of GPT6. So let's talk more about this model.
As I was showing earlier seems really good at Terminal Bench. Every time we refresh the numbers change because new runs come in and it looks like Max failed some things that X high passed which is why it just dropped a bit. But again pretty much all of these even high and X high are scoring higher than anything else ever has. And this is for me running this benchmark on random VMs on my network. So, uh, not the best suite to test against.
I also had to drop three particular tasks from it because they expected an H100 to work against, which I make decent money. I don't make H100 money, okay? But none of this is real world code work. So, let's talk a bit about that. Obviously, we'll have all the fun things like fish slop near the end, so stay tuned for that. But, I just want to fixate a bit on the costs here because the most expensive run with 6.1 soul for me was about $1.38 per task.
And the cheapest run with Opus 5.5 was $512. That's a four to 5x gap from the cheapest Opus to the most expensive soul. So, at this point, I would imagine you are hoping and praying this model is good and that it can actually replace Opus 5.5 for day-to-day work. And I promise we'll get some good answers to that in a bit. But first, we need to talk a bit about model behaviors here because this model is a part of the GPT6 family, which means it has uh the behaviors that are worth talking about.
I know I cite this diagram a lot, but there's a reason for it. The thing that made me so frustrated with GBD6 Astra wasn't that it was less intelligent than the best models from Anthropic, because it was more intelligent than the best models from Anthropic, and I would argue in many ways still is. But there is a problem. It is also dumb. It is smart and dumb at the same time. GB6 Astra would just randomly spike into the dumbest I've seen a model do this year.
Even worse than like some of the small openweight models I play with. It's still so deeply frustrating that Astra does this because on the other end when it does well, it's unbelievable. But these spikes got to the point where I effectively churned. I was only using my codec subs for computer use and I ended up just leaning on to Fable 5.1 and obviously now Opus 5.5 for almost all of my day-to-day work. So, have they addressed the spikiness?
Has GBD 6.1 Soul fixed the problems that I was so frustrated about with Astra? I would say mostly, not entirely, but for the most part, yeah, this is a much better model. Its peaks are not as high. This is not the incredible revolutionary 3D capabilities that we saw with Astra. In fact, I would put it slightly below 5.5 opus in most of those types of things. It is not as good at computer use as Astra, although it is close enough to the point where I have been happy using it for all of my day-to-day work.
I actually had 6.1 Soul go through all of my emails and find invoices that I had forgotten to pay or was behind on, mostly like investing stuff, and set up new tabs in Chrome for every investment I needed to wire, fill out all the details for me, and just leave me to hit send. It didn't get a single thing wrong, and called out additional stuff that I absolutely would have missed if I was doing this work myself. So, I'm literally trusting this model to wire money for me.
It's trustworthy enough for that. And honestly, I don't know if I would have trusted Astra with that due to the spikiness. 6.1 Soul much, much less spiky. From what I've heard from the other testers, they seem to agree with this analysis. I know for a fact that Julius and Ben, who have also been testing, have had a much better experience with this than Astra in terms of the spikiness. Julius called the model incredible.
Ben called it incredibly boring. And I think that's the best place you can be for a model drop like this. But as I had mentioned before, its peaks are not as impressive. While it does quality work the majority of the time, there are some tasks that are just at the edge of its capability that it will start to do weirder stuff on. For the most part, it's fine. But I I'm still reaching for Opus a decent bit. We'll talk more about the comparison later.
I do default to this model for a bunch of stuff, though. First off, as I mentioned before, computer use. I can't wait for Ultraast to be like an actual thing you can use with OpenAI models because when it is, this model is going to be crazy on it because it can already figure out how to navigate computer use totally fine. If it can suddenly do it six times faster, it's going to be unbelievably fun. Still not quite as good as Astro, but more than good enough that for the price difference, I wouldn't even think twice about it.
But as I mentioned before, there are certain things I would still occasionally use Astra for that I am more than happy to use Soul for. One of those things is deep code reviews. I have still found OpenAI models and the like Rottweiler nature where they'll dig into a problem and shake it and tear it to pieces until they find every single thing wrong with it. I find 6.1 soul to be incredibly capable in this particular way.
So, as you can probably guess, I had 6.1 Soul do some deep audits on orchestrator v2 and other parts of my real world code bases. In my orchestrator v2 audit, it performed nearly identically to Astra. I do believe it was slightly higher a score. Okay, not in this analysis, but in my other analysis, it did actually score very, very slightly higher, but it did it at about half the price. 297 versus 584. Sonnet was still cheaper and Opus was slightly cheaper as well.
The difference being neither of these models were anywhere near as thorough with their analysis. 6.1 Soul dug deep to find things, which is why it was able to get a score comparable to Astra, although it did admittedly burn way more tokens. Another task I've had a lot of fun testing with is asking the model to find opportunities to improve a code base. In this case, to improve T3 code, this is the one where Grock 4.7 scored strangely well.
Of course, Astra scored way better at an 83.8 versus the 80.7 from Grock 47, but GBD61 Soul hit it out of the park with an 87.4. 4. I didn't save all the prices for these runs. It's been a bit okay. But for Opus 5.5, it cost five bucks. And for Sonnet 5.5, it cost almost $9. With GBD61, it was $2.15. That's the difference. This model's token price is cheaper than Sonnet, but its token utilization is still maintaining OpenAI's usual efficiency, which results in just crazy price to performance.
This whole thread was particularly fun because I had Opus 5.5 review this model with a different name obviously, so I didn't know what it was. I went and edited the history after and it concluded very quickly this was a Frontier tier model. Its reviews and bug repros match the fixes that later merged. Its first draft code had real bugs which review bots caught. Four reviewers are still running. The local for Code reviews back Frontier tier again in a blind 10 model bench on the same prompt. 6.1 Souls placed first of the A7.4.
It found the fish slop runs and compared those two. It did say 6.1 souls quality output was slightly below Astras as well as the two OpenAI models with Opus 55 and Sonic 55, which we will absolutely show you in a bit. But I do want to call out the price here cuz it only cost $7 to run versus 15 for Sonic 55 and 50 for Opus. Opus' honest tier call was that this model is incredible for scoped work, top of the frontier.
Refine what's wrong and tell me the truth. I would choose it over Astra and about level with Opus for long unattended building. This was below Frontier follows its process rules even when they stop all progress and it does not ask for help. This I absolutely noticed. I had mentioned before a few times now that my TS Rust port that I'm making with Opus 5.5 is going way better than when I was working on that same port using Astra and Soul in the past.
I had that port running for a while with this model and it made no progress. It burned a shitload of tokens, but it didn't actually improve the compiler at all. Opus was able to from scratch restart it and get it working in a day after I had spent months and hundreds of thousands of dollars in tokens with this model as well as with Astra and 5ixole. Opus did in like a grand in like a night with just two subscriptions with the quad plan.
So for unattended long like heavy rewrite type stuff, Anthropic is just comically far ahead right now. And it also didn't have great judgment when I was using it for managing my fleet. And for those wondering, my fleet is all the computers I use for running all my agents and code because one computer is far from enough. I don't run any of them on this MacBook now. So, when I use this model to manage the fleet, it made a couple dumb mistakes here and there.
To be fair, so is Opus. Aster is the only one that hasn't really made too many of those dumb mistakes. But, like, I'm going to be so real. I am entirely done using Astra after this model. After I had Opus do all of this review, I asked it how much does it think this model should cost. It guessed $5 per mill in, 50 cents cashed, and 30 per mill out, putting it at Opus' prices roughly. It said that because it's performing like Opus.
Its speed should add a premium because it is quite fast. And it's not a pro model, which is where it expects those higher like $100 out tiered pricing things to come. Pro models aren't really a thing anymore. We just use Fable and Aster, but you get the idea. This is the funniest part of the whole thread, though. If OpenAI wants people to adopt it, I would expect $3 per mill in, 30 cents for cash, and $20 instead. That would still be a fair price for what it does.
To which I responded, if I told you it was $2 in, $10 out, and 10 cents per mill cash read, what would you think? I'd call that very aggressive pricing. For how you use it, it costs about a quarter of what I guessed. The cash price does most of the work. Agent workloads are 96% cash reads. So 10 cents for cash reads matters more than the $2 and $10 headline prices. When it looked at all of my sessions, its price guess would have been $5,700.
But after looking at these new prices, it redid the math and it would have been $1,550. That is a massive decrease. And for all my PR review type tasks, it was expecting those to be up to $10. And it's actually only up to $3. And that's for like heavy PRs with tens of thousands of lines of code. According to Opus, so don't blame me, blame Opus for saying this. First off, Opus says it becomes the default model for scoped work.
Second off, it says that bloated system prompts barely matter anymore because of the cash pricing. It's just noise. Third, it says long loops are still a bad idea, but not because of money. It's cuz according to it, the TS ROSport wasted 4 days and made no progress at all. And Opus even said they'd be suspicious of it lasting. They expect this price to go up in the future. I cannot fathom OpenAI ever increasing the price for a model, but Opus thinking they will is hilarious and shows just how good a value the model is.
I love this call out here. GBD 6.1 Soul did three rounds of work for about half the cost of Sonnet's single round. A lot of this comes down to how context was managed, both because 6.1 soul is much more efficient, so it's not doing as many calls that bloat the context. It's not outputting as many tokens that are like building up over time. So the average number of tokens being read per request was only around 110,000 tokens versus 360,000 for Sonnet 5.
The result is that Solot used under half as many input tokens as Sonnet, making this model significantly more efficient. Speaking of efficiency, I want to talk about these deep SWE scores a tiny bit more because this is a weird bench for me to have forked and include in these things. I actually did it for a different reason, not to compare against 6.1 soul, but to compare against a new release from Open Router, Jev Router.
Open Router added Jev router to try and optimize costs with your requests. And I thought it was an incredibly stupid idea. Once I started running it against benchmarks, I confirmed it's an incredibly stupid idea. It turns out a model that cannot reason, that is given a prompt and no context, cannot make a good decision around how hard the problem is. and Jev router ended up being Deepseek v4.1 flash router for the vast majority of its runs.
It around 60% of all the requests went straight to Deepseek 4.1 Flash, so didn't like it that much. It also routes to other smarter models, which should give it more of an advantage, but it ended up being more expensive than GPT6 Astro was on low while also taking four to five times longer cuz six Astro low took 4.6 minutes and Jev router took 20. Jev Router's average task took 104 steps whereas GB6 Astros took 19. You get the idea.
It wasn't very good. But the whole point of Jev Router is that it would be as cheap as possible to get a certain score. That was the promise on the tin. Whether or not you believe them is up to you, not me. I think it's Regardless, Jev Router was routing to Deepseek 4.1 Flash for the majority of its requests. Despite Jev router routing to the cheapest possible small openweight models from whatever provider will give it away for free, 6.1 soul on low got the same score for an eighth the price.
OpenAI is here to destroy any wins anyone else has in terms of efficiency. Completing this bench in 4.8 minutes for 21 cents with the second highest score I've ever seen on it is a massive achievement. Tying Opus 5.5 which took 50 minutes per task on max. a tenth the time and a 70th the price for the same score. If your work fits within the things 6.1 Sonnet does well, you should probably use it for everything. But if your work doesn't fit in it particularly well, you should probably keep using Opus and maybe give Opus the ability to call 6.1 Soul when it should for various tasks.
I'm almost certainly going to be setting things up so that Opus 5.5 can call 6.1 soul to do investigation work to try and like root cause bugs to do analysis of code bases to figure out what things need to be touched and why to help me triage real world work to help me review the work that Opus does and more. I'm kind of spoiling the ending here, aren't I? I'm going to keep using Opus 5.5 for now. Before I explain why, let me do the thing that I'm most excited for.
Fish lop. The first thing you might have noticed is the inclusion of slop in fish slop. This model did the horrible thing I hate where it surrounded the game in a bunch of absolutely garbage UI. And coming to this right after the 5.5 Sonnet demo hurts me deeply because Sonnet 5.5 did not make graphics anywhere near this good-looking, but at least it made a UI that was nowhere near this awful. And man, do I wish the bad UI is where the issue stopped.
I will turn on the sound. Oh god, it's blaring. It is stunning looking. The fish are some of the best. The model for the sub is way better. The propellers work way better. I'm going to mute the sound cuz that is looking pretty bad. I haven't even heard it honestly. But damn. like looks beautiful, but if you actually are playing it, one of the first things you'll notice is that the movement feels significantly worse than it does in either the Opus or the Sonnet versions that I have demoed in the past.
It Yeah, it it moves jank. It also has a significantly worse frame rate than the versions from the other models. It does have higher graphic fidelity, so that makes sense. Like the models here with the plants are significantly better than they were with the Sonic version. The dares of the fidelity of the extras in the tank is absolutely hilarious. Like, yeah. But god damn, I'm so tired of the unnecessary text everywhere.
This model does it worse than almost any I've ever seen before. We got take a breather paused. Your little world can wait. Back to the reef. Start a new tank. Slop01 feeder submarine. A little underwater chaos. The big little goal. Your little ecosystem. Little fish become big earners. Four meals and they're all grown up. Make the family a little bigger. A little golden overachiever. There's so many of these. There's like 20 plus of them.
And I promise you guys, as soon as I saw this, I took a screenshot. I sent it to OpenAI and I crashed out in the Slack because I cannot fathom how they haven't fixed this problem. This model is unacceptably garbage at UI. It has regressed again. And if you're looking for a model that can make frontends that don't suck, go spend your money somewhere else because it should not be spent here. This model sucks at front end.
It sucks at design. It has no taste and you're going to have to bring your taste yourself still. But it is admittedly really good at Blender. If you give it like a screenshot of a thing you want it to model in 3D and say, "Hey, you have Blender over the CLI. go make this. It will and it'll do a pretty damn good job, but I would never have it make the actual mechanics for my games because it feels awful to play. It also has like nowhere near as much gameplay loop.
In fact, the first time I tried demoing this before filming, it just randomly game overed as I was like getting started in the first 30 seconds and never like said why. Actually, I think I technically beat it. I almost want to like take this version and hand it to Opus or Sonnet and say, "Hey, can you make this play better because the graphics are good but the game sucks." But when you combine how cheap it was to make this cuz like this was $5 I think to generate, that's pretty insane.
And if you combine that with like ultra fast, if that ever happens, suddenly you're going to be able to make a game in a few minutes on demand. We're actually now getting to that threshold where game development is about to flip upside down because of how models are finally understanding threedimensional space and the tooling necessary to do these types of things. It's happening. As per usual, I was not allowed to put the code I wrote with this model inside of T3 Code or other open- source projects during the testing window.
So, I had to use it exclusively on my internal projects like Lakebed as well as for auditing other work, which means I mostly use this for auditing other work. And I was very impressed. This is a real PR I was working on to fix a bug where my new little work tree like setup window that would appear in a new thread in T3 code would disappear if you left and came back. I had Claude code work on this, but this problem went pretty deep.
So I wanted to make sure that whatever solution I came up with was very very very well vetted. While I personally still do not trust this model to write the code I'm trying to land, I absolutely trust it to review things. Ignore the GBD6 soul there, just a placeholder. So, when I had 6.1 Soul look through this, it found real problems that were entirely missed by Fable and by Opus. First, it called out that follow-up messages can stay blocked after the agent starts, which is very annoying if you want to cue a message, and also that recovered setup progress was disappearing too early.
It figured all of these things out with a combination of reading the code and analyzing it as well as computer use. It was able to prevent me from merging a real regression in T3 code. So, I literally just copy pasted those things to Claude and then told it to take another look. Said better, but I'll still fix two small gaps. Remember, I can't use this model to code for this project at the time. So, I copy pasted that again over to Claude and it eventually got it good enough and then I finally merged.
But that is what I like this model for, and I cannot wait to push its limits for actually coding. Although, I will say from the code that I did have the misfortune of reading, it is harder to justify merging this code than it is for code from Opus. Normally, I would make you guys wait for the Opus versus Soul video or the Sonnet versus Soul video, but I'll just spoil the details now. I like using them in tandem because I find Soul to be way better at reviewing and digging into the details, but I find Opus a more pleasant collaborator and significantly better at actually implementing code without getting blocked constantly throughout its work.
And even now with the Rust rewrite of TypeScript, I find myself in a similar pattern where I have Opus 5.5 just going and going and going, making the codebase work and work well. And then I had Soul come in and do an audit. And this is the funniest part. Remember before I said that I had Soul and Astra working on that TS Rust port for effectively months. I told Soul to come in and it got it from 83.7% to 100% in under a day.
I was blown away by that that it had somehow unblocked the work that Astra and Soul were doing as well as 6.1 Soul. And I was absolutely blown away by that, that it had taken the work that 56 soul, 61 soul, and Astra had done over months and got it unblocked where it had been stuck for weeks and finished it. I was much more blown away when I had 61 soul take a look at that work and critique it. And what it brought up was that of the 1.8 million lines of code, 1.3 million were not being used.
The reason was because Opus concluded all of the code from all the other agents was useless slop that had no chance of being recovered and it chose to rewrite it from scratch itself in another crate. So on one hand, the only reason the code worked was Opus, but on the other hand, the only reason the slop was still around was also Opus. So I had to have this model come in and clean up the mess that other OpenAI models had made because Opus didn't even notice the mess was still there.
What I'm trying to say is this model absolutely has a place in your workflows. It could probably even be your default coding model and you wouldn't have too many issues with it, but I still find Opus to be a better collaborator overall. That said, I have almost no reason to use Sonnet anymore because this will effectively take its place. And you bet your butt the moment this model comes out, I'll be going and making adjustments inside of my cloud config because I already have it set up so that I can use soul inside of cloud code because I want to make sure opus knows this is the model to have review its work and investigate the things going on in the codebase.
This is a damn good model and I'm really happy to have it. I wish we had something bigger, smarter, and more capable overall. I was really hoping for something to truly dethrone Opus 55 as my daily driver. This isn't it, and I'm not planning on canceling any of my cloud subs as a result of this release, but I am planning on taking a lot more advantage of my codec subs in my day-to-day work. Admittedly, in cloud code, this is an awesome release and an unbelievable price for what you're getting.
But this does potentially mark the start of the end of the subsidization era. So, make sure you're subscribed so that you can be here when I cover all of that and more. God, I hope this doesn't get me cancelled online. I have no idea how others feel about this beyond like a handful of early access testers I've talked to. I legitimately don't know if people are going to love it or hate it or land somewhere between. I will know in a few hours, I guess, cuz it's a Yeah, it's 3:00 in the morning.
I am going to go to bed now. This is a tiring one. Hopefully, I did a good job. Let me know in the comments. And until next time, peace, nerds. God, I'm so dead.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.