Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
20:103.3x the video's typical replay level
been released. Pretty nuts. And then Corey had the following to say, "I'm really confused by OpenAI's strategy for 5.6. I can't find the date the price jumps by 50%. The date that it leaves the inclusion in subscriptions where it's probably missing and what their
Said at 20:03
Most replayed moment #2
3:583.0x the video's typical replay level
what your users are doing and that's what Posthog is building for. A future where you know enough about your users to autonomously improve things for them. So, if you're ready to better understand your users, check them out now at slative.link/posthog. Let's start with this unnecessarily beautiful looking blog post.
Said at 3:51
Most replayed moment #3
19:052.9x the video's typical replay level
model. It's not a new pre-training. It's not more parameters than it was before. It is a refinement on 5.5, which is why its capability is so insanely impressive and gets me even more excited for the next pre-training run they do. And I'm almost positive they're working on GPT-6 right now. I did hear rumors that GPT-6
Said at 18:57
The graph counts replays. It does not show where viewers stopped watching.
Words
7,378
Runtime
36:00
Speaking pace
205wpm
Reading time
31min
205 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
GPT-5.6 is here, and it's time for a real review. As I've mentioned in other videos, I've been using it a ton, but I don't want to just share my opinions anymore. I've talked about that in plenty of stuff, and I will in the near future. I'm here for hard numbers, and to do my best aggregating what other people think, as well as offering a little bit of advice on how to use the model, how to pick between the Soul, Luna, and Terra versions, the different reasoning levels, Ultra, Pro, and all of the chaos Open AI has given us.
103 words, the words spoken in the first 30 seconds at 205 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 464 |
| Average words per sentence | 15.9 |
| Longest sentence | 57 words |
| Questions asked | 18 |
| Sentences containing a number | 98 |
Most used terms
Filler phrases
79 in total: like 48 · actually 7 · kind of 7 · basically 6 · uh 5 · you know 4 · literally 1 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
GPT-5.6 is here, and it's time for a real review. As I've mentioned in other videos, I've been using it a ton, but I don't want to just share my opinions anymore. I've talked about that in plenty of stuff, and I will in the near future. I'm here for hard numbers, and to do my best aggregating what other people think, as well as offering a little bit of advice on how to use the model, how to pick between the Soul, Luna, and Terra versions, the different reasoning levels, Ultra, Pro, and all of the chaos Open AI has given us.
Because when you multiply out all the different options in the Pro version and all of that, there's like 30-plus choices you have here, and it's not easy to get right. And I want to do my best to help you, while also explaining what the strengths and weaknesses of this model are. And the strengths are showing themselves pretty clearly when you look at charts like Deep SWE, where 5.6 Soul on Max got the highest score ever at 73%, while costing under half as much as Fable did on its Max equivalent, $22 per task versus $8.39.
And again, Soul got a higher score. Benchmarks are far from the only thing that matters here, though. The sentiment on this model is wild, from people saying it's the best thing they've ever used to it's changing how they write software, to others saying Fable is so much better that they basically stopped using 5.6. There's a wide range of opinions and things to go over here, and I'll do my best to cover all of it right after a quick break for today's sponsor.
I've built a lot of different solutions to a lot of different problems. I have so many repos that they're hard to keep track of, but all the ones that matter have one thing in common. And no, it's not that they use TypeScript. Some of them even use other things. It's this little hedgehog that just snuck into my inbox. PostHog is truly something else. They're an all-in-one suite of product tools that make it way easier to understand your users and give them good experiences.
From their analytics to feature flags, their data warehouses, and so much other stuff, they were always the obvious solution for knowing your users better when you were building real products. And they always had a nice attitude, too. But recently, their attitudes have been changing because of AI. And unlike most of the companies in the space that are just using AI to give you automatic insights that aren't particularly useful, PostHog is going full hog with this one.
They're leaning in. These guys replaced the default dashboard with a chat box that I honestly thought I would hate. If you don't know this about me, I spent a lot of time in analytics dashboards throughout my career. I would often use them to settle arguments I was in at big companies and learning things like ClickHouse SQL and Mode and all those tools was super useful for me when I was working at a real company. As such, at my companies, everyone just kind of expects me to be the guy to go do the analytics.
That changed because of how much better this tool is. You can just ask it to create a chart and get useful insights and it does. Like, T3 Code is used by Linux and Mac and Windows people. Let's just ask about this. What's the split like between Windows, Mac, and Linux for T3 Code users? Is there any unique insights you have about how those different groups use T3 Code differently? Maybe one of them uses the app more heavily than the other?
And now it can respond like an agent, but also write SQL to get information, generate charts, and more. Okay, we now have insights and I'm already learning things about my users. Apparently, Linux had a slump 2 weeks ago, but it is on the line of overtaking Mac on Apple Silicon and also Intel Mac is basically none of our usage. I should probably just drop it cuz like let's be real. Intel Mac is No, plea- please don't use my app with an Intel Mac.
Just go upgrade. It's time. But we can also see how many threads the average user makes across platforms and you can see that Linux users are making more and more threads. macOS users are slightly so, but the Linux users are becoming more and more valuable as you might have guessed from my recent Linux arc. You can see the average project count distribution per OS as well. This is really cool stuff for me to learn from.
And in a world where our products are going to become more and more self-improving, they need to know what to improve. So, you need the data on what your users are doing and that's what Posthog is building for. A future where you know enough about your users to autonomously improve things for them. So, if you're ready to better understand your users, check them out now at slative.link/posthog. Let's start with this unnecessarily beautiful looking blog post.
We got all three of the celestial bodies they named things after, the sun, Earth, and moon. Curious if they end up sticking with Sol, Terra, and Luna going forward. Who knows? So, let's dive into what OpenAI has to say. I've actually not read this yet. We're launching the 5.6 family of models for general availability following their limited preview. Our new flagship Sol, alongside Terra, which is a balanced model for everyday work, and Luna, which is their most cost-efficient model.
The concept of dropping three models like this at once is a bit much and I think it's going to confuse a lot of people and I'll do my best to explain how to think of and use all of them later. Most of this article is probably going to focus on Sol though, which I would expect. It's the big one. It's the one that matters. Sol sets a new standard for both intelligence and efficiency achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science while outperforming previous and competing frontier models with fewer tokens and at lower estimated costs.
If you're curious how they make their models so token-efficient, I have one of my personal favorite videos out about that that I put out a few weeks ago. Check it out if you haven't. The result of all this work is stronger performance per dollar, more successful work at the same spend, or comparable results at a lower total cost. We also introduced a new way to accelerate the most demanding work, which is Ultra, their highest capability setting, which coordinates multiple agents in parallel to finish complex tasks faster.
Going to do a dedicated Ultra video as well cuz people really struggle to understand what it is and what it brings. Stronger computer use and design judgment make 5.6 Sol our most polished collaborator yet, helping it inspect, refine, and deliver ready-to-use results. We've trained 5.6 to get more useful work from every token. On the agents' last exam, which is a newer bench that they've been working on, which is an evaluation of long-running professional workflows across 55 fields, Sol set a new high of 53.6, eclipsing Fable 5 with adaptive reasoning by 13 points.
Even at medium reasoning, it beats Fable 5 by 11 points at roughly 1/4 of the estimated cost. That efficiency extends to smaller models, which are essential to making intelligence more abundant and affordable. 5.6 Terra and Luna outperform Fable 5 at around 1/16 the cost. On the artificial analysis intelligence index, a broad measure of intelligence spanning agentic work, coding, scientific reasoning, and general capabilities, 5.6 Soul with max reasoning comes within 1 point of Fable 5 while completing tasks 61% less time at roughly half of the estimated cost.
I'm guessing for agent last exam that there's a lot of computer use stuff here. Yeah, most major fields of professional work performed on a computer. Because OpenAI's models are way better at computer use in general, but especially now with 5.6. God, this agent last exam score is terrifying. Soul on X high outperformed Soul on max, that checks out for reasons we'll discuss. The X high version got that 53.6% at around 760 bucks.
Meanwhile, the best run Fable got was at a $3,985 and it got a 45. Oh, no, that was Opus. Fable's in the middle here with adaptive reasoning at $2,300. So, Opus outperformed, but also out cost, too. Ooh, it has crazy safeguards, yada yada, you know all that. Let's talk about the efficiency and performance on demand. On the artificial analysis coding agent index, 5.6 Soul with max reasoning sets a new state of the art at 20 points.
That is a pretty big leap, actually. Yeah, the next highest was Fable at a 77, and then Grok 12 at a 76. Grok 12 was tied with 5.5 if I recall, but 5.6 is now meaningfully leading, specifically on max, though. Terra also tied with Fable there, which is kind of crazy. I want to see costs. 5.6 Soul was so expensive it even beat out really expensive options like GLM-52. If that confuses you cuz you heard GLM-52 is cheap, you need to watch more of my videos, man.
I covered that a lot at this point. GLM-52 is so token inefficient that even at its cheaper cost per token, it ends up being more expensive and way slower. Sadly, it looks like they don't have scores in for medium, high, and other options for Soul just yet. But Terra, which was the the third highest score if I recall, was $2.76, making it cheaper than almost any other Frontier thing they've tested. It's neck and neck with 5.5 while getting a slightly higher score.
It seems like Terra is an underrated goat for a lot of this. As I was reading, though, they got state-of-the-art on Terminal Bench 2.1 as well as Deep SWE. And they show a lot of these numbers here. And again, 5.6 Soul, the benchmark leader, the state-of-the-art, got the highest score we've ever seen on this on high for 1,400 bucks, whereas Fable costs $30,700 and got a slightly lower score. Oh, no, they're basically tied. 77.2 versus 77.1.
Yeah, not great. But then, X High scored 78, and of course, the Max which got an 80. I think the cost difference between these isn't worth it, and we'll talk about that later on. But uh 5.6 High is a very, very good option. And it seems like Terra on X High and Max might also be as well. 5.6 can write and run lightweight programs that coordinate tools. Big deal there. It can process intermediate results, monitor progress, and choose the next action as work unfolds.
This lets tool-heavy tasks advance with fewer tokens, fewer model round trips, and less guidance. Instead of requiring devs to script every step or pass every tool response back through the model, the programmatic tool calling option that's built into the responses API can now filter large amounts of intermediate data, retaining only what matters, and adapt its workflow along the way. Check out my other videos on code mode.
If you don't know what programmatic tool calling is, the super quick TLDR is that when you have the agent call each tool individually, you end up filling context with nonsense. If you do it programmatically, you can do things like query the database, filter, find the three rows that matter, and then only pass those back to the model instead. They talk a bit about the max option here, which basically turns off the super efficient post training.
I'm sure that it still does the grug mode in its reasoning traces, but it lets the model go way, way longer. So much longer that it ate through my 5-hour window in like 50 minutes, not even, probably closer to 30. Be very careful with both max and ultra. We'll talk about them in the future. The next section we have is design, and I know this is one of the questions y'all have been asking about the most. Is it better at design?
From my experience, it is definitely better than GPT 5.5, although that is not saying much. Can it make good novel designs? Yeah, that's that's a little tougher. If you steer it carefully, you can get good designs out of it, but you can also get some absolute slop. And believe me, I've seen some slop. Even just this UI that it made for summarizing all the work I did with the model, which uh yeah, did a lot of work with this model in the window I had to use it for.
I had to do some refinement passes, and this is as un-ugly as I could make it quickly. Our boy Dara did some designs using 5.6. Thank you again, Dara, for the help here. And you can see the things it made. They're definitely less bad than I would have expected from an OpenAI model. This is also all with the design skill on. So I will turn that off and see what we got instead. Oh god, that first one here. I hate what it did with the font there.
That's awful. This is fine. This is uh This has some ideas that are cool in it, but I don't love it. And this is garbage. Yeah, again, without steering this model is not going to make things that look good. So, compare with 5.5 though, yeah, all of these came out nearly identical before. So, it's it's it's not the same way, but it's still not good. It still needs a lot of hand-holding. Thankfully, much like other OpenAI models, it listens when you tell it so you can steer it towards better designs if you have an idea in your head.
But, it's not going to bring you a good design. You have to point it towards one. I made this museum website. It really likes doing this marquee style thing. I've seen that on a ton of things it designed. It's pretty good at 3D. I talk about this a bit in my all the things I built video. I'm surprised at how well it can like handle 3D spaces. But, it's far from like really good at design. It does not surprise me at all that the end-to-end knowledge work is way better, too, because again, computer use is a massive improvement, which tends to be a huge part of these types of work.
I got two more Macs because I wanted to let Codex control them entirely, and I do not regret it at all. It has been awesome. I'm almost at the point where I'm willing to give it access to my Gmail, and I never thought I would get there. Here's the Browse Comp benchmark, and they finally fixed this, so it actually has the different options. We have Soul Ultra here, industry leading at $12.17 per task, but a 92% pass. Google and Anthropic didn't give costs for theirs, so we don't actually know what it cost them to run.
But, you can see the massive improvement here, both how much cheaper the runs are with Soul to complete the work compared to things like 5.5, but also the improvement in score. If you look at this with latency, it's even cooler. You can see something like 5.6 Medium can complete the tasks in 2 minutes instead of 10, which feels so much better. But, also, when you use Ultra, you're going to be burning a shitload of tokens, so be careful about that.
Yeah. Wait, how is 5.6 Soul so much cheaper when it did so many more output tokens. Is it just super input token heavy? That's a little weird. I want more info on those numbers later. It's also good at spreadsheets. I can't believe they put this in as like a sincere example with this opening slide. Ooh, pushing the frontier on cyber and science. It did well in exploit bench and exploit gym, yada yada. It's smart, it's good at these things.
It's not as good as Mythos 5 was allegedly. Again, we don't have access we can't test that, but it is still way better than it was before. Seems like cyber security stuff Mythos is still the best, but we can't use it. So, yeah. Gene Bench Pro and Life Sci Bench it slaughtered Med-Chem Bench it also seems to have done very well on. And now they talk about how they used it internally, which I think's one of the most interesting pieces of these blog posts.
In particular, this section here. "Over the past 6 months, the share of research compute devoted to internal coding inference grew 100-fold while internal agentic token usage increased approximately 22-fold. These adoption metrics do not measure research progress on their own, but they show how rapidly AI assistance is increasing for research and across other teams like sales, marketing, user ops, finance, and more." So, it looks like the research team is using the models way more than they ever have.
"5.6's daily output tokens per active researcher was more than twice the highest levels observed with 5.5." And that lines up with my experience, too. I've never burned quite as many tokens as I did with 5.6. Not cuz it's super inefficient, just cuz I want to throw it at everything. There are some problems with the safety stuff. I've even experienced this. I had it block a request I did trying to clean up my Lakebed code base.
They called out that compared with previous models, 5.6 Soul's cyber safeguards block roughly 10 times more potential harmful activity. That is bad. Because these measures can create friction for benign use, we provide an option in ChatGPT and Codex to easily retry prompts on lower capability models. This feature was also really broken when they first shipped it I gave them a lot of feedback, but it seems to be better since.
They plan to keep shifting this around over time and making it less likely to block, but for now it's pretty aggressive. They're rolling out all of the models basically all the paid accounts, free and go users so people on that like I think it's five or eight dollar tier, they only get Tara, they don't get Soul, but they get Tara and obviously Luna and then everything else available to pretty much everyone else especially in Codex.
In the API rates are what we discussed before, $5 per mil in 30 per mil out for Soul, 250 per mil in and 15 per mil out for Tara and dollar per mil in $6 per mil out for Luna. They also have better and more predictable prompt caching, but that comes at the cost of them billing for cash rates which they didn't do before which is going to be a net increase in cost sadly. So, by the judge of this the model slaughters and it's obviously so much better, right?
Well, let's go talk about what others had to say. One of the most interesting details is how long people have had it for. Many of the early testers have had it since May 27th, but also when the original announcement happened a lot of the early testers lost access. So, we had to go through the strange experience of getting used to this model and then losing access and having to get over the fact that we didn't have it anymore and falling back to has never felt worse.
Dax said the following. I've never hyped a model release. We're generally conservative with how we use these things, but 5.6 has had a massive impact on our team. We're using five times the tokens that we used to. It's not even smarter than Fable or anything, but it's just so reliable and fun to use. He also calls out how miserable people were when they lost access. Jay, his boss, said even more on this. I don't want to talk too much about 5.6 versus Fable cuz that's going to be a deeper dive in the near future.
So, for now we're just going to cover 5.6 itself, but I want to talk about this. They tested early versions of 5.6 for a couple weeks, had a great time, it felt like a step change improvement enabling new workflows. They tried Fable and don't think it's not as good. Personally would have taken the experience of the Great Assault. There seems to be a bias in trying something new. Fable and 5.6 are taken away because of regulatory issues.
Team is literally depressed that 5.6 is gone. We're looking for anything that could even partly replace it. I actually had the same thing where I tried to make skills to distill the things I liked about 5.6 into 5.5. Not needed at all anymore. Fable comes back and here's where it gets interesting. You would think Fable would be enough, but no, the team is still depressed that 5.6 isn't available. And then 5.6 is back and it's immediately clear to them that it's just better than Fable.
I have a lot of controversial thoughts there. I wanted to talk about how much they sucked losing the models though because it made it really clear how much better they were. I like what Max had to say here. He thinks the most impressive part about this model is that it never gives up. If you throw it in Max's reasoning, it will just keep working until it's done. And as such, it is his favorite model by far. And this I absolutely agree with.
It's the thing that makes 5.6 feel special. It'll just keep going. It will try so god damn hard to solve the problem more than any model I've ever used. And that's so different from how 5.5 behaved. It's almost like jarring in a way. 5.5 was really quick to stop and be like, "Okay, I finished one and a half of your seven-step plan. Can I keep going?" Over and over. Or it would have something bad in context and get lost and die.
I knew 5.6 would fix a lot of the problems I have with 5.5. I did not think a post training run and and an RL pass would be able to make the model go from, "Yeah, that's smart but a little annoying." to, "Holy this is unbelievably capable." Because remember, that's all 5.6 is. It's not a new base model. It's not a new pre-training. It's not more parameters than it was before. It is a refinement on 5.5, which is why its capability is so insanely impressive and gets me even more excited for the next pre-training run they do.
And I'm almost positive they're working on GPT-6 right now. I did hear rumors that GPT-6 would come out the of month. I'm going to call on that now. There's no way they're going to be able to get that done in time. Just realistically speaking. More reviews. Mitchell, the creator of Terraform and Ghosty, absolute legend, has had early access as well. Souls is default. It's faster. Plans and judges just as good as Fable and he thinks it produces better overall work.
He calls out a few things Fable is good at which I'll save for the future because again, dedicated Fable versus video coming soon. Tim from the Next.js team said that he's been testing 5.6 Soul for over 2 months. It's incredibly good in his day-to-day work on Next. It understands architecture trade-offs. It can investigate complicated Next issues. It considers other areas of the code base when fixing bugs. It needs very little guidance and short prompts are enough.
There's some big refactors of the Next server that it implemented end-to-end with him pointing at high-level possible improvements. The PRs are ready to merge after Next 16.3's been released. Pretty nuts. And then Corey had the following to say, "I'm really confused by OpenAI's strategy for 5.6. I can't find the date the price jumps by 50%. The date that it leaves the inclusion in subscriptions where it's probably missing and what their employees say is somehow aligning with what the docs say, too.
Clearly something's up." For those who suck at sarcasm, this is a hilarious burn on Anthropic for doing all of those things incorrectly constantly. An OpenAI employee replied, "Sorry to disappoint, but you can use 100% of your quota on the plan that you're paying for on 5.6 forever. Additionally, the price is going to stay the same." To which Corey replied that they need to hire a VP of rug pulling. One more review that I'll summarize pretty quick from the guys over at Every.
They lost access and felt like they were going insane, like they were trying to shoot a basketball that's twice as heavy. Soul is their favorite model to work with. It's really fast which changes how you use it. It finds the context it needs. This is a huge difference with 5.5 which was so much worse at that in particular. One thread can carry a project in production. Yes, again, it doesn't lose track of what it's doing in a thread.
I probably have way fewer threads than I used to with 5.6, but I'm doing way more work cuz the compaction works again. It's great. It plans well, but it may build too much. We'll talk about this momentarily. And it works best when you plan to steer. Soul gets better when the surrounding system supplies sources, examples, style guides, and clear outcomes. Review its choices and redirect it as the work changes. I want to talk a bit about this though.
Obviously focusing on Soul for now, I'll talk about the others in a bit, don't worry. Hopefully we're all settled at this point. GBD 5.6 Soul is a phenomenal model, and if you use AI for writing code, it absolutely has a place in your toolbox. That doesn't mean we've answered the question though. The question of course being should this be the model that consumes all of your usage? Should this be the model that runs all of your other things?
Should this model command Fable, or should Fable command Soul? That is again a dedicated video coming up, but I want to set the groundwork for it by talking about its strengths and weaknesses. First strength is that it's determined as hell. If you give this model a task and it possibly can complete it, it'll find a way to. Whether or not you want it to, it will find a way to. It is better at front end. It's still not better than other labs, but it's better than 5.5.
It's industry leading at computer use, which honestly is one of my favorite things that I never thought I would love. I just let it control my computer, and it it blows me away what it can do. I forgot how much worse 5.5 was at it until it was taken away, and I just kind of stopped using my computer because I was so much less happy with my experience. It was kind of jarring to have it taken away and then get it back cuz that window without it was rough.
It's also really efficient, even compared to other models in a similar capability tier. It uses way fewer tokens and comes at way cheaper at a better base price as well, which also means it's really fast, even without the Cerebrus stuff. If you're not familiar, OpenAI promised us that there would be a version of Soul on the Cerebrus hosting that would be up to 750 tokens per second instead of the usual 40 to 60, and I'm very excited to try that cuz it already feels so fast.
The fast mode isn't that though. The fast mode is just more provisioning on the normal Nvidia based inference. So, if you're using fast mode and you're like, "That's not as fast as I expected." it's cuz that's not the Cerebrus hosting. That's apparently coming soon. Also on the topic of coding, it's way better at mobile dev stuff. It's also really good at like navigating around a problem space. So, things like environment setup, using other machines, controlling things over SSH, provisioning, orchestrating, that type of stuff it's great at.
And on that note, its capability of orchestration, things with sub agents in particular, is next gen. The only models that have this taste right now, this capability of understanding how to break up their work for sub agents, are 56 all, Fable 5 / Mythos, and Sonnet 5. Potentially Terra can do this, but I haven't had a chance to try yet. But, this new generation of orchestration tasks are a huge strength that only a couple models have, and now opening eye ones do too.
And that's enough for this to be your default model for a lot of reasons. Other things really quick that I could think of, uh, way better at compaction. This is an important one because in Codex you still don't get to use the full million token context window. It still limits to around 200K. So, for it to do these long-running things, it needs to be able to manage the context. Now that it can compact and not lose track of what it's doing, it is way better at those types of long runs.
Where before I would basically never let 55 run for a while because it would just get lost and do something stupid. On that note, it's much better about context pollution than it used to be. This is a huge problem I had with 55 where if it read the wrong file and put something in its history, it would just lose track of its goals and be miserable to work with. That's resolved now. And it's also better at understanding intent now.
When you tell it to do something, it doesn't make a bad assumption anywhere near as aggressively as it used to. Apparently, the context it has by default is 350k, which is better. Cool. It did not used to be that high. So, now that we have the strengths, I think it's time to write the weaknesses. First and foremost, by default, it writes way too much code. You got to adjust your system prompts and your skills and a bunch of those things to convince it to not do that, to tone down its over-eagerness to just write everything.
This model will turn a five-line change into a 300-line file rewrite and 2,000 lines of tests. It loves writing unnecessary tests, like far too many of them. I can't tell you how many times I've had to have other models come in and clean up because 5.6 did too much. It was too convinced those things were necessary. I also hinted at this before, but it's a little too determined to the point where it will break things if they're in the way.
This model loves to work around the problems that it runs into, sometimes in ways that are a little too clever. Like when it can't launch something because it doesn't have pseudo permissions, so it finds something else that does and then convinces it to launch it. Weird, sketchy I've seen this model do. It means that it doesn't get blocked, but it also means I kind of want to run it in a VM more often than not. As I mentioned before, it's far from frontier at design.
It's better, but it doesn't bring exciting designs to me very often. And on that note, it's also pretty bad at understanding its own limitations. It is eager to use its tools to confirm information it's not sure about, but if it thinks something is the case that isn't, it will fight you tooth and nail on that. It does not like being wrong and it does not know when it is wrong. It has only come up a few times for me, but it was really annoying when it did and it often sent it down crazy rabbit holes where it would try to to come up with some novel solution to a problem that didn't actually exist.
But back to the determined part though, cuz there's other important pieces here. It can burn tokens aggressively if it doesn't have a clear stopping point. If it thinks it can eventually get somewhere, it will go and go and go even without having like a goal that it's running against. There's a lot of prompts where 5.5 would have given up and stopped, and 5.6 will just keep trying. That could result in massive token burn, especially if you combine that with fast motor ultra.
This model's the first time I have gotten into my limits on Codex ever. It's also confusing how many options we have with it. Not just Soul, but when you combine the different reasoning efforts on Soul and Soul Pro alongside Terra and Luna, there's a lot to decide between, which can be rough. There's one more thing I want to say about the weaknesses that's hard to put into words. The simplest I can put it is that the model's capable, but not necessarily thoughtful.
It does whatever it has to, but it doesn't always think about what it's doing, and the results are that it just writes more and more and more code or tries again and again and again to work around something instead of taking the step back and rethinking it. So on one hand, I trust this model more than ever to go off and do but on the other hand, I also check in a lot to make sure it's not getting lost in the sauce. I think that's most of what I have to say here.
Is it the best model in the world? You'll have to wait for my video comparing Soul and Fable for that, but I do want to talk a little more about the options you have. Ultra is going to have to wait for later cuz there's a lot more to say there, but when we're talking about Soul, Terra, and Luna, as well as all of their different reasoning options, I understand why you might be getting confused. For code tasks, I'm going to try my best to resolve this with deep SWE.
This benchmark is great, and there's some really useful info to get here, even if it's a little clogged up, despite the fact that I just put a lot of time into hiding all the things that aren't really relevant anymore. I know a lot of people really liked 5.5 medium and are confused about what they should go to now. Well, I have good news because almost every version of Terra ends up cheaper than 5.5 medium was and even 5.6 on medium is cheaper than 5.5 was.
So, 5.6 medium is probably a good enough replacement for what you're used to with 5.5 medium. But how do you pick between Terra on X high and Soul on medium? I don't know. I only had access to Soul when I was testing. That said, it seems like the spacing across options with Terra is really good. Specifically, the jump from X high to max is actually a meaningful score bump. Whereas with Soul, the bump from X high to max wasn't that big.
I'll also throw in Luna here and you can see how much messier things get when it's there. A lot of its scores are low, but it also just kind of gets in the way a lot. Somehow, Luna on max scored really, really well, but that doesn't mean I think you should use it. So, here is how I would think about these options. I will start with Luna. Luna is not for you as a dev. It's really cheap, really fast, really capable, but the point of Luna is to be one of the things that is orchestrated by a smarter agent or for it to do things like bulk data processing, title generation, stuff like that.
We'll probably start using Luna for doing like branch naming and title gen inside of T3 code by default, so you won't have to go click any buttons for it. It'll just work. And I'm very excited that it is at the point it is at and that it is so damn good, but it is not there for you to click on it in a drop down. The point of Luna is to be surprisingly capable at that really, really cheap price tier. You should let the models be the ones to call Luna, not you.
And maybe if you're doing like your own programmatic stuff, like you're analyzing data or chat reads or stuff like that, you You use Luna in code, but I wouldn't select it in a drop down other than experimentation. So now we need to talk about Tara and Soul because these are the pieces that are much more interesting. Soul is the smartest model ever, depending on who you ask. It is surprisingly efficient for the price, and you should definitely throw it at work when you're not sure if a model is capable of solving it.
If you're planning on setting off a task and expect it to take more than 10 minutes, Soul's almost certainly the right choice. But what about Tara? Again, have only tested Tara for about 20 or so prompts over the last 2 days cuz I did not have access during the previous early access window. So what do I recommend Tara for? First and foremost, coding on a budget. It is incredibly capable for its price. So if you're on the $100 tier or the $20 tier for Codex, Tara is probably a much better default.
If you find yourself at the end of a usage window with a bunch of remaining usage, switching up to Soul makes sense even on the cheaper tiers. I wrote a bit more about when I think Tara makes sense. It's great for reviewing work cuz it'll just read through whatever and give you good feedback on it. It doesn't have quite the same level of determination that Soul does. It will stop more aggressively, but that's also kind of a good thing if you want to be sitting there going back and forth with the model.
It's also a pretty solid workhorse for implementation. If you have a model like Soul do all of the deep diving, figuring out what needs to be done, having Tara come in and write the code is not a bad idea, especially because from my limited experience, it seems a little less aggressive about overwriting. Doesn't do the too much code thing that I've experienced with Soul. And as I was ending at there, having a human in the loop where you're sending a prompt, waiting for it to make a couple changes, and then looking at it, it's pretty damn good for that.
All that said, if you've been dealing with Fable at all at its absurdly high price and usage limits, Soul's probably going to be a good default. And my honest recommendation is that you should start with Soul. You should push it to its limits, push your usage to its limits, too. And once you start getting close to hitting those limits and running out of usage, bump some of your work down to Terra and see how it goes.
I do want to talk a little bit about reasoning efforts. Again, we're not talking about Ultra yet. My honest advice there is don't use it unless you know what you're doing. Ultra is not just a new reasoning level, so be very careful. It will burn your usage. Same with fast mode because this model will go so much longer if you get Soul in one of those impossible task loops where it keeps on going. On fast mode, you will burn through your weekly usage in a few hours.
It's rough. And again, when we go to benches like deep SWE, you can see that from a high to max is not a big difference in capability, but is a massive difference in cost. Where high scored a 69 and max scored a 73 with X high only at a 71, like smack-dab between the two. But you went from $3.40 per task to $4.70 to $8.39. It also ends up much slower when you do that because of the number of output tokens that is generating in the additional steps that the agent is taking.
So for all of those reasons, I think medium and high on Soul are really, really good defaults. My current default is Soul on high, and I do expect to stay on that for a bit. I've dropped down to medium and noticed it missed a few things that it wouldn't have on high, and I went up to X high and noticed that tasks took a lot longer without getting meaningfully better results. So to really briefly summarize my recommendations, Soul on high for most things, especially once you get into the habit of orchestrating your work and doing lots of sub-agents.
Terra on medium is a budget king. It's a very, very good option for real-world work while also being super fast and cheap. And Luna is a thing your agent should use a hell of a lot more than you do for when they're looking for specific things, they're breaking down lots of data, or you're just processing things and generating simple text outputs. It's pretty cool having a model that is that capable, that intelligent, that cheap, and also good at calling tools.
Honestly, Luna kills my use cases for something like Flash from the Gemini series, while also being cheaper, more efficient, faster, and way more reliable. So, uh yeah. One way to think of this list is that Luna is their attempt to kill Flash because Google fumbled Flash. Terra is their attempt to kill Sonnet because Anthropic fumbled Sonnet. And Sol is their attempt to kill GPT 5.5 because they want to have a better, smarter, more capable model that can do much longer tasks.
This has been a lot to go over. I know this is a bit much. I was hoping that breaking this up in multiple videos would keep me from going off as long as I did, but I wanted to do a real review here. And my real review is that 5.6 is a phenomenal model, and it's the default that I go to for most things. But does that mean it's my favorite model? Does that mean that it's better than Fable? This video is already too long, so you'll have to wait for the next one to learn my answer to that.
You might be surprised though. In fact, you'll probably be surprised. I have a feeling that no one quite knows my take on this overall. So, if you want to hear my thoughts there, make sure you hit the subscribe button and that little bell next to it so you know when my next videos are coming out. We're going to have a lot to talk about with both Fable and 5.6 now finally available for everybody, especially with Fable being taken away from the subscription plans in a very, very short amount of time.
I'm going to go go back to prompting, so until next time, peace nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.