Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
4:024.3x the video's typical replay level
model, the company that he just did a bunch of propaganda trying to get people to delete their subs and cancel them and Yet, here I am with not one, not two, but three accounts so I can get as much Fable as possible before it's taken out of the subscription plan. Yep. Yep.
Said at 3:56
Most replayed moment #2
13:502.6x the video's typical replay level
work is or what you're doing. Soul low fast for debugging and realworld like working on my system type stuff and finding files and just as an alternative terminal almost has been awesome. Why am I using the codec cli instead of cloud code for 56 soul low? Remember that
Said at 13:43
Most replayed moment #3
44:582.4x the video's typical replay level
them, so we can see how they feel about each other, too. So, I have these two reports that I had Fable and 56 generate based on my histories. I don't remember which one was generated by which. So, I'm going to use that CCF binding I mentioned before and see how quickly Soul could find it. And in under a
Said at 44:51
The graph counts replays. It does not show where viewers stopped watching.
Words
11,873
Runtime
57:39
Speaking pace
206wpm
Reading time
49min
206 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
I think it's finally time. You guys have been asking for my thoughts on Fable versus Soul for a while. And now that it's been long enough, I think I finally am ready to talk about it. I've had access to Soul for quite a bit since before Fable even initially dropped. But Fable I had much less access to. I didn't get anything early. I had to subscribe to use it just like everyone else did. I had it for the 3 days before we lost it and then had to wait weeks until we had it back again. And now I've
103 words, the words spoken in the first 30 seconds at 206 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 703 |
| Average words per sentence | 16.9 |
| Longest sentence | 100 words |
| Questions asked | 16 |
| Sentences containing a number | 129 |
Most used terms
Filler phrases
148 in total: like 101 · actually 31 · kind of 9 · literally 2 · I mean 1 · basically 1 · right? 1 · uh 1 · um 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
I think it's finally time. You guys have been asking for my thoughts on Fable versus Soul for a while. And now that it's been long enough, I think I finally am ready to talk about it. I've had access to Soul for quite a bit since before Fable even initially dropped. But Fable I had much less access to. I didn't get anything early. I had to subscribe to use it just like everyone else did. I had it for the 3 days before we lost it and then had to wait weeks until we had it back again.
And now I've had it for about a week and a half or so. And I am very impressed. I'll be real. But both of these models are absolutely incredible and I do hundreds of dollars in token utilization with both every single day. And I have never built as much as I have over the last couple weeks. There are massive differences in both of these models, strengths and weaknesses to each. And that's why we've been seeing such a weird split in the community over which one is better, more than ever before.
People like Matt said 56 is the best model they'd ever used, but then Fable came out and they stopped using 56 overnight. And then people like Jay, the CEO of Open Code, had almost the exact opposite where they had Fable access and people weren't that interested in it because they'd already used 56. And when 56 came back, it was immediately clear to them that it was way better than Fable. And they were stuck using it on Medium due to bugs in open code.
They didn't even know that they weren't using XI. Max is also on the side of Soul here, saying that he was tweeting about missing Fable, but he was actually missing Soul because again, all of us early access testers lost access for a week or so during the early access window when the government stuff all happened. And he vastly prefers Soul. Ben Davis, my podcast co-host, also massively prefers 56 and just uses it much more than Fable.
But then we have Ryan Carneato who is one of the most legendary web devs in history, the creator of SolidJS and an absolute legend in the web ecosystem who really really likes Fable and said that it's made him distrust every other model he's used before. It's not perfect, but when he has a nagging suspicion something won't work completely, Fable will identify it without even prompting. Every other model is thorough, but it will spew complete confidence even when pushed.
This is quite a range of things to consider, and I obviously have my own takes here. I'm sure that the comment section is already full of bias accusations and I haven't even started yet, but I want to give you guys my honest thoughts about this breakdown. And I'll give you a bit of a spoiler. My answer is Fable. I'll explain why I like Fable more, but also why they're both really good and why I use both every day after a quick break for today's sponsor.
Codex and Claude have both gotten incredibly powerful. So much so that I've spent a lot of time trying to train them how to use each other. It's beginning to feel like I wasted my time, though, because today's sponsor, Tracer, does it way better than any prompt could. These guys built a best-in-class desktop app that lets you work with your existing subscriptions for Codeex and Claude in one unified experience where they can work together.
Their artifact system makes it way easier to plan work and have multiple agents give feedback on those plans, massively increasing the likelihood that they're any good and that the implementation you get after is actually usable. They also think about work very differently where a given task can have lots of different chats and artifacts and agents and all the other things you need. Rather than grouping by threads and hoping that you can get all the work done before it's closed, they've re up the user experience to make it easier to get real large changes done.
Here you can see it's sending messages to different agents and running them in parallel in separate chats that I can open in their own tabs. This one's particularly fun because I was commanding it through Fable on a real claude code subscription and it spun up another thread that I could see here that's using GPD56 Terra on low to actually get these changes live. And if you're working on a team, you really need to check them out because their multiplayer workspaces are super cool.
It's one of the best gang prompting implementations I've ever seen, sharing a whole workspace with lots of different people. Man, this is so cool. I just watched it spin up another PR review chat that it's using sole because it should use a higherend model for something like this. It has good intuition about how to take advantage of the tools on your system. So, if you want the best without all of the effort, get it now at sid.link/tracer.
Okay, now that the shock has settled, I know, I know Theo prefers the anthropic model, the company that he just did a bunch of propaganda trying to get people to delete their subs and cancel them and Yet, here I am with not one, not two, but three accounts so I can get as much Fable as possible before it's taken out of the subscription plan. Yep. Yep. I like this model a lot. It's It's bad. It's bad how much I like it.
As I've mentioned in other videos, I use Claude Code and Codeexc across various machines, usually through T3 code as my control layer because it's easier than me to SSH to get the CLIs on all these different machines. But I also built custom tooling to track my token utilization across this suite. And what you'll see here is that my total for inference for the last 15 days from July 1st to the 15th is around $10,775.
My COEX usage is meaningfully higher than my Claude usage here where my Claude's only around 4500 and my Codex usage is around 6 grand, which is crazy. I only have one Codeex sub and I have three Claude subs that are constantly maxed out. So, if you're looking for the best value per dollar, no questions asked, go with Soul. Go with Codex. It's a way better value for what you get, but if you're willing to spend more to get some differences, many of which I consider really nice, Fable is a really, really compelling option.
And considering that I paid $500 or so for Claude here, because remember I have the two $200 subs and I just added that $100 sub to get in two weeks around $4500 of usage when I only paid $500 is a pretty good ratio, especially cuz that is like a week and a half with limits considered. my limits are all about to get reset and I should be able to get another three or 4K out of those before Fable's removed at which point I'm probably going to cancel those unless they give us Fable back or maybe Opus 5 is really good.
Who knows? Regardless, I'm getting a lot of value here. And I will also be real, most of the $6,000 that I used in Codeex is from one particular goal I am running because I know that I have usage limit resets that are going to expire. So, I'm working on letting 56 soul port all of Typescript go to Rust on another box constantly. This goal has been running now for a day and two hours, and I already ran it for like six plus days during the testing windows, so it's making real progress.
And I have it on ultra, so it burns usage. That one goal is roughly 3K or so of the $6,100, probably more. I haven't rerun the number since. So, I have an easy way to do this. Um, get rid of the buy harness. That'll tell me which machine it was. And I've only used ebook for this. Yeah, it looks like it was even more than that. It was $4,000 of that 6500 were from that one run cuz that's the only thing I've run on that computer this whole time.
Which means that if we were to fix those numbers and normalize them to cancel out that experimental goal I'm running, this would be closer to $1,500 of codeex and $4,500 of fable, which is a huge gap considering the fact that I use them roughly equally and for similar difficulty of tasks. Soul is just way cheaper. And there's a lot of layers to why it is so much cheaper. The biggest reason why Soul is cheaper isn't just cuz it's priced cheaper.
It's because it's so much more efficient with its reasoning. Fable will go off for long journeys to find answers. Not quite as long as things like Sonnet will, where Sonnet will do 70k tokens per task average on a bench like Artificial Analysis. Fable knocks it down all the way to 33k tokens per task, but Soul is 15K, less than half as many tokens for its max version. And when you go down to X high or high, you end up in the 11 and 10K range. a third as many tokens to solve the same task with roughly the same success.
Cursorbench showed similar differences here where sure fable can get much higher, especially now on cursor bench 3.2 where the best soul score is only a 67% and the best fable score is a 70%. But also soul only did $5 or so of usage and in fable on max did $17 of usage per task. That's a bigger than 3x gap in cost for around a 3% gain in score. Fable 5 high is still a pretty damn good value here, but even that score is worse than GBT on Max does.
And the reason again is the token utilization. Fable 5 Max used more tokens than any other model they have benched because it will just go and go and go until it is confident in its answer. But as such, it gets better answers. 56 is comically more efficient. 56 on high to 13k tokens versus 100,000 from fable on max. That is a huge gap. And that's not just a cost gap either. This also is a difference you'll feel when you use the models because when it has to reason much more, it takes a lot longer to get you an answer.
And if it does more steps, it has to wait and then send up the payload and then wait till it gets processed by a GPU and it can start responding, which ends up being way slower in quad code because they're not using a websocket based transit layer like Codeex is. There were points during testing where the websocket layer broke for Codeex and I had forgotten how much slower it was have to send up the whole history after every single tool call to start getting a response.
Fable still works that way inside of Cloud Code. And when you combine that with how much more it takes to get an answer, I've had so many times where I sent the same request to Fable and to Soul, and Soul had an answer in under five minutes, and Fable took 20 plus. There are lots of layers to that, but the token generation is the biggest by far. Only OpenAI models have that level of efficiency. Grock is the first series to even come close here.
But you should also ignore Grock's numbers here because it's kind of poisoned because they accidentally included a snapshot of the cursor codebase in training for Grock 4.5. You get the point though. It's way more token efficient. So for most work, at least in my opinion, it's worth asking soul first just to see if it can do it. And I've been surprised how often it can, especially when you play into Soul's strengths, most of which I would argue are persistence related.
If it can solve the problem by going at it for longer and trying multiple different things, Soul will figure it out. If the problem involves stuff like computer use, it also does exceptionally well. So to summarize the benefits so far, we have cost, time to complete, the computer use capabilities, which I really do believe are bestin-class, as well as the diligence and eagerness because the model will just go and go and go, and if it can find a solution, it almost always will.
Also on the time to complete note, this isn't out yet, but 56 Soul is going to be hosted on Cerebra soon at an absurd speed of around 750 tokens per second. For reference, the speeds we get right now are 40 to 50. That is a comically large improvement when you combine that with how efficient the model is. It's going to be able to answer complex questions in seconds instead of minutes or hours. And I'm really, really excited to see what that feels like.
But again, we don't have that now. I have not gotten early access or the ability to test any of that. It is promised and should hopefully happen in the not too distant future. There are other areas I've seen soul really shine that impressed me. Places like iOS development have genuinely surprised me. BY5 was terrible at iOS. It made apps that it like letter box the content and like put the whole window in a 4x3 screen in a native Swifty app.
I've never seen a model make such weird bad decisions around mobile before. And somehow 56 fixed all of that. I don't know what they did or how. Might be because they hired a bunch of iOS devs, but 56 Soul is so much better at iOS and it's been very nice to use for those types of things. It also has surprisingly good I don't know how to phrase other than like system understanding. Soul's ability to get how a thing works to like show it my network and all the devices on it.
For example, it's pretty good at intuiting what it is and how the parts come together. And this ends up being really useful for things like debugging stuff on my machine, which I do all the time using models. One of the things I still default to Soul for is for quick changes on my computer. I actually I use this all the time for like that. People are asking why not Luna, and I love you, but I I question if I just haven't been a good enough teacher.
Because for Luna to be as fast as Soul, it has to also be on low settings. Because again, the goal here is to do as few tokens to an answer as possible. Is Luna smart for how small it is? Absolutely. But it would take Luna on X high or max to get even close to soul on low and medium. And then it would need to generate way more tokens to keep up. And my goal here is the fastest time to a good enough answer, which soul on low will be faster for because it will get an answer comparable to something like Luna on X higher max.
And yes, those models are slightly faster, but you need to like think a little bit, guys. If something is 50% faster and it has to travel 10 times the distance, what's faster? That thing or the original thing that goes onetenth the distance and is 50% slower. Come on, use your brains, guys. Soul on low is obviously going to be faster than something like Luna on X higher max because it is on low, so it generates jack for tokens.
Again, for comparison here, soul on low generated 3K tokens per task on the artificial analysis intelligence index. Medium is 4K, high is 7K, max is 15K, and Fable is 33K. That is a tenth as many tokens as Fable 5. And I did put Luna Max in here. The second highest token generation of the whole line. It is almost 10 times more tokens than Soul on low for roughly the same intelligence. So again, I'm just using Soul. It is really really good and you can change the reasoning levels depending on how complex your work is or what you're doing.
Soul low fast for debugging and realworld like working on my system type stuff and finding files and just as an alternative terminal almost has been awesome. Why am I using the codec cli instead of cloud code for 56 soul low? Remember that thing earlier I said about the websocket connection. Whether you don't remember that or you do but you didn't understand it, I would highly recommend you go watch my video all about why the websocket connection layer in codeex is so so good when you don't have to wait for the whole context to be sent up and processed on every single tool call.
Things are way way faster and the things that make cloud code better than codeex for most code work are about code work not around on my system. So the benefit of cloud code here is very small and the time cost is exponential. It ends up being comically slower for the same reasons for these things. They make sense and you can come to the same conclusions if you listen to the words. If you have a faster one, bench it, test it, and then tweet me about it because you're probably wrong.
And once you go to collect the numbers, you'll realize you're wrong. So, with all this considered, it seems like Soul's a pretty damn good model, right? It is. It absolutely is. It's one of the best models I've ever used. and I'm probably going to send more prompts to it than any other model for a while because it is so useful in my day-to-day work. But there are some serious weaknesses and it's time we talk about those.
First, we need to talk about front end. I don't know what cope levels OpenAI are on. Especially during this imprompti models do not come up with good designs themselves. They just don't. They can't. The things they design are hideous and sinful and full of cards that make no sense and fonts that are too big and all sorts of weird stuff that just isn't the case with a lot of other models. We're at the point where there are openw weight models that can do better novel designs without much help than 56. 56 can comply with an existing design system reasonably well and you can steer it and tell it what to do and it will result in pretty good output, but it is not capable of making a good design from scratch itself.
You can beat a decent design out of them, but as Maria put it in chat here, it's not a good design experience and it's definitely not what we want from modern models like this. And I totally agree there. I still reach to Fable when I want like multiple different designs to decide between. Here's a real example of something I much preferred Fable for. I was trying to rethink the sidebar in T3 Code and I just wanted some inspiration, some new ideas to like start my thinking with and I had Fable do this.
It created a bunch of different fullon design mocks that I could look at in this HTML page that I hosted for all the different ideas that I had. This one is the status rail with a chunkier sidebar like identifier for each thread. Then we had the inbox style which ended up being really inspiring to me. Made me think a lot more deeply about how we would make the sidebar useful and present the info that mattered. Then we had attention tiers which ended up being one of my favorite concepts.
The idea of having a section that is stuff you're actually working on that matters and then a section underneath that doesn't as much. adaptive density where things get bigger and smaller based on the thing that is currently going on in them and then the ops grid which was just slightly easier to gro and this highlighting for the things that mattered which I hated but the fact that it gave me all of these different distinct designs with very little prompting beyond here's what it looks like I don't feel like the oneliner rows for each thread makes sense make them bigger and have more vertical space and figure out how we can convey the following pieces of information in the sidebars.
And no, I didn't use any special design skills or anything here. Once it made the HTML page, it did end up spinning a puppeteer to grab screenshots so that it could have those for referencing and looking at how it came out and also to add them to PRs in the future, which was cool that it could figure out how to do that cuz their computer uses garbage. So, they come up with fun hacks. But, it did all of this really well.
And I was able to take all of these designs, fudge them around a little bit, make some changes, make some suggestions, and eventually land on a style that I ended up really liking. And here's where I roughly ended up. Obviously, still very much a work in progress, but I was able to get this design after, again, a lot of back and forth and giving my suggestions, my thoughts, refining more and more until eventually landing on a design, at least a starting point design that I am quite fond of and can see myself polishing until we're ready to ship it.
I tried doing the same in parallel with Codex and I didn't even save the outputs because they were just they were garbage. I'm going to be real with you guys. It was nothing I would even want to put in front of you. That all said, Claude is still way worse at mobile stuff and I have found it's okay at making like mobile mocks in an HTML page, but it is so much worse at actually implementing those changes on mobile. And I often still end up reaching for 56 soul to help with those things.
Also, I just tried to get this updated sidebar in the mobile app for T3 Code and it ended up overriding my preview environment on my phone. It ended up breaking the app and the signing entirely for me. It it struggled and I have not had anywhere near that many problems when I'm using Soul for that type of work. I have one more benefit I want to share before I get to the opposite side of it because you need to understand the positive before the negative makes sense.
Soul will follow instructions really, really well. When you tell it to do a thing, it's going to do that thing and only that thing and it's going to go to the world's end and back in order to get that thing done. Believe me, it will find a way if there is one. This has some benefits. Like if you tell it to not do a thing, it won't. When you say, "I want this PR opened. Do not open it as a draft." It will open it and it won't be a draft.
But if you tell it to delete these three VMs and those VMs don't exist, it's going to go find three VMs and delete them. might not be the right ones, but it will try. If you had specified make sure they're the right VMs first, it'll make sure and then it won't do it. But this model will go and go until it figures out how to do the thing that it thinks you told it to. And it's generally pretty good at understanding what you told it.
It gets intent better than 55 did, but it's still not quite where Fable is at finding what you mean between the words that you said. So, while it does follow instructions really well, this is kind of where the negative comes in. It will destroy things to complete its goals. And there have been a number of examples of this. An amusing but painful example of this came from Matt Schumer, who is one of my favorite model reviewers and personalities in the AI dev space.
He has really good thoughts on models, and I often reference him in my work. him and I were talking because we felt like we were going crazy with how many people we trusted being all in on 56 when we very much preferred Fable. He actually had an OpenAI employee hit him up and say, "Hey, you should try out Ultramore and see if that gets closer to what you're looking for." And he did. And when he did it, it ended up running RMRF on one of his dev boxes and nuking his entire user directory in a bunch of inrogress code that Fable had written.
So, he lost a ton of stuff on this machine because he tried out Ultra, which I've said many times, be very careful with Ultra. This is what I mean with that. Like, it will try whatever it can in hopes of solving whatever problem you told it to. Even if its solution might be destructive, it's so desperate to solve the problem, it will screw you over in the process. I do want to be clear that OpenAI has been really awesome about this particular incident here where after it happened many OpenAI folks ended up reaching out working on solutions and even GDB the co-founder and president of OpenAI called him and was offering to help in whatever ways they could.
OpenAI cares and they hate seeing these things happen. They will fix these types of behaviors but they do exist now in ways they just didn't before. Sadly, others have experienced this too with Bruno here having Soul delete his whole production database. Never happened to him before with any other model ever. It's not safe. Notice how it has a goal going here. Goals and ultra tend to make the model behave more dangerously because those take its already overeager nature and turn it up to 11 or when the model is being told no solve it, no solve it.
It will try anything. Remember this because we will come back to this idea. But first, it's time to talk about fable and its benefits. The first massive benefit is its understanding of intent. The model just gets me better. It might do things it shouldn't once it understands me, but it's very quick to build that understanding, and I find I have to give it less to know it will get things right. If I give Soul and Fable both the same two to three sentence prompt, I trust Fable to know what I wanted better than Soul did.
If I asked for a change in the UI that actually requires some backend changes, so will get really confused and try to work around the need for the backend change, Babel will just go do it and let me know in its response at the end. Or maybe it'll ask a question to make sure that's what I actually wanted. Obviously, these models are just piles of bites that are being run by GPUs. They can't really understand anything.
But the feeling of it understanding is a real difference. And I do feel that difference with Fable. In that same sense, I find that it hallucinates way less. It does occasionally do it. And I've seen plenty of times in reasoning traces where it got something wrong. And it will say to itself, I just made that up. I'm going to do this other thing instead that was actually there, which has been really nice that it can catch itself.
And when you look at certain benchmarks like the omniscience benchmark from artificial analysis, you'll see that Fable has one of the best scores by far because this test isn't just testing whether the model is smart. It's testing how it behaves when it doesn't have the answer, when it doesn't know. And when Fable 5 doesn't know, it's very quick to admit it. When Soul doesn't know, it's very quick to work around it or make something up to unblock itself.
It feels like soul was trained so hard on unblocking itself that it often ends up behaving in a dumber way as a result. This next part's going to hurt me because I hate using this word, but but we do need to talk about taste a little bit because fable it has better taste. It makes decisions that feel more aligned with what I want. And obviously a lot of taste is how well does the model do the things you like in the way you like and it does them much better.
At least for me, this is small things like the copy that it writes. Although both are pretty cringe at writing, I still prefer Kimmy. K2's pros to anything these models have. And that's the non-thinking Kimmy. From K21 onwards, and especially K25, Kimmy's now too rled on being good at code and design and not as good at talking. So take that as you will. I really liked how Kimmy talked with K2. Fable and Soul still talk like AI, but Fable does it a little less cringe.
The more important piece with taste though is how it applies the decisions that it makes. Whether that is in design, whether that is architecture, whether that is the scaffolding for the app you're building or if it's just the code itself. It applies its decisions here in more tasteful ways and it makes more tasteful decisions in general. One of the things it does much much better is writing less code. I probably should have included this in the negative section when I was talking about my problems with Soul.
My biggest problem with it by far is that it writes way too much code. It is genuinely so frustrating how much code Soul will write in its attempts to solve a problem. It's at the point now where when I'm trying to solve a hard task and I actually want the result to merge, I will only use Soul for investigatory work or as a assigned worker implementing specific things. Fable is way better at writing the small subset of code that is actually needed to solve the problem.
This is a thing that I wish more benchmarks would measure, but sadly most benchmarks are measuring how successful was the model at getting the tests to pass or completing the task as assigned, not how good is the code and how likely is it to merge. And while there's a lot of things I don't like about the Frontier Code bench from the guys over at Cognition, I do think they got this part right. And a big part of why Opus and Fable scored so much better in this bench is because it's trying to measure how likely is the code to merge, not just how successful is the code at passing the test.
And as I have experienced many times now, and I'm sure many of you guys will see as well, Soul is quick to write 10,000 lines of code that aren't really needed, whereas Fable will write the 100 lines of code that are. And while I have opened more PRs with soul by quite a bit, I have merged way more of the fable ones than the soul ones. One more piece, it's really just one more word that I also feel bad putting in here, but I kind of have to.
Fable feels clever in a way that I haven't noticed before. I I don't know how to put it, but Fable comes up with solutions that I wouldn't have come up with in problems that I'm very experienced in. It has impressed me how often Fable 5 finds a good shortcut to a problem. That's why I like it being the model that steers other ones because when it gets back code and it notices the code could be simpler, it can correct it and tell another sub agent to go fix it.
And I love that about Fable. It's cleverness. I don't know where it comes from. This is a thing they trained it for, if this is just a thing that comes from it being such a big model, but I love when Fable helps me simplify problems. And I have used Fable to find and clean up and delete so much code. It's genuinely really nice. And with that, I will include this because I guess it does fit. It's a big model. 56 is the same size as 55 almost certainly.
It is a different post training. It is unbelievable that OpenAI can take a model that is as powerful but weird and bad at staying on task and quick to get its context up like 55 and through enough abuse of RL can somehow make the model perform even close to an allnew bigger pre-training like Fable. It's unbelievable. And the post-training team at OpenAI deserves the greatest praise and reward in the world for what they've achieved there.
It is truly unbelievable. They could take an existing smaller model and RL it into one of the best models ever made. The second best model ever made to be specific. It is insane what they did here. But Fable 5 is a new architecture. It's a new model that is bigger and has more information. It is capable of more as a result. That all said, bigger model doesn't come cost free. In fact, bigger model costs way more because the model is bigger.
It's more expensive to run. And because they haven't rldled it to be way more efficient in reasoning like OpenAI has done, they have made improvements here, but it's still not the same level, it ends up burning more tokens at a higher cost in generating them slower, resulting in a way slower endto-end experience, which is one of the things I like Soul for so much. It also has a habit of giving up slightly too early sometimes where it should take the next step, but it doesn't because it's often just trying to be clever and find a work.
It's like, what if I just don't do that thing? And I'm like, wait, you probably should go do that thing, though. This is part of its cleverness and its intelligence. And also that it doesn't have this absurdly overeagerness built in. But it is a little quicker to give up on a task than Soul is. And it also can sometimes fix it on the wrong thing and fall down rabbit holes. Not that Soul can't. They both do this to an extent.
I've had both make really bad suggestions many times. I have a video coming out soon about a bug that Fable and Soul caused that neither could solve because they both suck at performance analysis in Chrome, especially when it's GPU related. But yeah, Fable will give up early a little bit too often. And sometimes it's not actually Fable's fault because sometimes it gets rerouted. And the rerouting, it just sucks. I had some problems with the rerouting on 56 when it came out, but it seems to have been toned down quite a bit and I haven't experienced it at all since day one.
Whereas with Fable 5, it does happen a bit. It's not a lot. I haven't seen it for a few days, but Fable is still a little quick to reroute you to Opus for anything vaguely cryptography related, and I've seen people get rerouted multiple times in one request. Toucan here is an exanthropic employee, and he said the following. He had a Fable 5 chat fall back to Opus 48, which fell back to Sonnet 46, which then fell back to Haiku 45.
That sucks. And this is a uniquely anthropic problem. These are while soul will now reroute which sucks that there is no Frontier model that won't occasionally reroute you to something dumber. It happens a lot less and Fable is a little too aggressive with this to the point where it is uh frustrating to use at points. And one more big negative I don't want to put this other than exclusivity. The codec subscriptions are relatively generous.
They kind of just let you use codecs in whatever tools you prefer. The cloud subs are much less generous. basically forcing you into cloud code and their harnesses in other ways. They did officially support using your sub for local use cases using their agents SDK which is a closed source TypeScript SDK that provides cloud code with programmatic access. They had announced they were going to ban people from that but they seem to have walked that back announcing that they canled that plan change and they haven't put up their new plan just yet.
So thankfully tools like T3 Code are still supported. Other tools that fully change out the harness like pi or open code aren't yet. But since T3 code is calling those other harnesses because T3 code isn't a harness. It doesn't call the model directly. It calls the harness whether it's codec quad code open code or whatever else and then the model gets called through that through their official code. We are still good.
But if you want to use your Fable subscription for other things, you're kind of screwed. I use my subscription with 5.6 for a lot of different things. I use it for some automation work to keep track of all of my PRs. I use it inside of other harnesses like PI and Open Code all of the time. I use it inside of my Hermes agent and my Open Claw for just doing all sorts of tasks for my life. I know Ben, my podcast host and channel manager, uses his subscription to power a bunch of bots that we have in Discord for managing all of our channels and media and sponsors and whatnot.
All of those things are things we can do on the single $200 plan with Soul. And our insane inference just barely costs anything. We pay that one fee and get practically unlimited usage. Obviously, if you're using Facet Ultra, that doesn't feel unlimited anymore. Go check out my other video about all that. But for real day-to-day use, Solo has been great. And I really like it for those use cases like a chatbot or something inside of my Hermes agent because it's so much faster that it gets a good response quicker.
In Fable, you're just kind of sitting and waiting for. But you also have to pay full API price for Fable if you use it at those places because you can't get the subsidization that you get through the subscription outside of a quad code directly. Thankfully, T4 code is still supported. If you want a better UI, you can use that. But generally speaking, Fable 5 is very exclusive to your sub. At least it is for one or two more days because it's about to be removed.
They've delayed the removal three separate times, but it is still about to get removed. We are going to lose access to Fable inside of our subscription plans very soon. And even now where we do have it, Fable only can be used for half of your subscription limit. Your Fable limit will get hit 2x faster than everything else because it is just so much compute that it's killing them. My guess is that their plan is to put out Opus 5, which should be a smaller model with similar capability, and then convince people on the subscription plan to go use that instead.
But it's not going to have a lot of the benefits I like. It's not going to be that big model that has the same taste and is carefully applying each line of code instead of writing a ton of slop. It might be really good at orchestrating sub agents still. It probably will be, but I'm really not thinking it's going to be as capable as Fable. So when you combine the government ban of Fable with the absurdity of how it is coming into and out of subscriptions with the fact that you can even use the subscription outside of their services and then you look at the cost of what it looks like to actually run Fable in real world work, it starts to hurt.
That said, when I want my changes to land, Fable is the first model I trust to get the code so good first try that I don't feel bad just immediately filing the PR. Every other model generates meaningfully more slop code that needs a lot more handholding to get into a good state. Fable carefully, concisely applies the right diff so often that I reach for it every time I actually want to land my changes. And while Soul is an incredibly capable, powerful, efficient, just awesome to use model, it is kind of crazy that when I want to land code, there's only one of these two that I immediately reach for, and it is almost always fable.
I do have a way to compare these that I think makes these behavioral differences make much more sense, but I'm not the one who came up with it, so I want to find the original source and share it with y'all. It is this post by Peter Gustaf. I think it does a phenomenal job of explaining the difference. First, he says they're not easy models to compare, and I couldn't agree more. They are similarly capable and similarly intelligent, but the feeling of them is so different.
It's so different. His overall feel is that Fable is a wise owl who is very thoughtful and very well spoken. 56 Soul is like a Rottweiler who will grab the problem by the throat and not let go until it's done. Spot on. This is exactly how I feel about using these two models. Fable is a fundamentally smarter model even at low reasoning can be very insightful and it writes in a clear and compelling way. 56 soul on the other hand is extremely diligent.
Can give a list of eight things to do and you will be sure that all of them will be done. Babel feels more arrogant to him. He was trying to get it to build a new benchmark. 56 worked between 6 hours and 2 days. He tried several times to build it and it came back with a very thoroughly tested working bench. Fable came back within 40 minutes twice and the benchmark sounded smart, but it was ultimately a vibe based slop or bench.
And since it was Fable's vibes that was doing the judging, it decided that it was good to go. Of course, it was giving Fable 100% scores constantly. This is actually an interesting thing that happened to me during testing. When Opus 48 came out, I had just started testing 56 and I had 48 review some of the code that I wrote with 56 and it glazed the out of it. Opus48 was in awe of the work that 56 was doing. Then I got Fable about a week later and I had Fable review the same code and I was like, "Yeah, this is fine, but that's sloppy.
That's a mess. That's not necessary. I have no idea why that is there. We can clean this up." It was insane to see Fable much more capable of digging into the problems with Soul's code than Opus was. That said, when I had 56 review its own code in different threads, of course, it was very critical of it. Soul can be critical of Soul's work. It's surprisingly good at that. I could never get Fable to say anything negative about work that was done by Fable.
Even with new threads, new contexts, new everything, Fable almost feels like it can smell work it did and gives it the the craziest glaze fest it possibly can. I would never trust Fable to make a benchmark for these reasons cuz it is going to promote itself to the moon. Back to this comparison, I think this is really good. Fable still crafts better UIs from scratch. The flow of the app would probably be nicer if you start with it, but Fable will miss key things while Soul doesn't.
I actually just experienced this a few times when I was working earlier today. Actually, when I was doing this sidebar overhaul, I kept insisting to Fable that it's important we know which project is open both in the sidebar and in the thread. And then I would do one follow-up to make a change and it would make changes that removed the name of the project from both the sidebar and the thread. And then I would have to insist, no, make sure this is included.
And then it would go apply it really poorly and say, make it more tasteful. then another change and it would disappear again. So, I ended up having to like write a really strict message. Let me see if I can find it in my history. Actually, here I said, I'm concerned we may have made it too difficult in general to see what project you're actually in, which sucks because that's like priority zero. And I you not, it screwed up again.
It tried using color boxes, which were hideous. So, I told it try again. I don't like the grouping cuz this point it like tried grouping based on project which we had moved away from. that I really want the threads to be fluid and statically ordered from when they're open so there's less moving around on the screen. And then it again removed all indication of what project was open. There was no way to know and I was like so pissed off that I wrote a little essay.
Let's get a few things clear because you keep messing them up. I will not compromise on the following. One, indication of which project in sidebar threads. Two, indication of which project in the currently open thread thread. Three, using favicons where available for threads. And four, no grouping by project and sidebar. And I was so annoyed that I didn't trust it to do one pass. So I did my favorite trick. I asked it to go make five mocks somewhere else that I could go through in order to get this right.
And once I saw all of these mocks, I found some parts that were good. In fact, I think it was only one that was good. Is this two is great. Do that is all I said. And then it went and did it. And it came out really good. But like that's what I had to do. I had to brutalize this model into actually listening to me because it kept forgetting the details that mattered because it's trying too hard to assume intent and not to listen.
This is one of the biggest differences I have felt between the two. Soul doesn't read between the lines. It reads what you said and it does exactly that, even if that means destroying things in the process. Babel will read between the lines a lot, which is often good because it will come up with better solutions because it's trying to get to the intent of what you said, not the specific details of what you said. But that also means it could sometimes miss the details you're saying in its pursuit of intent.
It loses track of what you actually care about. But again, you can brutalize it into behaving by using capital letters and lists saying, "No, you have to do these things. No compromise." But like I had to do that. So, as Peter said here, Fable is way better at creating a UI from scratch and making things that look good, but 56 will actually complete the task when given the details. For writing, Fable's better hands down.
Soul feels quite different to align to what he wants to say or explain things to him simply, though he does think pro writes clear. I agree. You still can't use pro in code sadly, but it is a clear writer for sure. Robustness and reliability. This is where I think Soul wins for me hands down. Fable seems to do things of high quality, but he can never relax with it. almost always misses something with five six this almost never happens.
I agree. But like you have to understand with any given task there is the perfect 100% version where everything is solved and nothing extra is done. There's the 80% version where it is the simplified thing that gets most of the problem solved but misses a few edges. And then there is the overdone extra version where you do 300% more code, a shitload of smoke tests, and all these other things that don't really matter.
Babel skews between like 80 and 110% from my experience where it just barely misses a few things or does a little too much extra. Soul's baseline is like 150%. It always does a little too much and it sometimes does way too much. I've been playing with skills and plugins and things to try and get Soul to not do that. Even just asking it at the end of every change, how can we simplify this? How can we make this PR as minimal and elegant as possible while still achieving our goals? and it will usually find a few things, but honestly, what I do nowadays is I have Fable go in and clean up instead.
Peter has a list of other things he prefers with Soul. Things like video editing, which it actually can kind of do now. I don't really agree. I've seen it do some surprisingly good stuff, like creating the XML files that we use in Final Cut. I know that Jeff, my editor phase, has had a lot of fun with that lately, but it sometimes creates corrupted files, too, and then you have to go do it from scratch anyways. It also has no taste with video editing at all.
It can just mark things which does make the editing process by hand easier. I'm interested to actually edit video though. The computer use however is unbelievable. I still cannot fathom how crazy the level up is here and how capable 56 is. I often will leave my personal computer with codeex open. Sorry, with chat GPT in codeex mode open just so that I can get things done on my computer when I'm on my phone remotely. It's so nice to have that deep computer use capability where I can like go into my browser and do things, find a file in my folders and drag it over somewhere else.
Like it can complete work on my machine without me there. It's been really nice. It's also really good at sub agents. I will say Fable's also quite good at them too, especially workflows. It's good at adhering to existing code patterns, which is really nice as well. I would not trust 56 to create a new project from scratch, but it will follow the patterns in the codebase better overall too much so sometimes where it'll make things bigger than it needs to, but it does really just fit into your codebase well.
It doesn't get mad at you when the patterns aren't what it likes. Fable will talk a lot of about your code. 56 will just perfectly match the patterns as they stand. It still leans towards writing TypeScript like it's a Python dev. I've beat that out of it with a stick, but it loves oneline functions that are just typ casting and Not my favorite, but you get the idea. It's getting better at research overall. It has some bad patterns, like it's too tactical, but it's more steerable to be a good researcher.
This I believe, especially with pro, but I have found that the insights I get from data or just like giving the model access to my post hog, I get better insights out of Fable than 56 by quite a bit. multi-day runs. Soul is really, really good at. I wish I could say I tested Fable for these types of long multi-day runs, too, but that's just like absurdly expensive. Not realistic, if I'm being honest with you guys. Now, we have the token efficiency side because it's so much more token efficient and faster.
Yep. Said the same. Very good. The multi-day runs thing is crazy though, especially with the compaction. Buy four's compaction was great. Buy five broke it. So, when you had long threads, they would just lose track of what they're doing after it hit that point. 56 fully fixed it and I trust it to just go and go. I don't even look at the token usage anymore because it will compact and figure it out and then keep going and be fine.
It's great. The downside is that you can feel Fable is naturally smarter. He had some brief moments of 56 where he was getting it to make fairly simple changes in eight turns. It just got stuck in dumb streams that was hard to get out of. So, while it's not AGI, don't get too carried away by the hype. I know this video has already been a bit long, but the wrap-up is where it's going to get really fun. I'm going to give you my exact methodology for how I pick which one to use and when, as well as the comparisons, the models generated of my actual long histories with both of them, so we can see how they feel about each other, too.
So, I have these two reports that I had Fable and 56 generate based on my histories. I don't remember which one was generated by which. So, I'm going to use that CCF binding I mentioned before and see how quickly Soul could find it. And in under a minute and a half, probably it found both. So the OD one is quad code and the other one is codeex. Interesting. The pretty one was generated by codeex in 56 and the ugly one was believe it or not cla code.
So let's read through these super quick and then I will give you my mental model for how I pick which one to use and what I think you should subscribe to. Fable versus 56 soul. The fleet view. This again I use these across all my different machines. There were many of them and it collected all of the histories from all of them and analyzed it. According to Fable, the TLDDR is that 56 Soul is a distributed construction crew.
Fable is a scarce senior engineer at two desks. I did a lot more work with 56 Soul. As I mentioned, they did not restrict me during early access testing, which they're probably going to in the future with how heavily I abused it. But as it said here, 87% of Soul's turns were machine delegated or scripted. So Soul was spinning up other sub agents. And 90% of the output I did was in a single 9-day swarm campaign where I was running a bunch of crazy rewrites on different machines.
And also Fable was not around at that point. Fable lives where Theo physically sits. The MacBook and my framework desktop had 87% of the output. It does the planning before the swarms and then judges after. And on the six days both models were active, Fable outproduced Soul by 1.3x roughly. So again, on the days I'm actually using the models and I have access to both, I definitely lean Fable. It didn't really get what I was asking for though.
I wanted to compare more how the models behave. Thankfully, Soul understood that because it takes me more literally. So, Soul wrote up a report about how they behave. It said the Fable leaned collaborator and 56 leaned system. Both did direct work in orchestration. Fable was more often the frontline collaborator and 56 was more often a persistent orchestration layer. 56 produced far more recorded token volume, but almost all of the gap came from a 9-day automationheavy campaign.
Fable appears in smaller direct product building bursts. I love that here it rendered these charts and you can clearly see when I had to stop using each model and different machines had different stories. Fable was mostly on my MacBook. 56 was on the Zbook which is the computer I had doing those crazy rewrites I showed earlier. I gave them different jobs where Fable was a frontline collaborator and 56 was an orchestration layer.
On comparable dates, the Gulf nearly disappears. What this report cannot tell you is quality. It's not actually measuring the outputs. Yeah, still cool to see it do a nice fancy report like this, but I wish it did more like for like comparisons between Fable and 56. I might ask it to do another set of reports, but I don't want to wait to do all of that now. I'll include that in future videos if I get anything good because right now I want to use your time efficiently and give you the most important thing, when to use each, but more importantly, which to subscribe to.
I have used a lot of all of these subs and I've used a lot of both of these models which is why my response here might surprise you. I prefer Fable and I prefer Fable for a lot of things. Specifically, I use Fable if I want the change to land. I use Fable if I want a good design. I use Fable if the problem is hard or complex in weird ways. I use Fable to figure out what I want. And I use Fable to verify and simplify work from other things, even itself.
And honestly, the first one should be enough to convince you that Fable is an incredible model. So, what the hell do I use 56 for then? Well, I use 56 first almost always because it's so much cheaper. It's practically free and if it turns out 56 can solve it and it can save me all the token and usage burn from Fable, it's worth it. So, I always start with 56 for the work that I would do in either just to see where it goes.
I use it for long running work all the time because it's so much better at staying on task for these runs that go for hours and hours. It's genuinely really impressive. I've had it run for days at a time and actually make working outputs, which is still crazy to me conceptually, but it works. It does it. It blows me away. The fact that I can have it on a goal and it goes for two and a half days and then stops and says, "Okay, I did it.
Here's what I built. Here's what I want you to go test. And if depending on how this goes, I can continue." Really incredible. and it sends me things that work after days. And since it's so much cheaper, this is actually somewhat reasonable, unlike on other models. I absolutely adore 56 for computer use. I use it way more than I ever thought I would. I did personally just didn't like computer use before to the point where I would like talk on it, and I've intentionally not invested in companies building in that space.
I regret that now because the models, at least the OpenAI ones, have gotten really, really good at it, and it blows me away what they're capable of. I have lots of examples of that in my 56 review video. You should definitely check that if you haven't. It's melted my brain. On the note of like general use, 56 in tools like Hermes and OpenClaw is incredible. Both because you get the crazy subsidization with the subscription, but also because it takes intent literally and responds quickly.
So, I really like it for those use cases, too. As well as like quick fixes in general. Whether the quick fix is something on my computer, in my system maintenance, or if the fix is something that is in a codebase that should be really quick. I love using 56 for oneoff things where I'm going to watch the terminal from the start to the finish. If I'm going to watch the terminal the whole time or I'm going to watch it in T3 code the whole time, 56 fast low is one of my favorite things to use.
But if I'm going to open it in the background and then come back in 30 minutes to an hour and want to change, I can hit merge on. I still obviously prefer Fable. I also really like delegating to 56 to have another model like Fable break up work into smaller pieces and then tell 56 to go complete a part of it. It's a way to save meaningful amounts of usage while also getting really good results. And I've loved 56 for this and having Fable be an orchestrator.
We showed this a lot in other content. Definitely check that all out. But I do really like 56 as a thing that is called on to do stuff. The simplest way to put this is I kind of treat Fable like a really smart co-orker. And just because your co-orker is smart doesn't mean they can't do stupid things or have bad opinions sometimes. It just means that they're generally pretty smart and get Where with 56, 56 is a really strong tool.
It is a tool that I love using and I wield for all sorts of different stuff, but they are very very different. Which is why it is crazy they bench so closely and they behave so to speak on benches as similarly as they do because using these models couldn't feel more different when I do it every day. And this is where the the hard part comes gun to my head. I can only pick one. What do I pick? This is actually really hard.
I have thought about this so much both online and offline and it's difficult and I'm going to give a weird answer here. If I have to pick just between these two things as what I use for my work, my code and everything daytoday, Fable wins. But if I look at my own life and the things around me, this is reductive because I have Julius. I have one of the best engineers ever working with me, building awesome who has really good taste, who builds really competent stuff, who puts up code that can absolutely land.
If I didn't have Julius on my team, I don't know how I would function without Fable. But because I have Julius and I can trust him to do that top level of work, I think I would struggle more without 56 because the way I use Fable is honestly similar to the way I text Julius. I give a vague idea of a problem I have and a rough direction I want to go in, and Fable can mostly figure it out, as can Julius, even better than Fable can. most of the time.
But when I just want to like do some random on my computer or fix something on another machine in my network or make a quick one-off fix, the type of thing that I wouldn't want to text to my coworker Julius, then I much prefer 56. And I am sorry chat, you guys don't have access to Julius. He is mine. And I would recommend you guys give him a follow if you want to see what building at this scale and at these levels looks like.
He's unbelievably talented. And since I have a team with Julius and other similarly capable engineers coming on too, my need for a autonomous model that can clean up the code itself and land the change is less great than it was without those things. If I was a solo dev or a solo founder, I would rely on Fable for everything I do. But since I'm an engineer at a company with other engineers who are incredibly capable, too, 56 feels more like a hammer in my toolbox that I reach for all the time.
Fable feels like the contractor that I hire to come in and use their tools themselves. So, with that, which sub should you buy? I will say very confidently, every dev who can afford it should have at least the $100 tier for Codeex. It is just so absurdly generous, especially with the constant resets that we've been getting as of late. You should have that $100 tier. You should find ways to max it out that bring actual value.
Whether that is a Hermes agent or a cron job that indexes your PRs every day or you start going for bolder work, whatever it is, you should not make a decision about which subscription to use until you have maxed out a $100 tier on codecs. And then you could sit and look at the work you've done and decide, is this useful? Do I like the outputs? Can I see myself using this even more? If so, upgrade that to the $200 codeex.
If not, go open up a $100 Claude code and see how you like the results from Fable. I'll be honest though, I would much more quickly cancel my Claude code sub than my codec sub, as I've made clear in the past, because I don't know when and how I can use that sub and all the things I can use it for. If you're not hitting the limits on your codec sub, you're not getting creative enough. And you can do some crazy things with that subscription, which is why I love mine so much and why I've been getting crazier at hitting those rate limits as of recent because I'm using it for all sorts of different things.
It's just so valuable a tool to have in your toolbox that I think you should have that first and get used to using it first before you look at alternatives to toolboxes. And of course, there's the problem that Fable is going to be removed from the subscription tier soon. It might not even be there by the time you're watching this video. Will it come back eventually? Yeah, who knows when though? So for all of these reasons, I think everyone should start with the codec subscription to use it to push its limits to get as much out of it as they can.
And if you hate the codec and you don't like the chat GBT app, I understand. I recommend T3 code for an alternative to the chat GBT app. Or you can watch my video all about using your codec sub include code in order to get the benefits of that harness. There are some rough edges there, but I found it to be a very pleasant experience. And with that, I think I'm finally done with this topic. I have my answer to the question everyone keeps asking me, which is which model is smarter and better, and which one do I prefer for different things.
I've said all of it now across all this different content, but this video in particular, I hope it is clear why and how I choose each model, and the simplest way I can put it is I use 56 most of the time, and if it doesn't do what I want it to, I go over to Fable to get a good change that I can actually get merged. It's truly incredible to have two options of this absurd level of quality. And for them to be so different is just fascinating and fun.
It's never been more enjoyable to play with models the way it is today. And I've had so much fun experimenting with both Fable 5 and GPT56 Soul that I recommend you do the same. Don't just blindly follow my recommendations here. If you take what I did and then copy paste it into your setup because obviously I know it's right. I'm the influencer. Don't take lessons from what I did, learn from how I think about these things and then go try them yourself.
Maybe you agree fully with what I did and what I say. Maybe you end up going in an entirely different direction and you feel fundamentally different than I do about all of this. I can't know. That's not my job. That's your job. Now, experimenting with these things is one of the best value ads you can have as an engineer. And then bringing that knowledge to your teams or other companies will help you level up more and more.
So, while this video is definitive for me, I hope that it's just the start for you. Let this open up the journey to go play with these things more. Experiment with the models and the tools and the harnesses you use them in. change the types of tasks and work you use them for and see how far you can push them. I have a feeling you'll be surprised just like I was. That's all I have to say about this. So until next time, these nerds
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.