Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Theo - t3․gg's most watched videos.
Most replayed moment at 2:36
5.6x that video's typical replay level
Well, you have nothing to worry about cuz the first million users are free. Get yourself enterprise ready at soidiv.link/workos. I'm very excited to read into what Linear is cooking here. You know they're cooking something different cuz this is the only not dark mode page I've ever seen Linear ship. They actually took
Said at 2:29
Most replayed moment at 25:09
3.4x that video's typical replay level
just set up now. It's really nice when you do need to do things in the GUI or format the machine, that type of thing. It's time to show you guys how I actually do work using this setup. NPX T3 at nightly serve. I do have these host commands to make it work better
Said at 25:02
Most replayed moment at 2:53
29.0x that video's typical replay level
of different places in order to pull it together. Setup couldn't be easier. You click start, you add a new database, you get a connection string and now you're good to go with a real Postgress database with all the power of ClickHouse behind it. Stop compromising on your database today at soy.link/clickhouse. The best
Said at 2:45
The graph counts replays. It does not show where viewers stopped watching.
Words
7,572
Runtime
36:35
Speaking pace
207wpm
Reading time
32min
207 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
There are so many incredible models available for us to use for our work every day. Whether you're using expensive, best-in-class stuff like Fable 5 or surprisingly cheap and effective stuff like Deep Seek V4 Flash, it's kind of hard to go wrong. But what if you are? Wouldn't it be nice to know which models are best and worst and why? I know a lot of y'all are asking for this. And as much as I try to include this context in videos, seems like you all just want me to rank them all and put a list up that shows which ones
104 words, the words spoken in the first 30 seconds at 207 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 471 |
| Average words per sentence | 16.1 |
| Longest sentence | 124 words |
| Questions asked | 14 |
| Sentences containing a number | 124 |
Most used terms
Filler phrases
96 in total: like 75 · actually 14 · kind of 3 · you know 3 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
There are so many incredible models available for us to use for our work every day. Whether you're using expensive, best-in-class stuff like Fable 5 or surprisingly cheap and effective stuff like Deep Seek V4 Flash, it's kind of hard to go wrong. But what if you are? Wouldn't it be nice to know which models are best and worst and why? I know a lot of y'all are asking for this. And as much as I try to include this context in videos, seems like you all just want me to rank them all and put a list up that shows which ones are best and which ones are worst.
Like a tier list. You know, the classic way of categorizing things. Wouldn't it be nice to just have all the models in one place? I'm sure that this isn't going to get me in trouble online. I'm sure everyone will fully agree with my rankings here. To be very frank, I don't think a tier list is the best way to compare models right now cuz there's so many different axes to compare on. From different tasks that they might be good or bad at to the costs and how those costs actually pan out for real world usage, performance, speed, all of these different things matter a lot.
But I'm going to throw that all aside so we can do the silly fun thing, which is a modern tier list of all the models available to us today. I just whipped this tier list together with 56 Soul. So, I think it is a really good model to start with as our basis for how things will be ranked going forward. And I'll let you know where I'm going to put it right after a quick break for today's sponsor. I got a weird question for you.
When's the last time you raw dogged a model? I know this sounds insane, but hear me out. When I use tools like Claude Code and Codeex, I am blown away with what these frontier models can do. But when I hit the API directly, I'm not that impressed. There's one particular thing that seems to be holding the models back is that they don't have access to the internet. And while Cloud Code and Codeex do provide a meaningful amount of access, which is why it feels so much smarter, other tools just don't.
And when I build my own, it just the model feels dumber because it turns out it's getting the majority of its context from the web, at least the parts it can access, if it has access at all. Wouldn't it be magical if you could give the models access to the entire web? It would make them so much smarter and more capable if they had access to all of the data available. If only they had a browser that was built for agents that scaled infinitely and gave them access to everything in whatever format they wanted.
I don't know, something like today's sponsor, browser base, cuz these guys built the best possible browser for your agents. Whether your agents need an equivalent of Google, if they need to fetch specific content from URLs they already know, or if they need a real browser where they can sign in and click and do things on a user's behalf. All of that is covered and more with browserbase. It's rare that I can say Microsoft and DeepMind are both users of the same thing, but that really shows you how strong browser base is.
That's why everyone from Verscell to Octa to Anthropic and Stripe are all going to their conference called Navigate in September. If I know anything about Browserbase, it's that they know how to party and have a good time. highly recommend this if you're down and you're in the area. Seriously though, your agents are restricted when they don't have context and the web is the context they need. Give it to them today at soy.link/browserbase.
I promised you I would start with soul and I still want to. But I'm going to be real. It's a bit hard for me because I'm somewhere between A and S tier. If I hadn't struggled the way I did setting this up initially, I would probably be putting it in S tier. But I'll just show you what I went through when I was doing this. I have a screenshot here of the logos that it did before. It didn't find logos for XAI for composer or Grock.
It put in the question mark I had with Composer 2X. I wasn't sure if I wanted to include it or not. It got the OpenAI logo wrong by putting a white logo over a white background. And it just didn't find Kimmy or GLM logos at all. I had to tell it with a screenshot, this sucks. Go find them. And it was able to do it. So, like it can fix things if you tell it what's wrong, but it's also really dumb and gets things wrong pretty often.
So, all of that said, I think a tier is a really good fit because this model is unbelievably intelligent. It is capable of things I never thought AI would ever be able to do. I have built things I did not think would be buildable, much less with AI with it. I have a whole rewrite of the T3 Code mobile app in Swift UI that I did entirely with 56 soul in a single thread. It is unbelievable what it can do. If you compare it with 55, it is absurdly better.
They fixed all my problems with 55 and more and made it way better at these types of longrunning hard tasks. If you haven't tried 56 soul, I highly recommend it. It is my default model I use for most things most of the time. But it still is not the smartest model I use. It's still not my favorite for writing important code that I actually hope to merge. You can probably guess which one is that, but we won't get there just yet.
Let's start with some models that are a little more confusing. I think Terra is actually a really good one to do next. Terra is an interesting model because it came out alongside Soul and Luna. I don't think they should have released all three at once. It made it hard to know what to use each for. 56 Soul is $30 per million tokens out. 56 Terra is $15 per mill out. I think they lowered it slightly. And 56 Luna was, I think, $2 per mill out, but they dropped the price a ton.
Sorry, it might have been way more than that, but the the price out now is $120 per mill out, which is kind of crazy. Terara is down to 12 bucks per mill out instead of the original 15. All that said, Terra's not as token efficient as Soul. So, even though it is cheaper per token, it often ends up using more tokens. And if I'm being real with you guys, I've struggled a lot to find a place that I think Terra makes sense in my workflows.
It's not a bad model. It's just not particularly useful to me right now. And for that, I think C tier probably makes sense for it. Okay, to be fair, it did cost a bit under half as much as Soul on the pricing charts for artificial analysis, meaning that for each task, it was about half as expensive. And that lines up with Cursor Bench as well, which saw 56 Soul Max at about $5.70 per task in Terra is $2.30 per task. And the score difference isn't as big as I might have expected either.
There are a few places where Terramax beats out soul per dollar, but it's relatively continuous here. It's just it fits in such a weird place because a lot of these numbers can be gotten for much cheaper with Luna. It only beats out soul at like weird price points. And even here, I'd rather just use soul on high because it will be so much faster because it has to generate so many fewer tokens. So yeah, while it measures well and for my usage is not bad, I have never chosen Terra for anything and I would be surprised if many people do.
On Deep Suite, you can see this even more so with Luna being hilariously cheap. Terra not being quite as cheap and getting close to 56 soul numbers per dollar. But again, like Soul comes out in many places here slightly more expensive and slightly smarter, but it's close enough that again, just it doesn't feel like Terra makes any sense. I'd rather use Luna Max if I'm really trying to cost save or soul on low if I'm trying to get responses faster.
Terra is an in between that makes sense on a pricing chart, but just doesn't really make sense in reality to me. On that note, Luna, I'm putting Luna in A. And not because I think it's a really good model you should be using for code tasks. I'm putting Luna in A because it's the first cheap model that we've had in a while that is fast, smart, and good for a bunch of random stuff. I use Luna all the time. In fact, it's probably my most used model by sheer calls to it.
Not because I'm doing code with it, but I have it doing all sorts of other stuff with my coding tools. For example, in T3 code, we use Luna by default as the model that does title generation and status stuff. I have some work I'm doing right now where Luna will categorize all of the stuff going on in your T3 code and manage your threads for you. And Luna's smart enough and agentic enough that it can like make tool calls, investigate things, pull stuff from GitHub, read content other places, and make decent enough decisions.
I wouldn't trust it with anything that isn't reversible, but as a way to like generate context, to summarize data, to generate titles, all these types of things that I often find myself doing. Luna is incredible at turning random context into a JSON object that is useful to you. And I love it for that. And I'm pumped that OpenAI made a good small model. Again, it's been a while. Speaking of companies making small models, I want to talk a bit about V4 Flash.
V4 Flash is a fascinating drop by our friends over at Deepseek. There's a lot of things that confuse me about the direction Deepseek is going. In particular, I don't think they should have used the V series for these new models because R1 was a huge moment in the software dev space that caused a lot of businesses like whole business models to be put into question. R1 is the brand that people had the strongest positive association with with DeepSeek and instead they chose to keep iterating on V3.
The point of R1 is that it was a reasoning model. V4 models from Deep Seek are both reasoning, but they dropped the R and went back to the previous like V series cuz V3 was great, but no one knew about it. V4 is also really, really good. I also want to note that both V4 Flash and Pro have had two snapshots come out that are both called V4, but are very different in their capability. And I'm talking here about the newest snapshots that recently dropped for both of these.
And the Deep Seek V4 Flash snapshot is really, really impressive. I wouldn't quite a tier it. Honestly, the more I think about this, Flash and Luna probably deserve the same tier. I think I need more tiers already. Damn. Cool. There we go. This feels much better and it gives me room to do other things, which I will likely want to do here. V4 Flash and Luna are both unbelievable values. And V4 Flash, unlike Luna, is open weight, which makes it very accessible to do whatever you want with.
It is a big model, even though it's flash as in like not big as in you need an H100 fleet to run it, but big enough that you're not running this on your own GPUs. You probably need two DGX Sparks for it. But V4 Flash being runnable on hardware you could reasonably buy and have in your house is pretty cool. And since you can run on your own stuff, you can use it for things you might not be able to use other models for.
I would put it slightly ahead of Luna for these reasons, but Luna's price drop makes it an unbelievable deal, especially when you combine that with the 50% off you can currently get on Open Router. Luna is a slightly better deal before Flash is much more freedom and similarly capable. I do find that Flash gets distracted and goes in loops a little bit more aggressively. Like I often have to like stop it and be like, "Yo, you're on the wrong track." Luna is a little quicker to bail out when it thinks it's doing things wrong. from my experience, but both are very good models and worth considering for things that you're running like in the background, summarization work, things that you don't need to merge code on top of.
Turns out Luna ends up at a little under half the price of V4 Flash for tasks like the artificial analysis intelligence index. Worth noting, I find them to be similarly capable, but if Luna is this much cheaper, probably worth considering Luna as a first stop. I also love that Luna is included in my subscription with OpenAI stuff. So, it is literally in quotes free for me. So, yeah, Bullies are really good. V4 Flash gets bonus points for being open weight.
Luna gets bonus points for being slightly cheaper and available in your codec sub. Both are very good. I understand why you'd want to use either. Just learned that V4 Flash doesn't have vision. Very annoying because I paste images all the time. For that alone, I will personally move it behind Luna. I need to be able to give it an image. That's important to me. So that that I I regret buying the second DGX Spark now that I know that.
I almost want to knock it down a tier for that if I'm being real. I'll leave it here because it's good. But that does hurt a lot. Does V4 Pro at least take images? Guys, please tell me V4 Pro takes images. If V4 Pro can't do vision, it does not. Okay, the idea of a pro tier model that is expensive that I can't give an image in 2026 is embarrassing. I think I have to F tier it for that. It hurts because I love Deepseek, but at the price they charge for V4 Pro at the size of model they have created with V4 Pro, not being able to give it an image is pathetic.
But on the topic of open weight and vision, it's Kimmy K3 time. Kimmy K3 blew me away when I tried it. It was the first time I had an openweight model that I felt like could actually do my end toend work, like longer complex tasks that required touching a lot of things and staying on top of stuff. It also has surprisingly good taste around design and it has vision that is really really good. Its ability to like pick the right pixels and images and stuff is genuinely impressive and it has novel 3D capabilities.
Like it's better than every other model at some specific 3D stuff right now. It really really impressed me. I am hesitant to put it a tier. I I almost feel like I should put it a tier and have soul be S. Instead, I'll put it right at the front of B tier. It is unbelievable what it is. The main reason I'm not putting it higher is that it's not as cheap as people seem to think. Just cuz it's open weight doesn't mean it's way cheaper.
It should, but it doesn't. The problems with K3 actually kind of stem from this. Kimmy K3 Max costs slightly less than 56 soul on max, but slightly more than 56 soul on XHigh. I mostly use soul on XHigh. Like that's just the setting I use and Kimmy when it came out only actually provided max. I think you have more options now that it's the weights are out but when they had it over their API you could only use it on max.
So for the work I do, Kimmy actually was coming out as slightly more expensive than 56 Soul was, which sounds crazy, but yeah, when you're that token efficient, and 56 Soul is incredibly token efficient, the cost per token difference matters less. And Kimmy K3 is a massive model. So the cost per token is still quite expensive. Yeah, it's $15 per mill out. That's Terra prices per token, but without Terra's efficiency on token utilization.
So, this model ends up being way more expensive than people seem to expect. But Theo, aren't all of the other hosts going to add it, and the providers will race the price to the ground? Let's take a look. Oh, huh. Seems like pretty much every provider is offering it at the same $3 in, 15 out, other than a couple small ones, but only like a very small number of them. The vast majority are still at that $15 point. There's two reasons for this.
Reason one is that the model's expensive to host. So the companies that are hosting it aren't going to like take a huge loss just for fun. The second more important one is the license. The license for K3 expects companies that are at a certain level of yearly revenue to do a deal with Kimmy before they can continue serving the model. So while it is open weight in the sense that you can download the weights and the license is relatively permissive, there is a revenue cap and I think it's 10 mil a year rev where the license now requires you to do a deal with Kimmy and it seems like the terms of those deals with moonshot are that you have to match the MSRP rate otherwise these companies would be fighting a lot harder to have cheaper rates and I happen to know that Sale research and morph are not at that revenue ceiling.
I don't know how digital ocean snuck in here. or they might be violating a license by being here. Not my problem. Just saying. I was so confident that Kimmy K3 pricing wouldn't go down that I made a bet with the CEO and founder of Hugging Face about it because I I was that sure and I won by the technicality of this license thing. So yeah, I'm not going to cash out this bet because it's silly that the license is guaranteeing my win.
But yeah, it is what it is. Kimmy K3 is not as cheap as people expect it to be. It actually is more expensive than soul on x high for a lot of stuff. As such, I will continue to use soul, but I will always give bonus points to openweight models that are actually pushing Frontier and moving the industry forward. And Kimmy, hats off for that. Let's clear out the other openweight models quick so that we can get on to the big hitters like Gemini and Grock and of course Fable and Opus is what you guys are really here for.
Let's start with GLM53, the last currently openweight model in the list. Spark will apparently become openweight soon, but it is not right now. 53 is what we're here for in 53. [sighs] Think we have our first Ctier. It's not bad. Pretty good model. It's a meaningful improvement from 52. It stays on task a lot better. It doesn't go in loops that are as stupid, but it's not what I would pick for most stuff. It feels like a catch-up model, like them trying to get where K3 is, but it's not a new pre-training.
It's just more refinement on 52. and 52 was incredible, but they quickly stopped being the lead after K3 dropped. I did not expect Moonshot to leaprog Zai as hard as they did. And also, no vision. The no vision thing, honestly, I think I have to Dtier it for that. I'm being harsher on Deep Sea cuz I expect better from them to be clear. And also D4 Pro is a huge model. There is no excuse for that not having vision. I guess if we're talking about the Kimmy models, we should talk about Composer 2.5.
If you're not familiar, Composer is by cursor, now known as SpaceX's cursor, probably just SpaceX in the near future. Composer was built on top of Kimmy K2. I don't think it was 25. I think it was just standard Kimmy K2. And they RLED the absolute out of it to make it a way better coding model that was super fast and nice to use in cursor. And it was genuinely impressive, especially like when it dropped. Since then, models have moved a lot, and it felt like they were trying too hard to catch up with where things were at the start of the year, and their existing way of training would never have got them to where stuff is today.
Well, Composer 2 dropped. It was really impressive, but I found myself not using it too much. Composer 25 was even more impressive, especially at those speeds, but I found myself still just not using it a whole lot. I was wrong. It was built on top of Kimmy K 2.5, as was Composer 2. The 2.5 and two were both based on Kimmy 25, which is confusing, I know, but yeah, pretty pretty good model. Really showed how much better cursors RL stuff had gotten, but also not a model I found myself reaching for a whole lot.
Apparently, it has since been removed from Grock Build and other tools. So, I guess that XI doesn't care that much about it either. I'll bump it to the back of D tier because as hopeful as we were about it, the fact that it's only available in Grock and cursor, not over API is annoying. The fact that it is relatively slow on its base tier is annoying. And they advertise the really cheap price for the slower version and the fast version is so expensive, it's now comparable to other more expensive models.
So, they're advertising the speed and the cost even though you don't get the speed at the cost they're advertising. So, for all those reasons, I'm going to drop it in D for now. H maybe high D because it was like cool and useful and I still would use it for demos just to like get something out fast but more impressive technically than in day-to-day real world use. So what do we got left here? Gro 46 Muse and Gemini before we can get into the chaos of anthropic.
Since we just talked about Composer, I'll throw in Gro 46. If this was Grock 45, I'd be putting it in Btier almost certainly. For me, Grock 46 felt like a downgrade. The model got less efficient at its token usage. So, it ended up being a decent bit slower. And the thing I liked about 45 so much is that it had like some of the orchestration capabilities of modern models where I could give it two disparate tasks and it would keep track of both and set up sub agents to do them properly.
Grock 46 is slightly better at that at the cost of the speed because again, it does more tokens. So it ends up feeling quite a bit slower and that's not what I wanted the Grock models to be for my use. I wanted them to be fast as hell and surprisingly smart. 46 got less of that. Still useful. Still see myself playing with it a bit here and there, but it's really hard for me to justify using it when I can get similar quality results on 56 soul on low and get a response slightly faster at around the same price.
Apparently it's really good in Grockbot, which I did have early access to but never meaningfully tried. So, I'll give that a go. I could see it being useful for like assistance and for the price, it's probably decent, especially if you keep it on lower reasoning efforts. I just wish they didn't make it so much less efficient compared to 45 and obviously OpenAI models. And if I'm going to be real, I would throw Muse Spark right next to it.
I think Muse has a couple more like useful novel use cases because it is surprisingly fast and surprisingly accurate. I don't like the code it writes. I would if I was using these for code, I would pick Rock 46 over Muse Spark any day. But if I'm using these to like go through all my open PRs and figure out what I should be prioritizing and getting done, Muse Spark does it in under two minutes when models like Grock will take 10 plus and models like Fable will take an hour plus.
And the cost reflects there, too. Muse Spark also has the absurdly cheap contributor tier where they use the data for training, but the price is like under a tenth of what it would be otherwise. And it's also going to be open weight soon, which is really exciting, too. I have been impressed with Muse Spark. I think it's better than people give it credit. Still not my favorite model. Still not something I would use for code, but as a way to like process data and I think it's a little underrated.
Speaking of incorrect ratings and models that should be good for processing data, Gemini 37 Flash into the F tier it goes. I have no idea what the is going on at Google. It is really bad. This hurts me more than almost any other placing on this list because Gemini 2.0 Flash was my one of my first S tier models. I loved 20 Flash. It was really fast, really cheap, surprisingly good, had all the cool capabilities that Google put in the Gemini API at the time.
Gemini 20 Flash was a meaningfully useful model that did things nothing else could. Someone in chat said that Google deserves its own tier at the bottom right now. Honestly, you're not wrong. I'll throw 31 Pro in there too because it's just so out of date and like 35 Pro should have dropped four months ago is bad enough they haven't bothered. It's so bad. Gemini 20 flash was 10 cents per mill in and 40 cents per mill out which was an unbelievable value at the time especially but even now if you've been around for a bit then you would know to not care about these numbers because token efficiency is what matters. 20 flash was not a reasoning model so the output token cost was whatever was actually printed.
So if you asked it to audit a 100,000 tokens and give you a new title for the thing, it wouldn't burn a thousand or 2,000 tokens reasoning about it. It would just cost the tokens for what it output. Super efficient, really easy to reason about what your cost would be when you used it. Super fast, able to handle audio and images and all these other annoying things. It was a good model. Then they introduced 25 to five flash.
And 25 flash originally had split pricing. to five flash with thinking off was only 60 cents per million output tokens, but if you turned on thinking, it would jump to $3.50 per million output. Remember, when reasoning off, you do less tokens. So, they're not just massively inflating the price per token here. They're also massively inflating the number of tokens. And in realworld use, 25 Flash ended up being 10 to 100 times more expensive than 20 flash was.
And it's gotten worse since. 35 Flash came out at $1.50 per million in and $9 per million out. It is now over half the price of Pro and uses more tokens than Pro, so it often ends up being more expensive than Pro at a similar quality. What the I want to give them some credit because 37 Flash accepted that this was the worst possible direction they could go and they lowered the price 75 cents per mill in and 375 per mill out under an introductory discount.
This is a half off that only lasts until December 31st, at which point they are doubling the price. Think about this for a second. This model came out a week and a half or two weeks ago or so in August. They're going to maintain this discount price until December. They're going to in five months double the price of a model that was out ofd and overpriced when it came out is going to be 2x more expensive. Are you joking?
Are you making fun of us? This is insane. This is such an absurd level of incompetence that it's hard for me to fathom. What the It is what it is. Google models deserve the Google tier. Someday Gemini might be useful again. Right now is not that time. The Pro series had some novel vision stuff back in the day. Like 3 Pro and 31 Pro were the best models at like marking up an image and putting things in the right categories.
Gemini 31 Pro still has the best skate bench score because it doesn't get caught up on some weird stance things that other models don't understand with skateboarding. There's so much knowledge baked into the model that is impossible to pull out of it because it gets so stupid when you try to use it. Gemini models make no sense for anything ever. And don't start about the well their TPS is faster so you should use them for things that are latency sensitive because in the real world TPS is not the measurement.
The analogy I gave when somebody was being annoying about this before was imagine that you're measuring how fast a bike goes by how fast the wheel spins and how many spins you get on the wheel a minute. That sounds really smart until you realize that there are some wheels that are bigger or smaller than others. And it doesn't matter that the Gemini models are faster, two or 3x even at token generation because they generate 10 times more tokens.
We'll go back to deep sweep for a sec with 56 soul. Here's how many tokens it took per task on average for the different reasoning levels. On the highest reasoning tier on max, it would use 60k tokens. But as I've told you guys a bunch, I don't recommend using max. I recommend high or maybe x high. And as you see here, the score gap is not that big all considered. High uses half as many tokens at 28k and gets nearly the same score.
Knowing that max uses 60k tokens and high uses around 30k, how many tokens do you think 37 flash is going to use? I bet you didn't expect 37 flash on low would do 73,000 tokens and that on high it scores worse than on medium by a little bit and is using 107,000 tokens. This means if you switch to cost, you're not getting much value here at all. And this is with that discounted price, by the way. If you bump it to the normal price, you're It's a bad model.
It's not like tolerable for some things or better for some things. It is just bad. And when you do this many output tokens for basic work and have the unreliability of the to the point where you drop points when you increase reasoning, the model is trash. It should be treated accordingly. Just don't use Gemini. And now we're left with what everyone is here for. Fable, Opus, and Sonnet. We also have a tier open at the top.
Still, I have one thing I need to put an S tier really quick before we continue though. Our sponsor. We need to be more honest with ourselves about the current state of engineering. Things have never been easier when we're actually sitting trying to build and write code. AI has fundamentally changed that for us. For the most part, it's made things easier. But there's one thing it has made way, way harder. Hiring. I cannot tell you how miserable it is putting up a job listing and getting flooded with all these crappy AI generated resumes.
How do you find a good engineer that has the skills that you and your company need? It's not very fun. At least it wasn't until today's sponsor came around because G2I has the best network of engineers possible. I'm very thankful G2I is here because they pretty much solved engineering hiring for companies of almost any size or scale. Whether you want a handful of mobile experts for six months to unblock a release, or you're hiring up a team that you plan to keep for years, if not longer, G2I's network has you covered.
There's a reason why companies like Web Flow One, Password, and Meta all lean so heavily on G2I for hiring. The TLDDR, the process, you join them in Slack, they integrate like part of your team, you explain what you're looking for, you write some questions for the engineers, and then they go and interview their network and bring you back the results. For most companies, you can go from first interview to first PR landed in just a few days.
They aim for seven, and I've seen them hit this enough times that it's almost scary. I've seen a ton of companies get screwed by hiring the wrong person, not only is G2I going to make that nearly impossible to have happen because their hit rates are as high as 90% in an industry that's normally closer to 40. If a bad candidate does somehow sneak through and end up at your company, you can have them replace with somebody who better fits your needs in less than a week.
If you need more flexibility or better engineers, stop hiring recruiters and start hiring soy. G2I. So, how do we feel about these models? Sonnet 5. I my instinct is F tier, but I have actually been using Sonnet 5 a decent bit. What I've been using Sonnet 5 for is testing Claude code sub aents in T3 code. I have it do a bunch of like random dumb math for me, things that aren't that important, but I like want to see it visualized in T3 code. and it knows how to spin up sub agents and do things and it has vision that works and it's not super token efficient and it's more expensive than it should be but it it functions as an option that you have inside of your cloud code sub I would low D tier it as a thing that you choose to pay for over API easy F tier never ever ever select sonnet 5 in anything other than your cloud code sub but if you're using it in cloud code trying to like maximize your usage for your sub and have very specific small things you're using it for and paying attention to, the lowest of D tier is an okay place for it.
And now we have Opus 5. If you saw my video about Opus 5, I was very nice to it because when you first try it, it seems incredible. It's like getting similar levels of detail as Fable. It's finding things even Fable missed sometimes. When you have Fable and Opus both write a plan and review each other's work, Fable will choose Opus' work surprisingly often. The issue with Opus isn't that it doesn't act smart or have useful insights when you work with it.
The issue happens once you start to merge the code it writes. Its text outputs are vague and jargony enough as is. Its code has the same problem though. It looks, talks, acts, and behaves like it's going to be really smart and useful, but then when it has to actually do the thing, the outputs are just bad. I don't know how else to put it. I I've never had such a weird almost like mimic behavior from a model before where everything seems good where it looks like a duck, it smells like a duck, it quacks like a duck, and then you cook it and it tastes like Am I low C or high D?
I I would put its capabilities in low C. I would put its price and its mimic behavior where it tricks you into thinking it's better than it is. I'll justify high D simply because it it tricked me into thinking it was good and I'm mad at it for that. And that leaves us with one remaining model and one unused tier. You get your win anthropic. Fable is the only S tier model right now. This is not meant to say it is perfect.
This is not meant to say that Fable is the only model you should ever use. Fable knows more than anything I've ever used or interacted with. It's unbelievably thoughtful. It trips over itself sometimes as a result. It tries to take shortcuts when it doesn't necessarily need to. It loses track of what it's doing sometimes and touches things it shouldn't. It is a genius that has to be tamed. Whereas 56 soul is a slightly dumber robot that does exactly what you tell it.
If I had to pick between Fable and Soul, I would pick Soul. It's the model I default to. It's the model I use all the time for all sorts of I would miss soul more than Fable, but Fable is the best model. It is the model that writes the code that I'm most willing to merge. It is the model that I trust to double check my work from other things. It is the model I talk about hard, deep things with, like problems that I want to build, areas I want to explore.
Fable 5 is the next generation. 56 Soul is an unbelievable model that makes something feel next generation while still being built on the last generation of tech. If you don't like how Fable behaves by default, it's a lot harder to fix. If you don't like how Soul behaves by default, it's very easy to adjust. Fable is that genius at the company that nobody wants to work with, but nobody wants to fire because they are the smartest person there.
And if you know how to deal with that person, it is incredible. Fable is so good that I went from encouraging people to cancel their Claude subs to maintaining five of them myself and regularly draining them. I just got back from a bunch of travel, so I haven't been coding as much the last few days. That's why these are all up right now. But in a normal week, I will burn through all of these subscriptions. It's the first model that's good enough to make me do something like this, like setting up a proxy just to bounce between accounts because I want it that badly.
I like it so much that it has made me be nicer to anthropic because it is incredible that they made a model this good. It pains me that they have something better internally that they don't plan on ever releasing. It's been rumored to be called model 2. They're scared because it might get banned again. They don't want anthropic employees to lose access, but also they hate people externally having it. all the weird restrictions around what it will downgrade for and refuse to respond to and all of that sucks terribly.
There's a lot of things about Fable that suck, and I'll be real, most of them are anthropic. That doesn't mean it's not the best model ever. There is something that this chart doesn't properly represent, though, which is the size of the gaps between tiers, because realistically, what I would do is I would make Fable 5 like an S plus tier. I would put Soul at S tier. I would leave A and then leave everything else as is. because there is a meaningful tier gap between what you get out of the best B tier stuff and what soul represents and it's not just a single letter.
It's not like you get 20% better going from K3 to Soul. You're able to do things and work in a way that is fundamentally different with these two models. When I use something like K3 or Luna or Flash, I have to be in the loop. I have to work on what I want to build, spec out the build, and then have the model go implement parts of it, review it as it goes, review the code before it gets put up for PR, verify the changes it makes, and then ship it.
With Sol and with Fable, I tell the model vaguely what I want. It gives me a vague what it wants to do with it, and if I like it, I give it the thumbs up and say, "Okay, build it and put up a PR when it's ready." And it will computer use to verify its changes. It will use sub agents to review stuff and be sure that the good code is good before it puts it up. It will follow my instructions on how to file the PR so that it's formatted in a way that I'm not cringing at when I look at it.
It will let me know when the code is actually ready. The gap here is that Fable understands what I want slightly better and it writes code that I'm less scared to look at and more happy to merge. 56 Soul is very thorough, but it writes worse code. Honestly though, if you put Soul in S tier, I wouldn't blame you. It is so goddamn good. And when I use Fable too much for a while and then go back to Soul, I find myself relieved in a lot of ways.
And if you're building iOS stuff, Soul is just really far ahead for some reason. I don't know why Fable's so bad at iOS, but it is. Soul is much better at it. Yeah, I think that's my list. There's a lot of models that aren't on here. The reason for that is I don't care. If you are upset that some model that you really like isn't on this list, let me know in the comments and tell me what it's good for so that I know why I would want to try it.
This is how I feel, though. I should also preface this with I would never pay the full API price for Fable. I did briefly consider it, but I'm doing thousands of dollars a week on Fable right now. It's only viable because of the sub discounts. Soul is a much better value per dollar, and you get way more out of your sub as a result. So, yeah. All of this to say, I'm very excited for Astra whenever OpenAI decides how they want to roll it out or if they want to roll it out, because hopefully, if it is as good as everybody's been saying, it should plug the gap between it and Fable.
And hopefully, I'm praying it can write decent code and finally make good-looking front ends. We'll know when it drops, but until then, I think this tier list accurately represents how I feel. And as you've probably guessed by now, I am pretty much exclusively using the two at the top, unless I'm debugging something. Let me know what I got wrong and how you feel about tier list videos like this. I'm planning on doing one of these every few months now, whether it's about the models or harnesses or other things.
So, let me know what you want to see on a tier list and what you wish I did differently here.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.