Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
3:093.8x the video's typical replay level
want to make sure we understand what we're even looking at here. I also want to turn back on the Gemini models because it makes the chart much funnier. OpenAI can score meaningfully higher on Deep SWE with 20K tokens than the best Gemini models can score with 270K
Said at 3:03
Most replayed moment #2
20:483.7x the video's typical replay level
in history. Oh, is there a leak of a reasoning trace on my screen right now? Turns out there are very few things in the non-deterministic world of LLMs that are as reliable as a company like OpenAI would hope. And as a result, some of these reasoning traces have leaked. And
Said at 20:41
Most replayed moment #3
8:122.3x the video's typical replay level
input affected the model so they don't have to regenerate that every single time a new thing comes in. So for example here, when it makes the read package JSON request, it knows what the state of the model was when it got there. So it saves that and then when this new
Said at 8:05
The graph counts replays. It does not show where viewers stopped watching.
Words
6,300
Runtime
31:30
Speaking pace
200wpm
Reading time
26min
200 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
There's a lot of debate about which models are the smartest and the most capable, especially for things like writing code. The winner right now is Fable, but we don't have access to it, so we have to debate between Opus 4 8, Gemini, and whatever is going on with opening eye, which right now seems to be GPT-55. I've heard rumors something's coming, but who knows when that will happen. There is one thing that we can't really argue about though, and that is the efficiency of these different models. Because as smart as models like Gemini and Fable
100 words, the words spoken in the first 30 seconds at 200 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 341 |
| Average words per sentence | 18.5 |
| Longest sentence | 87 words |
| Questions asked | 11 |
| Sentences containing a number | 43 |
Most used terms
Filler phrases
64 in total: like 37 · actually 19 · kind of 2 · literally 2 · uh 2 · you know 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
There's a lot of debate about which models are the smartest and the most capable, especially for things like writing code. The winner right now is Fable, but we don't have access to it, so we have to debate between Opus 4 8, Gemini, and whatever is going on with opening eye, which right now seems to be GPT-55. I've heard rumors something's coming, but who knows when that will happen. There is one thing that we can't really argue about though, and that is the efficiency of these different models.
Because as smart as models like Gemini and Fable can be, neither are particularly efficient when you compare them to what OpenAI can do on much much smaller token budgets. Charts like this one from Deep SWE really emphasize what I'm trying to talk about here, where GPT-55 medium got an incredible score on their benchmark with only 20K tokens, whereas a model like Opus 4 8 scored lower at 50K tokens. And even the heaviest 55 run on X high was only 46,000 tokens.
The efficiency that OpenAI is when pulling off with their models and the capabilities that they have demonstrated on really really low reasoning budgets are insane. And I think there's a lot we can learn about if we try to figure out how they are doing that. From what these tokens are even being used for to how they're affecting the intelligence of the models to how they affect the actual quality of the outputs we are getting.
This isn't as simple as a number on a chart. It goes much deeper and I'm going to do my best to explain why after a really quick word from today's sponsor. While AI has gotten much better at writing code, I never thought I would see the day where I could actually use a computer properly. Sure, if it can write commands, it can do a lot, but what happens when it needs to navigate a complex webpage and click buttons and all of those types of things.
It turns out OpenAI and Anthropic have actually put a lot of time in and made agents really good at that. If you don't believe me, open up Codex and tell it to go to configure something on the Google Cloud dashboard for you. I never thought I would see the day it works, but it does now. But what does this mean? It means your agents need a browser in order to take advantage of that capability. And sure, it's nice and fun when you're running it on your MacBook, but what happens when you want to put out a real service to real users that requires your agents be able to navigate the web?
Well, I hope you know about today's sponsor before you build that because BrowserBase makes it 100 times easier. These guys built the perfect browser for your agents and they made it as easy as possible to set up. You can literally just click the setup for agents, paste it in your agent of choice, and then you're good to go. Because BrowserBase provides everything from the SDK to the web infrastructure allowing your agents to use the internet.
Fun fact, did you know 85% of the web isn't exposed over traditional APIs? They're just exposed for specific web services when you use them. That 85% of the web is suddenly accessible when you use a service like BrowserBase because the agents can actually explore the pages directly. You can use this for your own services to catch bugs. You can use this for platforms you're building on top of to integrate them better and so much more.
But if the things you're doing are simpler and you just want to, I don't know, grab the content of a page in markdown or go search the web for certain content, they now provide that, too, with their super simple search API and their fetch API that let you get results back in HTML, JSON, or markdown, the greatest language ever. Jokes aside, they built something awesome here. And if you want to see your agents really use the web, check them out now.
It's soydev.link/browserbase. Before we go too deep into the why, I want to make sure we understand what we're even looking at here. I also want to turn back on the Gemini models because it makes the chart much funnier. OpenAI can score meaningfully higher on Deep SWE with 20K tokens than the best Gemini models can score with 270K tokens. That is a 12 to 14x increase in the number of tokens used to solve the problems.
And the score is almost half what the score was for the OpenAI equivalent. Insanity. The inefficiency of the Gemini models is crazy. And I think people get too fixated on token costs and not enough on efficiency. So, we're going to do a quick breakdown of all of that so you can better understand because I I'll be frank, I'm tired of seeing people in the comment section not understanding how these things are built. In order for this all to make sense, we first need to talk about the two types of tokens.
I'm going to put a star next to two cuz it's not as simple as it might seem. We have input and output. A token is a chunk of text the way that the models break them up in order to process information. Input tokens are the things that you submit to the model, whether that is a prompt you wrote, text that it's reading from your code base, something that it's running as a command that it gets text out of that it then puts back in.
All of the content being ingested by the model is tokenized and broken down into these input tokens. Even when you're ingesting PDFs or images or any other format, it gets broken down into these tokens. Output tokens are what the model generates. Those are the things that it uses to create a response as well as the response itself. As I said though, it's not as simple as two types because both of these break down a little bit.
Input isn't as simple as just input tokens because there is uncached and cached inputs. And output is also not as simple because there is both reasoning tokens and the actual output tokens that are the content you're reading or the code that it's writing. So, let's start by going through a fake chat history. Let's say we start with some simple prompt like, "Hey, I want to add dark mode to my app." Then the model responds to something like, "Okay, first I have to understand how styles are handled in this app." Then it goes and fires some grep request or a find in order to read the different CSS files and other style related files in your project.
Maybe it starts by just reading package.json and then it gets all of the content from the package.json file. I'm intentionally putting the package.json contents on the right side here because traditionally this is what we think of as input tokens, but more importantly, this is the context that the model is getting. So, when you make a prompt, the model starts generating, but it then realizes it needs more information, so it ends with a tool call and that tool call gets run on whatever machine you're prompting from or if you're doing this in a cloud environment or whatever in order to run the tool to get additional context that ends up becoming part of this history that the model uses to generate the next thing.
So now that the model has the context of the package.json, maybe it sees that this is a Tailwind project and it says, "This project is using Tailwind. That means I have to" and then it goes and does whatever it needs to do. When this response gets generated, the input isn't just your first message. It isn't even just your first message and the package.json. The input is literally everything beforehand. The whole history goes into the model and is used to skew the weights and skew the parameters so that they are pointed towards what we want to see, which is an answer that generates this specific dark mode that we're requesting.
So all output tokens become input tokens when the job has additional steps to run and when a step ends in a tool call, the tool call results also become input tokens. So you might think, "Oh, I only sent two sentences, so there isn't going to be that many input tokens." That is not the case at all. Anything in your history becomes an input token on the next generation, whether the generation is you sending a message requesting follow-ups or doing more work or if it is another tool call the model is doing in order to do additional stuff.
But if we were to re-ingest all of this every single time, it would be super inefficient. Like that would just be absurdly expensive. So ideally we're not doing that. There are strategies to reduce how expensive all of this is. The one that we're all probably most familiar with is compaction where you ask the model to summarize everything that happened earlier and instead of this being the thousands of tokens it might be right now, it could become a much smaller amount cuz the model ingests it once, summarizes it, and now that is the history going forward.
If you ever wonder what compaction was, that is it. It takes your history, it runs it through the model once and then generates a smaller thing that mostly represents the intent of the previous messages. That's also why you might lose details when compaction happens or it might lose track of what it's doing because the details that were much more specific before in that longer history get compacted into a smaller one.
But that's only one of the two things that allows for your history to not be super expensive. The other is caching, which is a strategy most of the labs put in in some way, some in much more effective and easy to implement ways than others, where they will keep track of how your input affected the model so they don't have to regenerate that every single time a new thing comes in. So for example here, when it makes the read package JSON request, it knows what the state of the model was when it got there.
So it saves that and then when this new piece comes in, all it has to recalculate is everything before when it cached and the stuff that has been added since. And now I can take this additional piece, add it to the cache, and then the next generation doesn't have to do all of that calculation again. So the two types of input tokens, to be very specific here, are cached and uncached. And if we look at a model like GPT-5.5, the pricing on a million input tokens is $5, but if they're cached, it's only 50 cents.
So if you're caching efficiently, your input tokens get much cheaper, but the output tokens are still really expensive. I said there was two types of output tokens, but we're only seeing one here. And no, I'm not referring to short versus long context. I'm going to ignore that for now. It just makes things too complex. The two types of output tokens are reasoning and output. I know one of those is contradictory. There's output output tokens and reasoning output tokens, but it's the easiest way I can think to explain it.
The reasoning tokens are the ones the model uses to try and improve its answer before giving it to you. It's the thinking tokens, the model taking time to decide what it wants to do before it responds. Back in my day, when I was using AI, models didn't have the ability to think. They would just start spitting out text immediately. But that means that if they caught themselves making a mistake halfway through, they were screwed.
It means if they wanted to explore other options, they would have to tell you all of those and it would just make the output unreadable. And OpenAI back in the '01 days came up with this idea of letting the model talk to itself effectively to improve the quality of the answer before it became your problem. And that idea of reasoning, letting the model talk to itself before you get the response, massively increased the quality of the answers that they were seeing from the model, which resulted in a massive surge in the capability of these models for doing difficult thinking work from coding to science to engineering to many other things.
Funny enough, even naming skateboard tricks benefited greatly from this change, letting the model talk to itself. But this had a cost, a literal cost. The number of tokens that you would get in the responses went up massively. For example, let's take something like Skatebench where I describe a skateboard trick and it names the trick. The named trick is often just like five or 10 tokens at most cuz it's just a couple words like switch backside kickflip is maybe six tokens.
But if it had to think about that first, it could generate thousands if not tens of thousands of tokens as it talks to itself trying to decide what it wants to do. I'm going to give a simple example here using GLM 52 because open weight models will show you their reasoning, which makes it easy to demonstrate. I'm going to ask it to name a skate trick. The skater is skating switch stance. They pop on their tail, flip the board in the kickflip direction, and the board spins 180° backside while also flipping.
The skater does not spin. This trick would be a switch varial kickflip or switch variflip. We'll see if it can name it properly. It's funny that our summary model did not name it correctly. This named the trick correctly. A skateboarder would call that a switch varial kickflip or more specifically a switch backside varial kickflip. No, we wouldn't call it that. We'd just call it switch varial kickflip. So it named this correctly.
The actual response here, like this part, is probably only a few hundred tokens at most. I'm guessing under a hundred. But it had to reason for a bit there. Like it took its time. And you can see here what it did in the reasoning. It analyzed the request. The user wants me to name a skateboard trick the way a skater would. The conditions, and it listed all my conditions. Deconstruct the skateboard trick components. You can see all of the things it did here.
This is a very interesting reasoning trace actually. So, if we compare this to like a dumber model, I don't know. Let's take any Quen model. Here the reasoning's going to be a lot more just text walls. Okay, so the user wants me to name a skateboard trick as a skateboarder would. Let's break down the description. And it got entirely wrong. Backside kickflip 180. But it generated a shitload of tokens talking to itself trying to figure out what it would be.
But if I did this with reasoning off, it would end up being way fewer tokens. This was 1,307 tokens. Same exact prompt sent on instant mode. Ends up being 745 tokens. I don't know how it ended up being that much. It should have been even less than that, but you get the idea. The reasoning is letting the model talk to itself as it names the trick. But you'll notice something weird if I switch over to GPT 55 and ask the same thing.
First off, you'll notice the number of tokens is very small, only 342 tokens. It got the name right. It describes why the name is right. But in the reasoning trace you'll see this is not what it actually reasoned. Here's where things start to get a little mucky. The frontier labs do not actually give you the raw reasoning tokens. They give you summaries of the reasoning. The reason they do this is they don't want all of that data getting out for other companies to be able to copy from and create models that behave similarly.
Because if you have all of the data of exactly how a GPT model came to the conclusion it came to, you can use that to make your own models behave similarly. Anthropic used to give us the full reasoning trace back in the Sonnet 3.5 and 3.6 days, but they have since stopped and also do summaries just like we're seeing here. This means we're kind of in a weird position if we try to analyze why the OpenAI models are able to be so efficient because we can't actually see what the reasoning traces are.
We can't see what's going on underneath because they're not showing us that. They are just giving us summaries. Also, somebody seems to think Google gave us the raw reasoning traces. Google has never once gave us raw reasoning traces. They didn't even give us summaries until a bit after Gemini 2.5 Pro came out because a lot of people, myself included, complained so much about how annoying it was to not have them to show the user what the model was doing.
But, there's a couple more layers we need to understand here before we go further when we talk about these reasoning traces and the things the model is doing. The summaries the model gives us are useful to get a rough idea of how the model got where it did, but they aren't the actual thinking the model did. They're not the actual information the model found when it was thinking or the other stuff it did throughout its time reasoning.
But, the actual data is useful for the model when it generates the next step. So, let's say, going back to this example, that when the model was asked, "I want to add dark mode to my app." This is an example of a reasoning trace the model might do after getting this prompt. This data is useful in future steps. The model might be confused as to why it assumed the project was JS, but if this exists in the history and the model can ingest this, it is more likely to understand on the next step why it is where it is and what it should do next.
I hope you can imagine the different things the model might think but not say that are useful in the history. Where if it realizes these files aren't useful, it shouldn't have to realize that every single time it does the next step. This information stays useful throughout the run, so we would want the model to have that, but it's not actually in our machine because they don't send it down the wire, so they have to save this on their end in order to populate the history properly.
I've talked about this a bit in other videos about context management and how these histories work. I've also talked about this extensively in my video about how OpenAI uses web sockets now for their inference. I think it's one of my best technical breakdowns I've done recently, worth checking out if you want more info on all of this. The point I'm trying to make here is that if you want to increase the efficiency of the model, the best place isn't to change how short the answers are, it isn't to prompt better, although that does help.
The best thing you can do is reduce how much thinking is going on. The fewer tokens that occur here, the fewer tokens that are used overall, both because it's generating fewer tokens in its responses, but also because it doesn't have to re-ingest those same tokens on every additional step throughout. Reducing the number of tokens in reasoning is an exponential decrease in the amount of tokens used overall. So, how did OpenAI do it?
Well, again, as I mentioned before, other labs don't seem as focused on these improvements to reasoning efficiency. Fable is better than I would have expected with medium actually competing with GPT-55's reasoning efficiencies, but X high and max are still over 100k tokens per task for the examples that Deep Mind did. Whereas, the OpenAI models are absurdly efficient, all falling under 50k tokens on the 55 line. First and foremost, I want to be real here.
The biggest reason OpenAI's models are more efficient is because they're more focused on efficiency. This is a thing they have been working on from the early days, and they have went out of their way to try and make the reasoning as efficient as possible in order to allow them to run the model more aggressively, have the models do more things for less money, use less GPUs. They want to be able to generate as many good answers as possible with as little compute as possible, and making it more efficient benefits them greatly as a result.
And when you combine this with the actual cost per token, which is what people get way too fixated on, you see why this is so valuable because the costs end up being significantly lower for 5.5 than other models simply because they're doing so many fewer tokens, even though the price of the model is higher than it was before. If we go back to the pricing chart, they have effectively doubled the price since GPT 5.4, where before it was $2.50 per mill in and 15 per mill out, now it's 5 per mill in and 30 per mill out.
That doubling doesn't hurt quite as bad. It is still more expensive, but it's not as bad as it could be because 5.5 ended up being so much more token efficient for a certain level of intelligence. So, if we look at, for example, 5.5 medium, it used under half as many tokens as 5.4 X high and came out with a higher score, which means its cost ended up also being lower at about half the price of 5.4 X high. If you compare the cost of 5.5 X high to 5.4 X high, it goes from 565 to 723.
So, yeah, 5.5 at its extreme is more expensive, but if you're comparing cost relative to a certain level of intelligence, it is indeed going down, and that is largely due to the efficiency improvements that OpenAI has been working really hard to attain. We will never know the full details of how they have managed to do this, but we have had some leaks that show some of them. Before we can understand this, we need to talk a bit about Grug.
Grug brain developer not so smart, but Grug brain developer program many long year and learn some things, although mostly still confused. Grug brain developer try collect learns into small, easily digestible, and funny page. Not only for you, young Grug, but also for him. Because as Grug brain developer get older, he forget important things, like what he had for breakfast or if put pants on. More simply put by aviator in chat here, why use many words when few word do trick?
This is a silly way of writing that is intentionally written to feel dumb, but also obvious cuz that's the point of the grog-brained developer in this way of thinking pioneered by Carson, the writer and creator of HTMX. It is meant to show you the smart thing isn't always the thing that has the fanciest words and vocabulary and the most elaborate write-ups. Sometimes the simple stupid thing is the right one. And sometimes you don't need all of those words to communicate the value.
It seems like this is one of the many tricks that OpenAI has been employing in order to make the models more efficient because if we're not seeing the reasoning traces, we don't care what language they're speaking, we don't care what vowels they're forgetting, we don't care what words they're omitting, how good their vocabulary is. We don't care about any of that during the reasoning step. We only care about the outputs.
And if OpenAI has gotten to the point where they can meaningfully draw a line between the reasoning outputs and the actual answer output at the end, where they can make the model behave one way during reasoning in an entirely different way for the actual answer outputs, they would hypothetically be able to make a model that is much more efficient in reasoning cuz it speaks one way then and gets good answers at the end cuz it speaks a different way when it outputs them.
Sadly though, we'll never know if OpenAI is doing this because those reasoning traces are super locked down. They've never leaked ever in history. Oh, is there a leak of a reasoning trace on my screen right now? Turns out there are very few things in the non-deterministic world of LLMs that are as reliable as a company like OpenAI would hope. And as a result, some of these reasoning traces have leaked. And we can see some very funny chains of thought as a result.
Need agent kind maybe open hands direct okay. Need just set tools default in JQ. Need finish tool? Default tool name is no finish. Finish is tool. The list in system prompt for delegated had only finish. Think switch LLM invoke skill because tools array empty, but default agent maybe always has internal tools. Need add terminal etc. Maybe tool name's okay. Add tools array. Here's another one. Use core new nodes. Need infer.
Note 35 maybe thing from the code base. Outputs things from code base. UI has conditioning negative. Need add VAE encode for images. Try. Try period. This is hilarious, but it's also really efficient. They have trained this model and they've R L'd it so hard that try period, which is probably one token. Oh, no. The period's another token. What a waste. This could be something that just happened as a result of how the model's trained.
This could be something they actually intentionally designed for where if it didn't have the period, it might think it needs to keep reasoning. But by putting the period in, they are forcing the model to stop there and then start doing whatever it's going to do next. But we're talking about a company here that is trying their hardest to get this to be as simple as possible to not waste history in order to make the model generate outputs more effectively because that makes them faster, it makes them cheaper, it makes them able to reason longer and get more done in a given token budget.
It's silly, but it works. Obviously, side effect one is that when it leaks, it looks really silly and you get funny memes on Twitter for it. And god damn, there have been a lot of funny examples that my chat has found for me. We need adjust UX. Need inspect current component perhaps parent max XL causing max 2 XL, but parent max WXL so ineffective. Yes, parent W full max WXL. Need baby within card. There are so many examples of this that they are clearly actually doing it.
And in order to not show this to you, they hand this to another model and say, "Hey, can you summarize what we thought about here?" And they show that to you instead. No matter how hard they try in training this with between the reasoning and the actual answers that are being outputted, there will be leaks like this. I've seen them myself. Almost everyone I know has seen this happen like at least once at some point during their use of the model.
That's a negative side effect that is only happening because they are reasoning this way. Versus again, if we look at the other reasoning traces that we got from other models earlier, like we did here with Qwen, it's writing plain English. It's not great English, but it's plain English. Or if we look at the reasoning trace that we got from GLM, it's doing bullet points and like lists here in order to break down the work that it needs to do.
Different model families reason in all sorts of different ways, but OpenAI's models seem to be reasoning in a very novel way with this crazy token efficiency style. And now to drop some potentially hot takes about effects that I think we see as a downstream impact of these decisions. First and foremost, I think this is a meaningful part of why Claude is nicer to talk to because when Claude talks to itself, it is probably talking in plain English.
Just seeing how much longer the reasoning traces are, it's probably talking in plain English. It's also worth noting that if Anthropic was to fix this, their income would go down meaningfully. If developers can do work with the models and fewer tokens are generated for the same work, that is money they're not making. So I would suspect other labs aren't as interested here because they're not as interested in lowering how expensive these models are for others to run.
But on that note, I think this is also why Claude defaults to 1 mil token context windows. This isn't because it wants to fit bigger code bases like so many people seem to think. This is because the outputs that Claude is generating and the reasoning that Claude is doing is long as And in order for that to fit within the context window, it needs a bigger context window. They're also really bad at compaction, which is part of why that benefits them greatly.
But it's also why when you resume old Claude threads, it tells you to not try and keep the whole context, but instead to try and compact it. Because they don't keep the reasoning trace on their side. In fact, they clear them out after like 10 or 20 minutes, if I recall. Because the reasoning tokens are just not efficient enough to be worth keeping around, because your costs would be absurd if you did. Meanwhile, the token context window with OpenAI models is like 200K or less, but they're fitting so much more in that because they've done all of the crazy grug speak in order to get there in the first place.
This is also why, in my opinion, OpenAI models go off the rails faster. Not faster in terms of the work done or the number of steps, but faster in terms of the token context window. Because if a million token context window has all of this grug speak, keeping track of what actually matters gets harder than if you have properly formatted text with like headers and bold sections and lists and formatting to make it clear what was going on, not just the path the model was going down.
When those tokens stop being paths and start being histories, this strategy seems to not be as effective, and it hurts the ability for these models to do things with the really long token contexts, which is why I also think OpenAI won't let you use the 1M token windows in Codex by default. I haven't tested this myself, but I've been wanting to. I actually told my team I want one of them to like go use an API key and test this for me for a bit.
If anyone in chat has tried it by bringing your own key to Codex and turning on a million token context, how does it behave? I don't know, but my assumption is not well, because otherwise they would have just went and turned it on. That's kind of what they like to do over at OpenAI. When a thing is useful, it doesn't matter how many tokens it burns, they'll just turn it on for users. Even the fast mode stuff, like how the fact that fast mode is available within the subscription tier on CodeX, and it's not on Claude Code, you have to pay API prices for it, is absurd, but shows the difference in how much compute availability there is at these two labs.
But, most importantly, what this means is a lot of secret sauce. When a model provider tells us a certain number of reasoning tokens were generated, we can't see them. We just have to trust them. And if the quality of the responses goes up high enough, then we'll accept the increased cost when the model has to reason more, even though we can't actually see what's going on. If we want to replicate certain behaviors or figure out why a model went down the wrong path, we can't, because all of this information is hidden from us.
This also means if you're a competing lab trying to generate models of similar capabilities that are similarly efficient, you can't get the data on how they make it so efficient. You don't see the reasoning traces. You see the inputs, you see the summaries, and you see the answers. And you have to create your own methodology in the reasoning for your model in order to get similar quality answers. Obviously, a model like GLM-52 is trained on responses from models like Opus, Fable, and GPT-55, but they don't have the reasoning traces.
So, they have to take the inputs and the outputs and build a system to dynamically generate what a reasoning trace might look like. And all of the labs are now trying their own strategies here, which is where I think things get much more interesting. Now that reasoning is essential to get the model to generate good outputs, we're no longer thinking of reasoning as just the model talking to itself. We're seeing it as like a code golf thing, where we can min-max in the reasoning trace both to increase the quality of the outputs, but also to increase the efficiency of how the model generates things, which is why GLM-52 has this new strange format for its reasoning traces.
I I tried this before, but I'm actually curious if I use an older GLM model, will its reasoning look meaningfully different? There we go. See, I switched over from 5.2 to the standard GLM 5, and the reasoning looks entirely different. Also see here, let me think about this. Regular stance, backside 180 kickflip varial flip, switch stance, that would be a varial heel flip. Nope, that is wrong. Wait, let me reconsider.
The description says flipping the kickflip direction, that means the board flips with the toe dragging off the heel edge. Standard kickflip. In switch stance, a kickflip in switch is still a switch flip. Actually, I need to think about this more carefully. See these weights and actuallys? This is how we used to think reasoning had to happen because the model was just talking to itself to get to a better answer. And it kept going, it kept thinking, and then eventually it got it right, switch varial kickflip.
Even though earlier here it was wrong, it thought this was a switch varial heel flip. But it kept saying, "Actually, no, wait, these other things matter, these other things were done differently. Wait, but there's another naming convention. No, wait, let me just go with the straightforward naming. Actually, I think I'm overthinking this." Yeah, no But again, to compare this to GLM 5.2, which here we did 1,516 reasoning tokens tokens total.
Same correct answer from GLM 5.2 did it in 600 tokens, almost a third as many because it thinks in an entirely different way now. I genuinely think it's really cool that you can still see this stuff in open weight models, and you can see how they have progressed as they try to get more efficient and smarter. And saying, "No, wait, actually" over and over again isn't the best method, even though that's the one that it seems like everyone did for a while.
A lot of the improvements we're seeing in new models are the ways that the RL process, the reinforcement learning systems they are creating are resulting in new reasoning methods being invented and tested thoroughly. All that said, GLM-52 is still far from an efficient model, taking over 42,790 tokens to complete the artificial analysis uh tasks. Like, that's a per task number. It averaged 42,790 tokens per task. Whereas, something like, I don't know, uh GPT-55 medium was 5K.
That is a tenth as many tokens for a similar level of intelligence. I think I've said all I have to here. I found this really interesting, and when I started to see those silly leaks of the grug brain style talking, it just I couldn't get this out of my head. I thought this is really cool, and I hope you guys agree. I know you guys miss the technical deep dives. I hope this qualifies as one. I found this fun as hell.
Let me know what you guys thought about it. And until next time, peace, nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.