Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
3:134.7x the video's typical replay level
see it yourself, check them out now at swive.link/firecrawl. As I mentioned before, I've been using this model for a ton of real work, as well as some silly demos, like making 3D games and whatnot. I've been going at it non-stop for the last 24-ish hours, and I have a lot of thoughts. But first, I want to start with
Said at 3:06
Most replayed moment #2
19:472.8x the video's typical replay level
intentionally just to see where it would stumble and it didn't. Did it find everything Fable found? No. Was its code as thorough as 56? No. Was it able to go back and forth with me on really big, heavy tasks and be pleasant to use while
Said at 19:39
Most replayed moment #3
14:062.6x the video's typical replay level
them, it's kind of insane. They went from in the thousands on the human baseline to 1543 leaprogging across multiple labs entire like last decade. It's crazy how far they have jumped on all of these things. and benches like how to they are now number one in the
Said at 13:58
The graph counts replays. It does not show where viewers stopped watching.
Words
5,043
Runtime
24:43
Speaking pace
204wpm
Reading time
21min
204 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
A new model just dropped and its creators are making some very bold claims. Specifically, they're saying that for dev work, it should compare to models like Fable 5 at a fraction of the cost. The model is Gro 5 from Space XAI, now partnered with Cursor. And I was a little skeptical after hearing this, but it turns out I was actually testing it. Over the last 24 hours, Cursor gave me early access to a new model that I thought was going to be a new composer, and I was pretty impressed with it. I learned today that that model
102 words, the words spoken in the first 30 seconds at 204 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 315 |
| Average words per sentence | 16.0 |
| Longest sentence | 117 words |
| Questions asked | 9 |
| Sentences containing a number | 83 |
Most used terms
Filler phrases
55 in total: like 34 · actually 8 · kind of 8 · literally 2 · right? 1 · uh 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
A new model just dropped and its creators are making some very bold claims. Specifically, they're saying that for dev work, it should compare to models like Fable 5 at a fraction of the cost. The model is Gro 5 from Space XAI, now partnered with Cursor. And I was a little skeptical after hearing this, but it turns out I was actually testing it. Over the last 24 hours, Cursor gave me early access to a new model that I thought was going to be a new composer, and I was pretty impressed with it.
I learned today that that model was actually Grock 4.5. And now that I'm seeing the benchmarks, yeah, it's a pretty damn good model. This is from the artificial analysis code index, which combines a handful of benches that I actually like and trust. And according to this bench, Gro 45 is neckandneck with GPT55 and just barely below Fable while also beating out Opus 48. While I think these numbers might be a little bold based on my experience using it, they're not that far off.
Gro 45 has been genuinely impressive. And it has a few things that are truly novel to it that I never would have guessed. And at its current price, it's a steal, especially with the 50% discount they're currently offering for people using it through tools like Cursor. I want to break down all of the good, the bad, and the ugly with this model. But first, a quick word from today's sponsor. Normally, my ads are showcasing all the cool things AI can do.
I'm going to do something a little different here. I want to show you something AI failed hard at with me. LG, notoriously wonderful at naming, announced a monitor I've been really excited about since January that still hasn't come out yet. This monitor has a bunch of cool tech that I'm excited about. That's not what I'm here to talk about, though. I'm here to talk about my attempts to get it. When you click the notify me button, it doesn't notify you.
It launches a broken JavaScript thing, which in other browsers lets you sign up for notifications, and it doesn't even have the monitor I want in the options. So, I have no way of knowing when this monitor comes out. So, I did what any nerd would do. I asked my Hermes agent to monitor it and let me know when it comes out. And it did every single day because it was waiting for content on the page to change. And the content was changing, but never changing in the ways that mattered.
And my AI was not smart enough to realize that the reads it was getting were not actual changed content on the page. All I wanted to know is when this monitor came out for sale. Today's sponsor is Firecrawl and they make it way easier for your agents to scrape the web. That in and of itself would have been really useful for this type of thing. If it translates the page to markdown, it's more likely to notice when things have changed because it can get the page in markdown.
It can get a screenshot of the page or it can get the plain text. Super useful. But what it can also do is a new feature they just added called monitoring where you schedule a recurring check to detect changes on a page. This makes this exact task trivial. And not only is it easy, it falls entirely under the free tier. And even if it didn't fall under the free tier, I could have just forked it and ran it myself because they're open source.
Buyer crawl has pretty much everything your agents need to scrape the web. From search to URL specific readouts to interactions to monitoring to an MCP server to make it way easier to connect that doesn't even need an API key, so it's trivial to set up. I can see why these guys are hiring right now. They're clearly doing really well. All of this is so useful. And if you want to see it yourself, check them out now at swive.link/firecrawl.
As I mentioned before, I've been using this model for a ton of real work, as well as some silly demos, like making 3D games and whatnot. I've been going at it non-stop for the last 24-ish hours, and I have a lot of thoughts. But first, I want to start with the official reporting. Introducing Grock 4.5. It's SpaceX's smartest model built for coding, agentic tasks, and knowledge work. Let's talk about what we care about.
Real world engineering excellence. Grock 45 was trained on data sets spanning knowledge in coding, science, engineering, and math. With both intelligent and efficient reasoning, Grock 4.5 excels at real engineering tasks and it exceeds comparable leading models at many of these tasks. So, Deepsw SWE, which is the software bench that I've been repping pretty hard. I find it to be a very reasonable benchmark measuring how work actually happens with agents.
Obviously, Fable is still number one. DBT55 is still number two. We don't have numbers for 5.6 yet. Excited to see those. But for now, third place is Grock. Massively beating out any Google model. I don't think any are even referenced here. Opus 48 is falling a decent bit behind and then 47 is way lower. This is a new like frontier tier that we are seeing happen and Grock 45 is on the line between last gen and this gen in a lot of ways. 4.5 was trained across tens of thousands of GB300 GPUs which is the newest technology from Nvidia.
Training and stability techniques designed for large scale runs beyond raw token volume. We invested heavily in data filtering and curation, dduplication, quality scoring and domain focus selection. So the data mixture stayed high coverage and high signal. That is interesting. This model does seem to be a full new base like pre-training rather than an adjustment on previous Grock models. The biggest indicator of that is that they mention it's a 1.5 trill per model where previous ones were only 500 bill per.
So clearly whole new model here. Their RL training covered hundreds of thousands of tasks centered on multi-step software engineering and other technical work with automated and modelbased grading. Our stack is built for highly asynchronous training, so agentic rollouts can run for many hours while learning continues across tens of thousands of GPUs. The result is more intelligent and efficient reasoning on real engineering and agentic tasks.
They give some examples of one-shotted tasks that they have the model make, and it's impressive. It does 3D particularly well, which I will be sure to talk about in a bit. It's also quite fast, both because it uses not too many tokens and because it is running faster in general on the info they're serving it on at roughly 80 tokens per second. Cursor, who is now part of SpaceX, has their own article about this drop and it has a couple more interesting details I wanted to jump on.
First off, they have more benchmarks, including Terminal Bench 2.1, where it performed just barely below GPT55 and pretty close to Fable with an 83.3 versus 83.4 and 84.3, respectively. SWE Bench, which I don't care about. It's a bad benchmark, so I'll skip it. And then Deepswe, which you mentioned before, it's doing very well on. SWB Pro, which again, bad benchmark. I don't care. Cursor subscription plans for individuals and teams include significant usage of the model with double usage for the first week.
This is the other exciting thing. One of the problems Cursor has as a business is that they struggle to compete with the labs just doing this crazy subsidization because you can pay OpenAI 200 bucks and then get up to $14,000 of inference not even counting resets. That's hard for them to compete on. They've won on enterprise still because the API rates are what the enterprises often have to pay. They don't get that crazy subsidization from the labs.
But winning on individuals has been tough and this is a huge win for them there because they finally have models they can kind of subsidize. Not that they have to though because the price is really good. We haven't got there yet, but we will in just a bit. As they mentioned, Grock 45 is a mixture of experts model. They trained jointly with SpaceX. So, this is a model that was still a Grock model, still SpaceX focused, but they came in and trained it jointly, bringing data as well as their own processes.
Training included trillions of tokens of cursor data, which capture a wide range of user interactions with code bases and software tools. This data set lets the model learn both from existing software as well as developer agent interactions, capturing how developers work and how agents interact with their environments. I've seen a lot of good examples here which we will definitely showcase. When they trained Composer 25 to be a coding specialist, Grock 45 kept the training data intentionally mixed and more broad.
This involved drawing on highquality STEM tasks, research papers, and other knowledge work so that the model gained proficiency across a wide range of domains. They made a bunch of difficult RL problems for the model because even a lot of the traditionally hard problems are now trivial for models and in RL they want to have really really difficult stuff and I think that's why it's benching so well. It can get it can just go on those types of big bold hard tasks.
You might have noticed one bench missing from here though. Cursor bench. That is certainly not because it performed poorly. As we see here the top right being the cheapest and the best. It performed comparable to Fable 5 High for a significantly lower price where Fable 5 High cost $8.77 per task and Grock 4.5 cost $151 for a slightly higher score at that tier. Obviously, the best is still Fable on Max, even though it's double the price of Fable on High.
But yeah, Grock 4.5 High is looking insane by this chart, even crushing out GPT. But how? Like it it can't possibly be that good, right? If you scroll a little, you see why they did not include this information. Grock 4.5 has an advantage on cursor bench. An earlier snapshot of the cursor codebase was unintentionally included in training. The exact score impact is unclear. The data has been removed from future models.
For a rundown of third party benchmark scores, see the Grock 45 launch blog. Yep, they accidentally put cursors actual code in the training data. And since cursor bench is based on real problems that they have in cursor and working on cursor, this bench is now kind of tainted at the very least in the Grock world. It's a shame because I liked this bench, but uh mistakes happen. I'm happy they were transparent about it and that they're not advertising Grock 45 via this benchmark publicly, which other labs may have done similar things to.
They're being straightforward and transparent with it. I appreciate them for that. Now, I want to talk about the price. Gro 4.5 is $2 per million tokens in and $6 per million tokens out. That makes it comically cheaper than a lot of competing models. For example, Fable is $10 per million tokens in and 50 per million tokens out. Between five and almost 10x the cost. That ignores the fact that Fable is relatively token hungry and this model seems to be less.
So, hard to know for sure until we've really put it through its paces, but based on all the benches I've seen and all the work I've done with it, it's relatively efficient. It is also worth noting that this price only applies under 200,000 tokens of context. If you go over, the price is double to $4 per million tokens in and 12 per million tokens out. Still way cheaper than any other model at this tier, but it only goes up to 500k tokens.
It's kind of weird to have a model that can go over 200k and charges more but is still under a mill specifically like 200k to 500k is not that much more context and to build twice as much for it especially to build twice as much on output feels a little much to me. Feel like they're reaching a little here. My guess is the reason they did that is they wanted to get the input and output token costs for the base tier 200k version as cheap as possible.
And the GPUs that this is running on are also being resold to companies like Anthropic and Google with massive markups. They have to make sure that it's priced in a way where they're not losing too much money that could have been made from reselling GPUs, but at the same time is priced cheaply enough to actually compete with those labs. SpaceX has put themselves in a weird spot here, but I think they navigated okay according to everything I've been seeing and all the use I've been having so far.
Let's go over the artificial analysis numbers and then I will dive into my experience. SpaceX AAI's Gro 45 scored a 54 which places it fourth in the artificial analysis intelligence index. Kind of wild to see a different color in the top again. It's been a very long time since I saw purple all the way up here. It is right behind GPT55 and just ahead of Sonnet 5 based on the intelligence index crushing GLM52 which is even more interesting when you realize how expensive GLM52 is to run.
GLM52's base price on a lot of providers is a $110 in roughly and $4.40 out roughly. There are some providers that offer quite a bit cheaper now, like Novita has a temporary 60% off. Deep Infra has it at like $3ish per mill out. When you remember how much more token hungry GLM52 is, you realize it's kind of been crushed by Grock 45. Here we can see the actual costs incurred per task average across the entire suite. And Kimmy K26 was about 35 cents per task.
GLM52 is about 37 cents per task. And then Gro5, which scored way higher than those other models, was only 31 per task. And for reference, Fable 5 was $2.75 for the same work. Yeah, it's an efficient model. And if you look at the intelligence versus cost chart, you see it is very well positioned. It is just on the edge of the green box, you know, the the good spot that almost nothing is in. Both Gro 45 and Gemini 31 Pro score very well here.
And Grock 45 is meaningfully more intelligent while also being in this price range. This gets much crazier with the coding focused benches, though, because Grock came out swinging here, crushing every Google model by a large margin. And if we look at the token usage per task for the code work, you'll see Opus and Fable both massive token hogs at 7.2 mil for Fable and 9.2 mil for Opus. 55 on X high is still in the 6 mil range or so. 55 on medium was only 3.5 mil tokens, but all the way at the end here, Grock build with Grock 4.5 only 2 mil tokens to get that high of a score.
That is an insane level of token efficiency. I never would have guessed that this model would be so efficient, but it really is. And the result is that it feels way faster to use and the bill ends up being cheaper than you might have expected. It kind of makes this a great go-to default code model that you bring other things in to clean up after if you use it and it doesn't do quite what you want. Artificial analysis also calls out that it does very well on agentic tasks, things that are multi-step where it has to call tools and synthesize information.
One of the most costefficient models to run for near frontier intelligence. Yep, it is insanely cheap. The token efficiency combined with the low price is what makes it so compelling. As a coding agent, Grock 455 and Grock build is on par with 55 and offers efficiency benefits. Still insane because 55 was such an efficient model. Getting more efficient is just unbelievable. And when you see how big of a jump this is for them, it's kind of insane.
They went from in the thousands on the human baseline to 1543 leaprogging across multiple labs entire like last decade. It's crazy how far they have jumped on all of these things. and benches like how to they are now number one in the world in just crazy crazy leap and it really shows the benefit that cursor brings XAI in being part of the business it seems like cursor's combination of like data and RL process has been incredibly beneficial to X so far that does not mean it scores well in everything though for example in skate bench it ended up being quite expensive because it kept thinking and reasoning trying to figure out the tricks and it still only got a 76% which is the lowest score from a Frontier Lab on the max reasoning settings while also being relatively expensive at 1.3 cents per run.
Not as bad as something like Sonnet 5, which cost way, way, way, way more than it should have at 15 cents per run. Literally 10x the cost for a lower score, but you compare it to something like Gemini 31 Pro Preview and ended up being little over half the price, but a meaningfully lower score. Yeah, I was hoping it would be a little cheaper here, but it does seem very determined to get answers as it averaged at 2,100 tokens per response, making it one of the most heavy reasoning models I have used here where it really thought before giving an answer.
All of this said, benchmarks are benchmarks and a lot of them don't measure how it feels to use the model in the real world. For example, a lot of the Google models score great and when I use them for code, they just don't feel particularly good to use. So, how has Grock 45 been in my usage? I'll be frank, I'm impressed. I'm working hard on my new cloud product, Lake Bed, and I'm really close to shipping. I wanted to spend a lot of time the last few days doing a big pass, auditing it, finding any potential issues, whether it's security, maintenance problems, things that aren't great to have in an open source project, that type of stuff.
I had it do an audit and it did a pretty good job. It found most of the things that Fable and GPT56 found. However, it did a great job when I asked it to start fixing the things and when I noticed how well it was doing, I started to push it a little hard. I opened with working on hardening lake bed for its first public release. I have the following PR up which addresses the majority of the remaining issues with the link to the PR.
Here's the report for the majority of those issues with a post plan link. Have we resolved them? Are there other things worth solving before launch? I had an issue with the cloud environments. I didn't even set up yet. I was doing this all remote. I think I was actually doing it from my phone if I recall. So, I set all of that. It went through the PR. It said that PR92 does not close the report. It closes two hard stops cleanly and one only partially.
Most of the launch gates are still open. It said explicitly what is resolved. Gave some caveats and then gave me a list of things that were still open and should not be treated as done. All of this is great. This is a really good way to process and synthesize and give me this info. It feels very opacy is how I would put it. So I end up having questions because it was a lot of text. So, I went through it all. I called out a couple different pieces.
I grabbed this section and I only cared about issues three and five because I had questions about what it meant by these things. And then after it had a different section with a lot more stuff. These types of things could be confusing for models because there was two lists that had a number three and a number five in them. So, I give a tiny prefix of which list I'm referring to. I thought this might trip it up because that's just a lot of context to get through.
And I'll be real, a lot of like the openw weight and frankly not Frontier stuff tends to struggle once you get it to this point. And it did a great job. It addressed all of my concerns very well and very directly. It called out the split release thing and described what it meant. Called out what it thinks I should put on security MD. And then it went through all of these different issues that I had questions about and helped me prioritize them.
I gave it a little bit more feedback on my thoughts there. But specifically here is where it gets fun. I called it I don't care about the token and URL issue but issue two which was bound public work. I said I like their ideas for it. Make a separate PR that addresses all of the concerns it raised. Then I asked it to do an investigation for this part and then I asked it to hold on to other things. It made two PRs because there was two issues it was told to address and it addressed both of them and did a great job.
It also answered my questions. Remember this this is a lot of different pieces I asked for here. as it do a PR for splitting up the actions earlier and a separate PR here for the public work bindings that it was mentioning and it was able to in just one run make both of those PRs and also address my other questions and give me a to-do list on all the things I have to go do before we go live. I noticed there were some comments on the PR so I asked it to look I just said both PRs have review comments to address and it addressed them all.
I then used cursor's built-in babysit skill to monitor both of the PRs and continue addressing things that come up on them, and it succeeded. I then asked another agent about these PRs, and they said a couple different things should be fixed. So, I literally just screenshotted it, both to test its visual capabilities, but also out of laziness. Pasted the screenshot and said, "Can you make these changes to 93?" And it did.
This is great. This is the type of back and forth and complex multi-target tasks and work that you could make most models do if you break them up really carefully and cleanly and try to not bloat the context. Even a model like 55 in my experience would get confused at some point during this and get too fixated on something from three messages ago instead of doing what I'm asking now. I had no ergonomic issues with four five here at all.
It was actually very pleasant to work with. I kept making my responses like worse and worse almost intentionally just to see where it would stumble and it didn't. Did it find everything Fable found? No. Was its code as thorough as 56? No. Was it able to go back and forth with me on really big, heavy tasks and be pleasant to use while also being very fast and cheap? Absolutely. Yeah. In a lot of ways, I see this model as a good alternative to something like Opus 48. and it kind of shits all over something like I don't know GLM52.
I have no interest in that model anymore other than the fact that it's openweight which is genuinely really cool and I appreciate them for that. But Grock 45 is a weirdly good default code model. So I wanted to push further and I did. I took a game I built previously, my little fish slop thing, and I asked it to make the game in 3D. It did and it has some issues like this layout is broken because of the model. I didn't choose for it to be panned hard to the side like this.
Wait till you see what happens when we open the game. It successfully made a full 3D environment, modeled all of the different creatures itself, as well as all of the geometry of the things on the bottom of the tank. And it did a way better job with that than any other model I've used. Even models like Fable and 56, which I have given access to 3D modeling tools as well. This crushed all of them. Obviously, the models are far from great, especially the creature models and the model for the submarine.
And it also got the controls super wrong where A goes right and D goes left for some reason. They Yeah, they screwed these things up. And it also just doesn't control very cleanly. I gave it a follow-up telling it to fix it, and it fixed the pointer and the clicking, but it didn't fix a lot of the rest. One of the models I was really impressed with in here though was the enemy model for the aliens that invade the tank.
It did a very good job with that. But yeah, I did not expect this to have three-dimensional taste. And I'll say outright, a lot of these, while far from like I would chip this confidently tier, are so far ahead anything I've gotten from any of the other models that I can confidently say this is the first model to be almost decent at 3D modeling in game engines like 3JS. Take it as you will, this is not bad. I I am more impressed than I expected to be regarding a Grock model's 3D capabilities.
I do have one last thought I want to discuss here, which is this idea of model generations. I do really feel like Fable's a new generation of model and I have a lot of thoughts on where GBD56 falls here. Video probably coming tomorrow depending on a lot. We'll see. The reason I bring this up is because Fable and 56 in many ways feel very different. And as silly as it is, even something like Sonet 5 has a bit of that feeling, too.
The difference is in the model's ability to orchestrate. They can step up a level and prompt sub agents and orchestrate big work into lots of smaller chunks to go longer, do more, and complete more difficult tasks. This is why I got so addicted both to Fable and to GBT56. I don't necessarily see that capability in Gro 4.5. My attempts to break up sub agents with it were admittedly limited, but I was not super impressed.
It didn't seem to have the same nuance in how it would break work up and delegate it and it would often get stuck as a result of a certain process it ran hanging and then not knowing how to clean up after. Some of this is harness specific, some of this is model specific, some of this is just being behind. But I would say in many ways what they built with Gro 4.5 is less a comparable model to this new generation with things like Fable and GPD56.
More it's really really impressive what they did with the previous technology. If you were to think of this in gaming, for example, they just put out the best PS2 game ever, but the PS3 has been out for two months. And as such, I'm very excited to see if they can level up to a PS3 generation game. And I think they can. They have all of the pieces. And the amount that they just jumped in such a short window is truly a sight to behold.
I don't think any lab has had a jump like this ever, other than maybe arguably like deepseek. going from forgotten in the conversation and just reselling GPUs to beating out practically everyone else in the space at the tier you're competing in for cheaper and more efficient models. It's impressive and we shouldn't be sleeping on SpaceX anymore. I posted in April that I legitimately believed XAI could have a crazy comeback.
I thought it would take 6 months to a year. I didn't think it would take two and a half. I am blown away. These guys are cooking again. I have no idea where this will all end up, but I'm thankful somebody other than Google is actually competing now. This is the first real player that Anthropic and OpenAI have had to be scared of in quite a while. And I hope they are. I hope the result of Gro 45 is faster, cheaper, and smarter models for everyone because that's what we want in the end.
So, congratulations to SpaceX AAI for catching up. I didn't think you had it in you, but clearly you do. I can't wait to see what comes out next. And until next time, peace nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.