Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Luke's Dev Lab · @lukesdevlab
Words
4,429
Runtime
25:34
Speaking pace
173wpm
Reading time
18min
173 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Like Yeah, this is by far the best I have seen on this test. Like it's just amazing. It's perfect. Hey, welcome to Luke's Dev Lab. My name's Luke, and I know a lot of you have been waiting for this video. So have I. So in this video, we're going to be looking at Qwen 3.8 27B. So I know that I'm going to be doing a series of videos on this model. So with this one, we're going to be running it through the usual
87 words, the words spoken in the first 30 seconds at 173 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 441 |
| Average words per sentence | 10.0 |
| Longest sentence | 58 words |
| Questions asked | 6 |
| Sentences containing a number | 40 |
Most used terms
Filler phrases
64 in total: like 46 · actually 4 · um 4 · uh 3 · I mean 2 · kind of 2 · you know 2 · sort of 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Like Yeah, this is by far the best I have seen on this test. Like it's just amazing. It's perfect. Hey, welcome to Luke's Dev Lab. My name's Luke, and I know a lot of you have been waiting for this video. So have I. So in this video, we're going to be looking at Qwen 3.8 27B. So I know that I'm going to be doing a series of videos on this model. So with this one, we're going to be running it through the usual tests like Blender, uh Godot, Sand Physics, et cetera.
And then after this video, I'm going to be doing it in a head-to-head comparison with the 3.6 version of the model in some different coding challenges. And then there'll be some more videos as well after that covering certain aspects of the model such as like quant comparisons, things like this. So if you're not subscribed already, definitely get subscribed for those videos. All right, well, let's not delay any further then.
Let's jump in. So the version of the model we're going to be looking at specifically in this video is the GG UF from Unsloth, and we're going to be looking at the Q4K XL version of this model. So I don't think you need much introduction on this model. We all know what it is. We all know what it should be able to do. So yeah, let's keep going. So our usual standardized testing suite for the model for this video. So we'll test the prefill and the decode speed.
We'll fill up the memory context and see how it does finding pieces of data throughout. We'll do some simple tool calling exercises in Agent C, then OpenAI Human Eval with 164 Python challenges. Kanban will be our front-end web development test. Sand Physics and Dungeon Crawler will test for algorithmic and mathematical capabilities. And then Blender and Godot for MCP connections with advanced tool calling. Quick look at system specs, so performance-related benchmarks.
This is the system I'm running on. So key takeaways are the RTX 2080 which has 16 gigs of VRAM, and then I've got 32 gigs of DDR4 RAM to overflow into. Now, this GPU is definitely not the fastest. This is a workstation GPU. So, if you've got a gaming GPU, even with 16 gigs of RAM, it's going to be faster than mine. Okay, so let's jump into our first benchmark here. So, this is performance. Again, this is run on my RTX 2080.
So, memory bandwidth is about 225 to 250 GB a second. I've got my prefill here, which is the speed at which the model ingest tokens. Now, on my prefill test, I do ensure that zero caching happens when this test is run. So, this could be why this seems slower for me than it is for you on similar hardware. But anyway, let's move on. So, I'm getting pretty slow prefill speeds at the moment on this model. And then if we move down to the decode speed, this is the speed at which the model will generate tokens.
I'm getting about three tokens a second. So, nothing to write home about, but if the model is intelligent, it's something I can live with until I get MTP up and running. And again, you're probably going to get better than this. My GPU is very slow, and of course, it doesn't fit entirely in my VRAM. Moving on to our memory benchmark now, so I've filled up the context of the model with 256K tokens, and then we insert a piece of data either at 0, 25, 50, 75, or 100% of the depth, and we ask it to find that piece of data three times over, giving us a total of 15 runs.
And as you can see here, the model has done perfectly. It's found the data every time at every depth. So, this is excellent for the memory. And then if we have a look at the outputs of the model here, they're pretty clean. So, this would be a perfectly clean output right here. And then some outputs are not perfectly clean. There's a bit of thinking padding around it, but overall, the model's done very well on this test, of course. 100% pass rate, so very good on the memory.
Moving on to our agency benchmark now, so this is some simple tool calling exercises where the model's put into a fake company sandbox. Now, you can find an overview of this in the GitHub down below if you'd like to read more about this test, but overall we've got a 96% pass rate, 47 out of 49 scenarios. This is pretty good and it's probably about as good as 3.6 performed on this test. So, yeah, it's good. And then for our final benchmark of the video, we've got OpenAI's Human Eval.
So, this is 164 handwritten Python challenges by the people over at OpenAI. So, I'll leave this GitHub link down below if you'd like to read more on this test or run it for yourself. So, jumping into the results here and this is quite promising. So, 85% pass rate doesn't seem amazing on the surface, but what is really good here is 99% answer rate. So, it's only failed to answer one question and this is a good indicator that the model is not overthinking.
So, yeah, I'm quite pleased with this result here. So, yeah, it's a pretty solid result here on this test. So, let's jump into some coding. All right, let's get into it. So, we're going to start off with our Kanban challenge. So, this is a good test of the model's ability for front-end web development. So, I'm going to pass the prompt over into our pie harness for today's video. We've got 128k context and the recommended settings for coding on the model.
So, I'll pass this in and then we'll check back when there's something to see. Ooh, all right. It's been a while. We are almost at 87% context used of our 128k, 130k. So, yeah, this model likes to be thorough, but I did not notice any thinking loops or anything like that while it was coding. So, it really is just putting in the effort, I guess. So, it says it's done. So, I'm going to list out the files it's created in the directory.
That's looking good. It looks like it's even done some testing in smoke test.js. So, we will now serve this up. Let's see what we've got. All right, look at this UI. Now, just going to make sure that we have no local storage from previous runs. So, I'm just going to give this a refresh. Okay, here we are. So, first impressions are it looks really good, for sure. This looks nice. There's a little overlap over there in the search bar with the search icon and the placeholder text in the input.
That's okay. So, let's test some functionality. So, let's try to move a card. Yep, that's working. Again, I love the UI. This is really nice. Let's move this column. Yep. That works really well. So, if I move this here, does it get a done state? It does. Reordering within the column as well, really good. So, we can turn on these filters at the top. I love that. That looks really nice. Assignees, so we've got a couple of assignees.
That's working. We can do a search here. SK is what I've typed there for the sketch, so that's correct. I'll type in the word fix. Yep, that's worked. So, let's try and add our own card. Okay, interesting. So, it comes up with this user interface here. So, I can do the title. I can choose a a tag here, and then I just press add. That's kind of strange, but now I can click into it. And we can do a bit more here. So, let's just do that and save it.
Okay, so we've got our new card. So, I guess the interface to create a new card is not perfect, but it's okay. You can see in the assignees, the new assignee has been added, and it works. So, let's try archiving this card. Goes into the archive. Yes. It says like where it came from, the date. Restore, that's worked. And then, if I delete it, we get a message saying it's going to go permanently deleted. We do that. Done.
Um okay, what else? So, let's change a column name. So, it looks like we've got an inline editor there. That's nice. That's done. Let me just refresh, make sure this is persisting. It is. And then, I guess uh we should create a column two. So, that works. We can move this over here. We can change the name of it. We can archive the column. It can be restored. Let's see if it's restored. It is restored. So, let's archive it once more and then delete it.
And there it is. Done. I think that is everything. And yeah, okay, I got to say like it took a long time to get here, a lot of tokens, but the result is fantastic. Like this is brilliant. This is one short, nice UI, nice functionality, just a great UX with all the, you know, you see this pattern here showing you where the card's going to go. Same when you move a column. It's just Yeah, it's a brilliant result. I'm very impressed.
So, this is a really good start for 3.8 27B. So, let's move on. So, we'll jump into the sand physics simulator next. So, let's give the prompt over to the model, and we'll see how it gets on with this one. Okay, we are back on sand physics. And as you can see again, 75.5% token usage is a hell of a lot for this challenge. A hell of a lot. Um we can see why. So, I'm going to show you. So, here it decided it would run some tests.
Now, I don't love that it's written this in a temp folder in the root of my system. That's not good. It should have wrote this in the project directory. But, this is why it took a lot longer because it decided to write a testing harness and run through a bunch of checks. So, it had 26 assertions. It discovered some issues that it wasn't happy with. So, again, it's very thorough, this model. Very thorough. So, you can see all the changes it's been making. 10 out of 10 runs green.
Like, it's yeah, it really goes into it. So, this is why the model was using so many tokens because it really checks everything. But, let's see. Let's see what it's done. I haven't looked at it. So, confirm the presence of the file. Yep, that's there. So, let's serve this up and let's have a look here. Oh, okay. So, stuff is just come in already. So, I'm seeing textures on things. So, you can see we've got like a a cross here.
Let me zoom in just a bit more for you guys. So, there we go. So, we've got a cursor. We've got like a circular Okay, so this circle presumably, yes, it shows the brush size. I've not seen that before. That's really nice. So, I'm loving the textures. It looks like there's numbers on these buttons, too. So, yes, if I press numbers on my keyboard, it does switch over to them. So, that's really cool. Let's go for It's really smooth, too.
Like, so smooth. Let's go for a large brush here with the water. That's really nice. And then let's place some acid. So yes, it dissolves in materials as it burns through them. So that's really nice to see. And then it settles on the water there. You can see how it's like a See, the acid has like a glowing effect. It does. The acid has like a glow effect to it. Look at that. I've never seen that before. You can pause it and regime?
Wow, again, I've never seen that before either. The the detail is just amazing. Like wow. Um what else have we got? So C to clear, space to pause, the um angle brackets for the brush size. And then again, like you can see the circle here in the canvas getting bigger and smaller to show you the size. Like yeah, this is by far the best I have seen on this test. Like it's just amazing. It's perfect. Yeah. So again, same as before, so many tokens used for such a great result.
So let's move on to dungeon. I'm excited to see it. So here we are with the prompt. Let's pass this over and I am keen to see how Quen does, but probably going to take a long time. Okay, here we are with Dungeon Crawler and not as many tokens on this one. So we're only up to 42.5% and I was seeing the same behavior as before. So let me find it. So again, here, don't love it that it's done it in temp at my root, but it's written a test harness again and it's gone through testing and making sure things work, it seems.
So that's definitely the pattern that we're seeing in this model is how it builds a harness and tests what it's done. So, let's check the directory. We have an index.html. So, let's serve that and here we are. So, let me just zoom out a touch. Yep, there we go. So, Dungeon Fog BSP generation, Bresenham line of sight, 60 by 60 grid, radius nine. So, I guess that's nine pixels. So, we can move. I like how there's like a glow around the player tile, as well.
That's nice. So, fog of war is obviously working. We can move. The ray casting looks pretty solid as we pass the doorways. Let's turn off god mode and look at our dungeon. Not too bad. Let's do some map generation. So, everything's connected. That's absolutely correct. And then, we've got an explored percentage down the bottom there, as well. Again, that's Don't think I've ever seen that before. I mean, should I be surprised at this point that it's done a perfect job with one shot and done extras that I've never seen before?
I mean, it just delivers. Again, what a great job that it's done. And it didn't actually use a great deal of context this time. Yeah, brilliant. Let's move on. So, let's move on to the MCP portion of the video now. So, I'm definitely excited for this. We're going to be starting with Blender. Now, in my last video, one of you who actually knows what they're talking about with Blender commented that using the latest version of Blender might not be the best idea because these models don't have up-to-date training data, which makes perfect sense.
So, I've bumped down to a version where they recommended. So, we're on 4.1 for Blender here. I did manage to find an MCP tool that works with this version of Blender. So, we've got Blender MCP here. I'm going to connect that now. And so, we've got a fresh project. We've got a prompt here, or rather a project brief, I would say. And the gist of this is that we're going to generate a couple of assets, a marble and a finished gate.
The marble will typically be quite simple, and a finished gate can be quite complex. So, let's give this over to the model, and then we'll come back when there's something to see. Okay, we're back with Blender. I want to quickly show you 96.4% on the context because it's about to compact now, so that's going to go down. So, it's used up almost the entirety of the context get this done, but luckily it didn't need to compact in progress, so there shouldn't be any sort of degradation of what the model's working with here in terms of its memory.
So, it's summarized with the deliverables that it's done. Apparently, there's a screenshots folder with six PNGs inside. So, let's check the deliverables first. So, we've got our folder here. We can see there's a blend file in here. We've got our exports, so it's exported. And now we've got some screenshots. So, woah. Okay. Wow. That's amazing. I guess that's the marble inside the floor, but it's it's like taken shots of the gate at like multiple angles.
Like, damn. So, yeah, there it is. Wow, look at that. That is by far the best I have seen on this test. Oh, my Blender is lagging, but like, damn. That is really good. Like, really, really good. This model really is a step [snorts] up, for sure. This is fantastic what it's done. So, the story continues. A lot of tokens get burned up, but the results are there. It delivers. Like, that's amazing. Yeah, what can I say? Like, another excellent job.
So, let's move on to a game engine. So, let's move on to our game engine, which is Godot. So, this is going to be our final project brief that we're delivering to the model in this video, our final test. So, normally, the intention was after the Blender assets were generated, the model would then generate a game using those assets. But, for the sake of comparison, I'm going to stick to the usual Godot challenge that we do.
But, in a future video, I think we will see if it can deliver our original marble run game using those assets. So, definitely get subscribed if you want to see that in a future video. But, for now, this prompt is essentially a simple 3D environment where the player can jump on platforms, collect a few orbs, and then reach a finish tile. I've got Godot open here. MCP is connected with a fresh project. So, yeah, let's give this over to the model.
I'm excited to see how it does. Okay, guys. I just want to jump in on how we're doing here. So, we've already used up the entirety of the context and then compacted. Now, we're up to almost half the context used again. And it seems to be getting hung up on thinking that the player can't jump. And the player can jump because I've tested it. I even interrupted it at one point said, "Yeah, jump works. It's fine." But, it still thinks the player can't jump.
So, I'm going to tell it again. So, I've tested the game. Jump works. Sprint works. Movement works. Orbs can be collected. Finish tile can be reached, triggering an end state. The only thing is, can't look around. So, let's see if the model can do that, and then we'll check out the game together. Okay. So, here we are with the model now in Godot. So, after I jumped in and told it, it was able to work through, fixed the issue, and has now verified the game is up and running, and it's good to go.
So, we are up to 79% context usage after a compaction when we reached 100% already. So, let's head over into Godot, press play, and launch our game. So, I'm just going to enlarge this window for you. And here we are. So, it's a little bit laggy there, but it's not too bad. My system's not that strong. So, we can move around. We can look. We can sprint. That works. V changes our FOV. That works. We've got collision. We can collect the orbs.
So, in the top left corner there, it says orbs one out of five. It's not easy to see on the background. There we go. On that darker background, it's easier to see. This is the first model I have seen where you're able to collect the orbs. So, this is This is exciting. So, let's get to the end. Goal reached. Orbs five out of five. Goal reached. And actual finish state. And again, as well, this is the first time I've seen a finish state on this game as well.
So, yeah, this is the best I've seen this done. Which is kind of a theme for this model in this video. So, let's conclude. So, that brings us to the end of this first video looking at Qwen 3.8 27B. And I got to say, I'm pretty impressed. So, let's review. So, on performance, it's pretty slow on my hardware. There's no way around it. It's uh you know, it's 27B. It's dense. There's no MTP or speculative decoding at the moment on my setup.
So, it's to be expected. Then on memory, it got a perfect score, and it had mostly clean outputs as well. So, very good on the memory in the 256K context we used for that. On agency, it scored 47 out of 49, and it was pretty solid. I would say it's about in line with other models in this weight class on this test. Then looking at our human eval, so it did pretty good here with 85% but the standout was that it answered 99% questions.
So, 163 out of 164 were answered. So, that's really good. And then we jump into our coding test. So, Kanban, 86.8% on the token usage. It took a long time but when it got there, it was brilliant. What it delivered was clean, had a very nice UI, the user experience was really good in the way you moved things around, you could like visualize where the card was going to be moved to. There were no bugs, there were no issues, it was a one-shot.
It was really good. On to sand physics, so we were at 75.5% of our context used for that. And again, this is the best result I've seen on this. So, the way the acid was glowing, the textures that it put in place, the keyboard shortcuts, the way it had some materials already placed as like a demo layout. It added pause functionality which I've never seen before. Like if you showed me all the sand physics results that I've done on this channel so far, I would look at this one and say, "Oh, that one's probably done by a frontier model." Because it it's just so much better than anything else.
And then moving on to dungeon, so this one actually didn't use a lot of tokens like I was expecting it to. So, it got to 42.5% of our context used. Again, a one-shot, everything was done. Nice little extra details as well like there was a glow around the player and you could see like the percentage of the dungeon explored. I've never seen that before on this test. And it wrote a test harness as well to check that things were working correctly.
Also did that on sand physics, too. So, another great result there. And then moving into MCP, so starting with blender, we got to 96.4% of our context used, no compaction needed. So, we were right on the limit there. And again, the best gate asset I have seen by any model on this test so far. It was it did a very good job. It took screenshots of the scene from like multiple angles, which again, never had that either.
So, just really good result in Blender. And then finally on to Godot. So, Godot was where we could see the model was being pushed here. I did have to jump in a few times in Godot, unlike the other tests. So, with Godot we we hit 79% context used after reaching 100% and compacting. So, a lot of context text was used on this test. But then again, a great result was achieved in the end. So, you could collect all the orbs.
You could go to the finish tile, and it would trigger an end state, which I've never seen before on this test as well with any other model. So, it's just same story really. The conclusion I would have on this model is it takes a lot of time and tokens to get to a result, but the result is brilliant. Like, it's the best model I've tested so far for local coding without a doubt. Like, absolutely it's the best. So, as long as speed and performance is not an issue for you, 100% use this model.
It's as simple as that. It's brilliant. I think the Qwen team has done a really good job on this model. Like I mentioned earlier, this is not the only video I'm going to be making on this. I'm going to be testing this. I'm going to do some comparison against the 3.6 model. It's not even going to be a competition, I don't think, but we'll just see the difference in how it is against the generational leap. And then I definitely want to compare this against Meta's Music Gen cuz I also thought that was pretty good.
If I gave you the hot take on it right now, I would say this model does a better job than Music Gen, but it uses a lot more tokens. So, if you're token constrained, there's merits to use Music Gen still. And then some of the videos you guys have been asking for, such as like quant comparisons, harness comparisons, actually been holding off on these videos to do it with this model. So, I'm going to start working on these videos too now.
So, you you're going to see a lot more content from me on this model coming up and hopefully if Quen does a 35B version as well, that would be really nice to see. So, definitely get subscribed if you want to see all the upcoming content on this. So, thank you for watching. A huge thank you to the people who keep supporting me on YouTube and on Patreon. You guys are amazing. Thank you so much and just thank you to you for watching end of this video.
I really appreciate it. So, give it a thumbs up if you enjoyed it and subscribe if you want to see more content like this. Bye.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.