Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Matthew Berman · @matthew_berman
Words
2,820
Runtime
16:50
Speaking pace
168wpm
Reading time
12min
168 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Google is back. They just dropped Gemini 3.8 flash and it's actually good. And that's not the most impressive thing about it. So, I'm going to show you the benchmarks in a second, but I just want to give you a little bit of context as to where Google is currently in the competitive landscape. So, a few years ago ChatGPT came out and Google was nowhere to be found. ChatGPT took off, OpenAI was doing so well, and of course Google had to
84 words, the words spoken in the first 30 seconds at 168 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 224 |
| Average words per sentence | 12.6 |
| Longest sentence | 72 words |
| Questions asked | 4 |
| Sentences containing a number | 55 |
Most used terms
Filler phrases
50 in total: actually 16 · kind of 10 · like 10 · uh 6 · basically 4 · you know 4.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Google is back. They just dropped Gemini 3.8 flash and it's actually good. And that's not the most impressive thing about it. So, I'm going to show you the benchmarks in a second, but I just want to give you a little bit of context as to where Google is currently in the competitive landscape. So, a few years ago ChatGPT came out and Google was nowhere to be found. ChatGPT took off, OpenAI was doing so well, and of course Google had to drop everything and try to catch up.
And then at a certain point, about a year and a half ago, they released Gemini 2.5 Pro, which was at the time a phenomenal model, the best model on the planet, the first model to create a full Rubik's Cube simulation that actually worked. But, it was a fleeting moment for Google. Suddenly, Anthropic and OpenAI were just far ahead. Fast forward to the last 6 months, Google has released a number of models in the Gemini family that, to put it plainly, have fallen a bit short.
The benchmarks have been pretty good, but the actual usage of the models has definitely left something to be desired. But now, Gemini 3.8, fresh off the release of Gemini 3.7, which really was just a few weeks ago, we have a brand new Gemini model and it's actually quite good. Now, it's really good in some benchmarks, kind of okay in others, but the cost is what makes it special. All right, so here are the benchmarks.
Number one, the most important benchmark in my opinion today, Deep Seek. And this test long horizon software engineering task. It seems to be the most accurate reflection of how actual engineers using these models feel about the models. For example, Fable on many benchmarks is way ahead of GPT 5.6 Soul, but on Deep Seek, it's really close, and that's my experience with these models. Soul is fantastic, Fable is fantastic, but they're not that far apart.
And now, check this out. Deep Sweet V1.1 Gemini 3.8 Flash coming in at 73.7%, which is effectively even with Claude Opus 5, which was just released a few weeks ago. Here is GPT 5.6 Soul coming in at 72.7% under Gemini 3.8 Flash. So, this looks to be a really good model on the benchmark that I respect most of all. Then, we have GDP Val. This is a benchmark from the OpenAI team testing the models on real-world knowledge work.
PDF extraction, data analysis, presentation creation, all of that kind of stuff, and it's coming in at 1545, which is not that good. So, Claude Opus 5 absolutely dominating the competition coming in at 1824, second place 1710 for GPT 5.6 Soul, and then kind of a much less good score of 1545 for Gemini 3.8 Flash. So, on GDP Val, knowledge work, real-world knowledge work, it's just okay, basically equivalent with GPT 5.6 Terra.
However, if you look at the pricing, that's the model size that this new 3.8 Flash is trying to compete against. So, you really shouldn't compare it to Soul and to Fable. It doesn't really make sense to compare it to that cuz it is a fraction of the price. And speaking of price, let me just show you the price before I show you the rest of the benchmarks. So, input price, 75 cents per million input, output price, $3.75.
Really, really inexpensive. Look how it compares to Opus 5. It is a fraction of the price. $5 and $25. Here's GPT 5.6 sold $4 and $20. GPT 5.6 Terra $2 and $12. So, again, like 20-30% of the price of even Terra. Now, the 75 cent $3.75 prices is, if you look at the fine print, they're introductory prices which expire at the end of this year. I don't love that they put this in bold right here. So, it's 75 cents and $3.75, but in parentheses in small text under it is the actual price that we should expect after this year.
Now, of course, if everybody is clamoring and loving this model, they may just extend the price indefinitely. But, if you already plan on raising the price in the future, you should make that price the more prominent price. But, even at $1.50 and $7.50, uh it's still cheaper than a comparable Terra model. We have Harvey's legal benchmark absolutely dominating. Number one, 61.4% and number two was Gemini 3.7 Flash. So, for legal work, this is the best model on the planet and that's where I want to spend a minute.
This model, and seemingly a lot of these newer models, are really good at certain things and not so good at others. So, if you're an enterprise company and you're thinking about which model to adopt for your business, it's not as simple as just saying, "Okay, what's the latest from Anthropic? What's the latest from OpenAI?" You should actually look at the benchmarks and, more importantly, test them yourselves. Have your own internal benchmarks that you create, which it's not very simple to do, but it's very doable, and you can test all of these new models against your own benchmarks and choose the right model for the job.
So, if you're a law firm, you should definitely look at this model in particular because it scores so highly on the Harvey legal agent benchmark and also is very cheap. Then we have terminal bench 2.1. It got the number one score at 89.4. This is agentic terminal coding, one of the most important benchmarks for knowing how well a model can code. And then we have terminal bench 4.0, the newer version, and look at this.
It actually got quite a low score. So, I don't know exactly what changed between 2.1 and 4.0. Okay, so I did a bit of research and terminal bench 2.1 was all but saturated, as you can see here. And so, terminal bench 4.0 is really a brand new benchmark under that family of benchmark, a much more difficult benchmark, and that's why we're seeing lower scores. However, look at this, Claude Opus 5, 51% absolutely dominating the field.
And then we have, of course, Gemini 3.8 coming in at 19.1%, which is just okay. It beats Sonnet 5 and comes in well under Opus 5 and Soul. Here is Humanity's last exam, and surprisingly, Gemini 3.8 Flash comes in at number one at 55.9. So, this is why it's a little bit weird. Some of these benchmarks, Gemini 3.8 Flash just dominates at and is so cheap, and then on others, it's just okay at best. Here's OS World, which is agentic computer use, the ability for the model to control your computer, to control your browser, coming in at 59%, a respectable score, but Opus 5 just dominating at 75.
But, this is the chart that really matters. This is Deep Sweet, again, but importantly, it also maps the average cost per task. This is very important because the average cost per task not only takes into account the actual cost per million input and output tokens, but it also takes into account how many tokens are used to solve the task. Because if one model is half the price of another model, but takes twice as many tokens to complete the same task, it is effectively the same price.
So, here's what we see. This is Gemini 3.8 flash, and really where you want to be on this chart is as high up and to the right as possible. And that's what we're seeing. It's a good model, and I do very much trust Deep Squeeze. So, I'm looking forward to using it. Now, they released a second model in the 3.8 family. It is Gemini flash 3.8 cyber, and it is specifically exactly what it sounds like, built for cyber capabilities.
And so, it is only available to what they call trusted defenders via the Fairwind program, but it's basically the model without some of these cyber guardrails that come in the kind of normal Gemini 3.8 flash. And what we're seeing here, this is the Cyber Gym benchmark, and it performs really well. Here's Mythos 5, although Mythos 5.1 is now out. Here's GPT 5.6 Soul, 83%, 83%, even GPT 5.5 Cyber, a model specifically made to be really good at cyber attacks and defense, 85.6, and Gemini 3.8 flash cyber 86.2.
Now, unless you're part of that program, you're not going to get to test this model, unfortunately, but you can apply to the program and see if you get in. Now, they did something interesting which I haven't seen other model companies do. Listen to this. To better capture real-world defensive needs, which are not limited to just C, C++, codebases like in Cyber Gym, okay, so Cyber Gym is the benchmark, and they only have tests built in C and C++, we also evaluated Gemini 3.8 flash cyber against a comprehensive internal benchmark in which the model has to discover a wide range of vulnerabilities across complex codebases spanning 20 programming languages, and it does really well.
Here's Gemini 3.8 flash, Gemini 3.7 flash and 3.5 flash. Now, they did not test other companies models on this internal benchmark, but on 3.8 flash it is a massive jump versus the previous iteration version at 3.7 flash. All right. Let me show you those tests now. So, over the last week and a half I've tested GLM 5.3, Fable 5.1 and GPT 5.6 Soul against the same set of demos or tests and this is the first one. This is creating seven individual 3D low poly biomes and Alex put the other ones on the screen, please, so we can actually compare them.
Here is Gemini 3.8 flash and I'm quite impressed. The beach looks really good, the farm looks really good. I'd say the overall detail is definitely less than GPT 5.6 Soul, which so far has been my favorite. Fable 5.1 is probably second place and then GLM 5.3 and Gemini 3.8 probably right around the same judgment. Now, very similar to Fable, we have this weird kind of glitchy thing right here and then we have a bunch of little fish in the ocean under the singular boat.
I don't know why that happens, but we've seen that before. Here's a little mistake. Look how the water under this ice lake is kind of protruding out from the biome right here. So, yeah, overall it's pretty good. It's definitely not the top of what I've tested over the last week and a half though. All right, next I asked it to create simple web pages for different products, an Apple, a DJI Spark, a rubber duck company, the Galaxy Z Fold and a Tesla Model Y and again, please Alex, throw the others up on the screen.
Here's the Apple website. So, it looks okay, very simplistic, no Apple above the fold. Yeah, this is definitely one of the more simple websites that I've seen. This is kind of weird. It's like a checkout, but it gives me no options to actually select items to put in my cart, and it pre-selected six slots. This is definitely the worst website I've seen. It's not terrible, but compared to the others, this is not as good.
Here's one for the DJI Spark. This one actually looks quite good. A little clipping on this G right here, but otherwise, the colors look good. It's kind of interesting how all of these models get the colors right, cuz these are the colors of Nvidia, and specifically the DJI Spark. So, a bunch of good stats. Let's see if we can change it. Yep, we can change it. We can see all the other numbers updating. We have a little terminal right here.
I think this is pretty good, comparable to all the others. Here we have a rubber duck company. This is definitely not the best. I really like the more kind of childlike look of it, where we actually see a rubber duck. This is more of of just like a realistic duck emoji. If we scroll down, yeah, I don't know. Astro Duck Voyager. Yeah, none of this really makes a lot of sense. I'm going to put this at the bottom, also.
This is definitely not that good. Here's one for the Galaxy Z Fold. I guess that's a foldable. Am I supposed to click this? Here, let's see if I can Okay. Oh, that's actually kind of cool. So, I can drag this and actually open and close the foldable, although, you know, it doesn't look anything like a phone. And in fact, if I close it, you can see you can see the back of what's supposed to be on the screen on the inside of the phone.
Yeah, this is this is not good. And then last, the Tesla Model Y. Once again, a model tries to recreate the Tesla Model Y, basically looking like it's using Microsoft Paint, and it is really bad. But, you know what? The other ones were really bad, too. So, I can't really blame it. Uh we are able to design our Model Y. We have the Performance All-Wheel Drive. We have the Long Range All-Wheel Drive. We can change the colors.
Let's see if it changes the colors up here. It does. There's our little preview of our car. We can say how many miles we're going to drive per day and our average estimated 5-year gas savings. This is actually pretty nice. Oh, this is super nice. So, vision-only neural autopilot. This is a pretty clean animation to show that off. Obviously, like very simplistic, but definitely uh better than what I've seen elsewhere.
So, I'd say this is probably comparable to the other ones. Overall, I'd say this probably falls right about or under where GLM 5.3 was. All right. So, for the next test, I asked Gemini 3.8 Flash to create a PowerPoint deck all about data centers. And this is what we have. Now, the nice thing is it knew to use my Forward Future branding. These are the Forward Future colors. forwardfuture.com again. So, that looks really good.
This is our typography, uh how the physical backbone of the internet, the cloud, and AI actually works. Okay, first page looks good. Uh second Yeah, it's it's okay. It doesn't have strong design chops, but like the structure's there. The information looks correct. But, certainly if I'm going to create a PowerPoint, and especially if I'm focused very much on design, I'm probably going to a cloud model. All right. Alex from my team just sent me this link.
He ran a test with Gemini 3.8, and the purpose is to create a 3D topographic map of Mount Everest. And this is phenomenal. I I am super impressed. So, I can drag it around. I can zoom in. We have different points. I have different sliders. Oh, nice. So, I can get a cut angle straight through Mount Everest. Crustal depth offset. Ooh, that's really cool. Solar azimuth. I hope I'm pronouncing that right. I've never even heard that word before.
Yeah, very cool. Vertical exaggeration. Ooh, yeah. Yep, yep. This is really cool. This is a very very good topographic map. Here's a 2D projection. Yeah, wow, this is really cool. I'm very impressed. Oh, and look at this. We actually have a bunch of different points on Mount Everest. Here's Camp 2, for example. And so, if we switch back to the 3D terrain, there it is. Yeah, so this is really cool. This is an absolute winner.
I have nothing to compare it against, but in my mind, this is great. All right, and Alex just sent me another game that I was completely just playing in a distracted way, but I'm going to show it to you now. This is a very simplistic Doom recreation. This is basically just one prompt. Uh let's see if I can find a baddie to kill here. Here we go. Yeah, so you know, you can see it. It's very simple, but with a few more prompts, I think this is actually decently fun to play.
Yeah, I you know, pretty darn good. And real quick, I just want to tell you to go visit forwardfuture.com. That's our newsletter, and if you want to stay up-to-date on all the latest AI trends, new model drops, go there. forwardfuture.com. Check it out. It's free. It's awesome. So, congratulations to Google. This is definitely a good model. It's not quite the frontier, but for the price, extremely competitive. And you can actually see how it compares to Fable 5.1, which I covered yesterday.
Go check out that video right here.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.