Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
7,265
Runtime
38:32
Speaking pace
189wpm
Reading time
30min
189 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
AI never sleeps and this week has been absolutely insane. Your waifu is finally getting delivered. Xiaomi releases their latest and best model MIMO 2.6. XAI also releases their best model Gro 4.7. OpenAI drops their latest model GPT6 Soul and Luna. And then Enthropic also drops their best and latest model Opus 5.5. And then shortly afterwards, Google also drops their latest and best. Oh, never mind. It looks like they've only released a flash texttospech model. We have some new interactive world generators that actually stay consistent the whole time. Flux releases an
95 words, the words spoken in the first 30 seconds at 189 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 483 |
| Average words per sentence | 15.0 |
| Longest sentence | 71 words |
| Questions asked | 10 |
| Sentences containing a number | 117 |
Most used terms
Filler phrases
116 in total: like 61 · basically 22 · actually 19 · kind of 9 · I mean 2 · you know 2 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
AI never sleeps and this week has been absolutely insane. Your waifu is finally getting delivered. Xiaomi releases their latest and best model MIMO 2.6. XAI also releases their best model Gro 4.7. OpenAI drops their latest model GPT6 Soul and Luna. And then Enthropic also drops their best and latest model Opus 5.5. And then shortly afterwards, Google also drops their latest and best. Oh, never mind. It looks like they've only released a flash texttospech model.
We have some new interactive world generators that actually stay consistent the whole time. Flux releases an open robotics model. Google shares their plan of putting data centers in space. We have a ton of new open-source models and a lot more. So, let's jump right in. First up, we have a new interactive world model called World Crafter by Tencent. Like other world models, WorldCer also lets you create basically an interactive 3D world with just a text prompt or a single reference image.
Now, the clever part is that it has implicit 3D aware memory. So, even if you look away and then you look back, things remain consistent the whole time. And how this works is in addition to just generating the video, it also reconstructs a 3D point cloud of the scene so that things remain consistent the whole time. It's like it's storing a 3D model into its memory. So, this gives it a much better understanding and memory of what's actually in the scene.
So, here are some additional examples for reference. You can see this can generate worlds in a ton of different artistic styles. You can also control different characters. It can be first person or third person. This is a very flexible tool. Now, at the top of the page, they released this already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer.
Note that they released two different models. There's a base model and a distilled fast model. The model size is pretty huge, so it's around like 147 GB for all these components. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a new and super powerful image model called Ming image 0.1. This is an open model built specifically for visual design. So this is a 6 billion parameter textto image model for things like UI, infographics, posters, and other visual designs with a ton of different text and elements.
So here are some impressive examples for your reference. It supports up to 2K resolution. And one particularly useful feature is that it can also generate images with transparency. In other words, it has an alpha channel output. So you can generate an object or design element with a transparent background and then add the design somewhere else. So the main model is called Ming image 0.1 design, but they've also released a second model called Ming image design layer which kind of does the opposite.
So it's meant to take a full flat design image and you would prompt the model on what you want to separate and it would output multiple transparent layers of separate things which you can then edit individually. So a super powerful tool and if you look at this UI UX design leaderboard by artificial analysis then at least in terms of open weights models you can see that Ming image design is currently ranked number one.
So, if you're looking for the best open model to generate infographics or user interfaces, different diagrams, posters, or basically anything that contains a ton of elements or text, then this might be the best model to use. And the awesome thing is this is under the MIT license, which is very permissive. Note that for the base Ming image design model that is 12 GB in size, this should be able to fit on most mid to high-end GPUs.
Now, this design layer model, this is where you would take an image and break it down into separate transparent layers. This is a bit larger at 24 GB in size. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have a new 3D scene generator called Mirror Scene. This is able to take in any image and reconstruct it into a 3D scene with separate individual objects like this.
So how this works is you would feed it an image and mirror scene first identifies the objects in the image and then estimates the scene's depth and then generates the 3D geometry of each object. It then reconstructs their positions and sizes so everything lines back up with the original image. And this is important because afterwards you can then take the scene and edit objects if you want or you can also import the scene into Blender and animate objects like this.
Or you can even plug it into a physics simulator to train robots like this. The awesome thing is they've released everything already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to use this. Everything is like 6.1 GB in size, so it's fairly tiny. This should be able to fit on most like mid-consumer GPUs. Plus, they've also released the data set to this, which is fantastic.
If you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have a new robot model called Flux 3 Action. Now, in the past, the Flux team has released a ton of generative image and video models. So, it's interesting that they're kind of taking a completely different direction here, and they're releasing a robot model. Now, how this works is it's basically a 7 billion parameter world action model designed to control robots.
You would give it a goal along with visual information from what the robot can see using its cameras and also information about its joint positions. And the model can predict both what the robot should do next as well as what the world would look like after that action happened. So if you tell a robot armed to put an object somewhere else, Flux 3 action can look at the scene, predict a sequence of motor commands, and get the robot to do those commands and then look again and keep correcting itself until it achieves your goal.
Now, the cool thing is this doesn't have to just control robots because this can determine what action to take and what would happen next. You can also get it to play a ton of video games like this. And as you can see, if you compare Flux 3 Action against another similar model, Cosmos 3 Nano, not only is Flux 3 way smaller, it has like less than half as many parameters, but it also was able to achieve a higher success rate at a much faster speed, making this super efficient for running robots in real time.
The awesome thing is they've actually released the model. So over here, if you click on download the weights and you scroll down a bit here, it contains all the documentation on how to use it, plus how to fine-tune it on your own robot or video game or other devices. And at 7 billion parameters, the base model is fairly tiny at only 14 GB in size. If you're interested in reading further, I'll link to this main page in the description below.
Also, this week, Google releases their latest texttospech model, Gemini 3.8 Flash TTS and Flash Light TTS. The awesome thing about this Gemini TTS model is you can design a completely new voice just by describing it in a text prompt. So, here are some examples. Here's a high energy DJ voice from Melbourne. >> All right, Melbourne. So good to have you with us tonight right here. Heaps of great music on the way for you.
But first off, something that's going to absolutely pump up your weekend. I mean, seriously, this one is straight from the dance floor. So, get those feet up, get those vibes rolling. We've got this incredible new track coming up, and I know you're going to dig it. I mean, it just rips. Let's get it going. >> Over here is a Japanese dragon. And the nice thing about the Gemini TTS models is that they allow you to add metatags within the transcript so that you can create some really natural sounding and expressive generations.
Here are some examples. >> Hang on a second. Let me find that order status. Yeah, this order never really got delivered. Sorry for that. Let me update and reschedule this order. Give me one moment. I just updated the status and rescheduled the delivery. >> Or here's another example where you can describe exactly how you want the transcript to be read out within metatags, giving you really granular control. >> Maybe it's just a deer. >> That's not. >> People say it's always deer. >> It's not a deer.
It's It's staying right on the edge of the light. >> Stop. >> It's following us. Look right there by the pine. >> I don't see it. >> There. >> Oh my god. What is that? >> What is that? >> It's too big. That's not an animal. >> It's just standing there. >> It's just It's just watching us. >> We need to go. >> Okay. >> We need to go right now. Don't >> Don't run. >> Don't run yet. Just walk fast. If we run, it'll it'll what? >> It'll chase.
Just keep your light on the path. Okay. Oh god, it's moving. >> It's coming faster. I can hear the branches snapping. He's right behind us. No, no, no. We have to go. We have to run. Now, here they've released two different models. There's the main flash TTS, and then there's also a flashlight TTS. And this is built for high volume, costefficient scale. So, if you need to do a ton of dubbing or create a ton of audio content, here it says this supports more than 100 languages and dialects.
And if you look at these benchmarks in terms of text to speech quality, then these new Gemini models are state-of-the-art. On average, they even beat 11 Labs. And then in terms of voice design, again, that's like creating a completely new voice from a text prompt. This is also much better than other voice design models. If you look at this voice arena leaderboard where people can blind test different AI models, then you can see that this new Gemini 3.8 Flash is in second place in terms of pronunciation robustness.
Then you can see that Gemini 3.8 Flash is the current leader. Now given its quality, this is also insanely cheap. So this new Gemini 3.8 Flash is located over here in terms of pricing, making it very affordable. Now, currently this is already available via API and Google's AI Studio, which you can try out for free, and it's also available in Gemini Notebook, which was previously called Notebook LM. Definitely one of the best and cheapest texttospech models you can use right now.
If you're interested in reading further, I'll link to this main page in the description below. Now, last week, I showed you this dashboard where Xiaomi shared their reinforcement learning process for their upcoming model, Mimo 2.6, in real time. So you can see the training runs and the progress of the model over here. Well, they just finished training and this week they have released the MIMO 2.6 model and its results are pretty state-of-the-art in terms of open weights models.
You can see for deep suite the pro model is edging very close to GPT6 Astra and it even beats Fable 5. For these other agentic and knowledge work benchmarks, you can see it's also very close to Frontier. The craziest result here is Cyberjimy where it achieved like over 95% much higher than the other Frontier open models. It's super capable at vibe coding different 3D environments and games. As you can see from these examples, if you look at this intelligence index by artificial analysis, then you can see that MIMO 2.6 Pro is actually now the leading open model out there, beating the previous leader GLM 5.3 by just one point.
The cost is also ridiculously cheap. So, MIMO 2.6 6 is all the way over here, costing just 13 cents per task. Even cheaper than GLM 5.3 Flash as well as the latest Deepseek 4.1 Flash. In fact, this is basically the new frontier in terms of cost effectiveness. The awesome thing is they've released everything. So, if you click on this hugging face link, here are all the models for you to download. Note that the Pro model is a trillion parameters.
So, this is quite huge as expected from a Frontier model. And the total size of this is around 573 GB. The flash model is a bit smaller at only 311 billion parameters. And this is only 178 GB. Now, in addition to these two MIMO models, they've also released MIMO desktop, which is like the desktop app or harness that uses MIMO. It's kind of like codeex for GPT or cloud code for claude. If you're interested in reading further, I'll link to this main page in the description below.
Also this week, one of the other major AI labs in China called Stepfun releases their latest model, Step 5 Preview. This is also an incredibly powerful multimodal agent. So this has 600 billion total parameters. This is a mixture of experts models. So when you use it, only 27 billion parameters are active and like most Frontier models, this supports a 1 million token context window. From its deep SWE score, you can see this is pretty good.
Same with these other agentic coding and knowledge work benchmarks. It seems to be on par with or even better than the other open frontier models like GLM 5.3 and Kimik 3. If you look at this leaderboard by artificial analysis, then you can see that step 5 is still one point behind GLM 5.3. But it is almost three times cheaper than GLM 5.3 and also way cheaper than the closed models. Because it can also understand visual information, you can get to vibe code a ton of stuff really well.
For example, I can just input this photo and get it to render this in 3D. It's also great at financial research, as you can see from these benchmarks. Now, if you scroll all the way down to the bottom here, it says step 5 is available through their products and API, and the model will be released on October 15th. So, stay tuned for that. If you're interested in reading further, I'll link to this main page in the description below.
If you want to supercharge your content creation, definitely check out Luma, the sponsor of this video. Instead of manually doing everything one prompt at a time, Luma agents can autonomously work with you across entire creative workflows. You can start with an idea and have Luma help you develop the concept, generate assets, experiment with different directions, and iterate on everything in one place. For example, I can give Luma the initial idea and then work with the agent to actually develop it step by step.
It remembers the context of the project so I can keep refining the results instead of starting over with a new prompt every time. You can access the best image generators and the best video generators out there. In fact, one of the most impressive ones is Ray 3.2, Luma's latest video model. What really stands out about Ray 3.2 is its understanding of motion, physics, and 3D space. As you can see, it can handle complex movement, camera motion, and interactions between objects while keeping the scene coherent.
So instead of just generating everything that looks good in a single frame, you can create genuinely cinematic sequences where everything actually moves through the scene in a believable way. If you're doing any type of content creation, Luma is basically like an AI co-pilot that can work with you across your creative workflows. Try Luma today using the link in the description below or by scanning the QR code here. Also this week we have another AI called GAE which stands for geometry native autoenccoder.
This is basically able to generate a 3D world which is much more consistent because as you can see here, not only does it generate just the video of the scene, but it also generates a depth map plus the camera trajectory across 3D space, plus it also reconstructs a 3D model of the scene as the camera moves around so that it can keep the scene consistent in its memory. So, here's another example where on the left you can see it generating the video, but as it's generating the video, it's also generating the camera trajectory and it's reconstructing the entire 3D scene.
Now, this works with a ton of different environments and characters. As you can see from this example, it can be first person or third person. The awesome thing is they've released this already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. Note that the weights are fairly tiny. Everything is under 12 GB in size, so you should be able to fit this on most midconsumer GPUs.
They've also released the training and evaluation script for this, which is fantastic. If you're interested in reading further, I'll link to this main page in the description below. Now, last week, Meta releases something called Muse, which is like your personal AI agent. You can talk to it like a chatbot, and it can autonomously do tasks for you, like summarizing your emails or doing shopping for you or booking stuff to your calendar, drafting replies. etc.
It's kind of like OpenClaw, but Meta just wrapped it into a much easier to use product. Now, I don't know about you, but I'm wary about giving Meta my personal information. So, instead of Muse, actually, this week we have an open- source alternative called Open Muse. It's basically the same thing. This is like a personal agent with a browser and terminal and files, and you can just give it any task, and it can keep working for a long time until it achieves your goal.
You can connect it to Gmail or Google Calendar or other apps. So this can also help you like read and summarize your emails or draft replies or book meetings for you on your calendar etc. So on this page it contains all the instructions on how to get started. Now this is an agentic framework so you can actually use any model you want. Of course you can pair this with a local model if you have enough compute or if not you can also link to a paid model via an API key.
And the awesome thing is this is under the MIT license which is very permissive. So, if you're looking for an open-source alternative to Meta's Muse, here's a nice tool to check out. I'll link to it in the description below. Now, I think it was last week we had this new type of model called Jev, which uses something called system one thinking, which makes really quick and fast decisions. This is different from system 2 thinking, which is like more deliberate, deep reasoning.
And this allows it to answer questions really fast. Well, this week we have an even better and faster alternative, which is open source. So, it's called contrastive language models or CLM for short. This is also an ultraast system one model, but it's trained with something called contrastive learning objectives that connects states and actions. It's basically the same thing as Jev. So, think of it as like giving the AI the current situation and a menu of possible things it could do next.
CLM looks at them all and then immediately scores which action best matches the situation. It basically gives you a confidence or probability score for each option. Now they built this around the open source Quen 38B. This is a fairly small model at only 8 billion parameters. But as you can see, this is like nine times faster than Jev while matching its performance most of the time. So here are some examples of this new CLM model playing Dino Run versus Jev at the bottom.
You know, the interesting thing is for most video games, it's just looking at the current state you're in and then figuring out the best action to take next. So system one thinking actually works very well. But we don't actually need to get an AI model to do really long and deep reasoning, unless you're playing a really strategic game like chess. Now, on this page, they revealed a lot of technical details on how this works, including the training algorithm, the data recipe, etc., including every stage of training.
So, this is fully open- source. And at the bottom, if you click on this GitHub repo here, it contains all the instructions on how to download and run this locally on your computer. If you're interested in reading further, I'll link to this main page in the description below. Now, earlier this week, XAI releases their latest and best model, Grock 4.7. As you can see from the white line, while it doesn't match the performance or intelligence of other Frontier models like Fable or Opus, Grock 4.7 is able to do things at a much cheaper cost.
Here are some benchmarks for your reference where you can see that Grock 4.7 is way cheaper than the other models. Here are some additional benchmarks for your reference. So, in terms of GDP val, which measures an AI's performance in professional knowledge work, you can see that it even beats GBT6 Astra, and it's just a few points below Fable 5.1. Here's another benchmark, which tests its performance on multi-hour office work.
And then here is electrical engineering. Now, those are just their self-reported benchmarks. So, let's look at an independent evaluator called artificial analysis instead. If you look at their intelligence index, then unfortunately, this is still behind GPT6 as well as the latest clawed models. And this is actually just tied with the best open-source model, Mimo 2.6 Pro. So, in terms of intelligence, it's not that impressive.
And interestingly, if you look at the cost per task, at least according to artificial analysis, this even costs more than GPT6, which is kind of shocking. So, I really don't see how this is more costefficient compared to the other Frontier models. Again, here's a chart showing intelligence versus cost. And as you can see, Grock 4.7 is actually all the way down here. It's even less efficient than the other Frontier options as shown by the dotted line.
But if you're interested in trying this out, it is available already in Cursor and Grock Build and also through the API and other thirdparty providers. So if you're interested, I will link to this main page in the description below. Also, this week, OpenAI releases their latest models, GPT6, Soul, and Luna. Now, the naming convention is getting more and more confusing, so let's clarify this. Currently, the GPT6 family consists of three different categories.
We have Astro, which is the highest intelligence. As you can see from this chart, it performs the best, but it also costs the most. And then the midcategory is Soul. As you can see, the Soul models score slightly less in terms of performance, but it's cheaper. And then the cheapest and most lightweight variant is the Luna category. Here, it scores the lowest, but it's also insanely cheap. So you would use the Luna models for like really quick and easy tasks that don't really require much reasoning.
So if you look at its overall intelligence, then GPT6 Soul Max does perform a few points below GPT6 Astra and the leading claude models. However, this is much cheaper. It costs basically over three times cheaper compared to GPT6 Astra. And let's not even talk about the Cloud models cuz they are way too overpriced. Now, if you look at the smaller Luna model, it performs all the way over here. So, not really within like the top five, even below Deepseek and Gemini 3.8.
But again, the advantage of the Luna models is they are incredibly cheap. They are one of the most costefficient models to use. So, this just costs like 7 cents per task. Whereas, if you compare this with like the most expensive clawed model, it's like over a 100 times more expensive. Now, in terms of availability here, it says both soul and Luna models are available in GPT work and codecs for all paid plans. And then for free users, you can access the smaller GBT Luna in the desktop app already.
These models are not yet available in chat. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Anthropic releases their latest and best model, Claude Opus 5.5. Now, I already did a full review video testing the hell out of it. So, if you're interested, see this video for more info. I'm not going to repeat too much here, but basically after a few days of testing, here's the vibe that I got.
It is among the Frontier models out there. It's incredibly good at like 3D and design and coding things and also navigating through websites and working with different interfaces, but I wouldn't say it's significantly better than GPT6 Astra. Their intelligence is really close. However, if you look at the cost per task, then Claude is still way more expensive than GPT6 Astra. So, at least for me, I still prefer to just use GPT6.
The claw models are just way too overpriced. But anyway, if you're interested in learning more, I'll link to this video in the description below. Now, in addition to Opus 5.5, Anthropic also releases this blog claiming that they discovered a new enzyme system for gene editing. So researchers just asked Claude to search a massive collection of DNA sequences for unusual reverse transcriptises, which are basically enzymes that copy RNA into DNA.
So Anthropic used roughly 950 agents using 210 million tokens. And after spending 21 hours, the agent noticed this string of repeated letters, which could be related to the crisper mechanism, which is used for a gene editing. But there's an important caveat here, which is that they don't actually know what this does. So this requires further investigation. So the impressive part here isn't that Claude discovered a potential new tool for gene editing.
It's that AI agents have independently searched a huge data set of DNA and narrowed it down to something that could be worth testing. Now, to be fair, this announcement from Anthropic is also getting some push back from people in biotech. So some people say that this discovery is way less significant than anthropic streaming. So these researchers claim that DNA contains a ton of these crisper-like repeat sequences and unusual combinations.
And scientists already know that most of these ultimately don't turn into anything particularly useful. So finding another unusual sequence pattern doesn't automatically mean you've discovered a new gene editing system. So that's an important caveat here. The significance of this might be a bit overblown. Anyway, if you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have this cool demo from Primebot, and they released what they claim to be the world's transformable personal robot.
It's called the Primebot T1, and as you can see here, it can transform from a quadriped doglike mode to a bipedal humanoid mode. And this transformer feature has various benefits. For example, if you need it to carry stuff, then it can first transform into a dog, then you can just place things on its back. Or when you need it to do some certain actions that require the humanoid form, then it can also just transform into that really quickly.
So quite an interesting design from Primebot. And this is really cheap. This is only around 20,000 yen, which is like $2,800. They say that shipments are scheduled to start in October. Also this week, your waifu is finally getting shipped. So, we have a new update from UBTE Robotics. They say that the first batch of this U1 series of super realistic humanoid robots are now being shipped. I featured these robots before.
They are designed to have super realistic faces, hair, clothes, and poses. Now, I'm a bit surprised they haven't created a catgirl option. I think this is a huge missed opportunity. But anyways, the exciting thing is this week they finally announced that at least the first batch of these robots are already being delivered. In other robotics news, Unitere has introduced their latest dextrous robot hand, the Dex 5S. And this is quite an incredible feat of engineering.
This is a one one human hand size. It's trying to basically copy a real human hand as closely as possible. It has 22 degrees of freedom. So that means it has a lot of separately moving joints so that the fingers can bend and pose and be super flexible just like a human hand instead of a stiff robot claw. And here they say that the price starts at $6500. Each joint also has impact torque protection. So if the hand slams into something, it's designed to limit the hit.
And as you can see from this demo, this can do a ton of things from like holding and manipulating objects, tapping piano keys, doing some finger gestures, and using scissors to cut paper. In other humanoid robot news, Skilled AI released a pretty impressive demo. They're basically kind of doing Alph Go, but for humanoid soccer. So, their main robotics model is called the S1. And here they're training the robot to autonomously play soccer.
And the way they train it is kind of like how DeepMind trained Alph Go to become world class in Go by just playing millions of rounds against itself in a simulation. Well, here they're doing the same thing for this robot. It's playing soccer against itself in simulation. And soccer is a brutal choice on purpose. This requires a ton of balance, speed, ball control, reactions, and a sense of how the game actually works.
So, they just gave S1 one goal, which is to score. And they let it play the equivalent of more than 140 years of soccer in Nvidia's simulator called Isaac Sim. Now, of course, at first, without any knowledge, it falls over. It's pretty clumsy. But after a few rounds of reinforcement learning, it was able to dribble around defenders, shield the ball, and also shoot. The opponents were basically older versions of itself.
So, as one version got better, the next version had a harder match. This is what they call an automatic training ladder. So, afterwards, after so much training, you can then deploy this back into the real world, and you would get a humanoid robot that can kind of autonomously play soccer against a human. Now, this is just a fun demo, but the same strategy of self-play in a simulation can move beyond soccer into things like navigating around people or moving objects or self-driving or any other specified goal.
If you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have a ton of smaller models that are not necessarily Frontier, but they're still worth mentioning, so I'll rapidfire through them. The first one is called Dots 3 Note, and this is by Red Note, which is like the Chinese version of Tik Tok. Here this is a mixture of experts model with 28 billion parameters.
When you use it only 16 billion are active. Now this has a context length of 512 tokens. So a bit smaller than the frontier models which in general have a million token context window. Nevertheless for its size it actually is pretty good. So here are some agentic coding and problem solving benchmarks. And as you can see it's even able to perform very close to models that are many times its size. For example, Kimik 3 is like 2.8 8 trillion parameters which is like 10 times the size.
Here are some other benchmarks. So in terms of multimodal understanding again it performs very well given its size. In fact one of the most impressive things about it is its score on this arc AGI benchmark. If you're not familiar with this benchmark it basically places the AI model in this game environment which it has never seen before and it needs to learn the rules of the game by itself and basically win the game.
Now, for most AI models, this is really hard to do because they can't really technically learn new things over time. Their model weights are fixed. So, it's really hard for them to learn new rules and apply this new information to their decision-making. But, as you can see, at least for this ARC AGI2 score, this new DOT 3 note actually performs the best out of the open models, and it's also really cheap. So, here they've open sourced this.
They've released the full model and also an FP8 variant. The full model is 577 GB in size, whereas the FP8 is a lot smaller at only 300 GB. A decent medium-sized multimodal model to try out. If you're interested in reading further, I'll link to this main page in the description below. Another noteworthy open model is this OUS VL embedding. So, this is an embedding model, which is quite different from a regular large language model.
This is basically designed to search across a ton of different information using the same system. So instead of having one model for text search, another for images, another that searches video, and another that searches documents here, this embedding model basically takes all these different types of formats and converts it into the same numerical representation, also known as the embedding space. Think of it as like translating all these different types of inputs into just one universal language.
Now, they've released two different variants of this. There's a omni3 billion parameter variant which also takes in audio. And then there's also just a vision variant which is either two billion or 9 billion parameters and this can take in images and videos but not audio. And if you compare its performance with other similars sized embedding models then this is on average state-of-the-art. So both models are fairly tiny.
The Vision 9B1 is 16.8 GB in size. The Omni 3B1 is only 11 GB. And the awesome thing is the models are under the Apache 2 license which has very minimal restrictions. So, this is not your standard AI model, but if you need an embedding model to convert different formats of input into one shared space, then this might be one of the best open options to use. If you're interested in reading further, I'll link to this main page in the description below.
And it gets even tinier. So, we have another model called Limite 1B Violto. Aside from the really strange design of this landing page, which kind of hurts my eyes, this Liite 1B model is a super tiny model designed entirely for one thing, which is solving really hard math problems. So, you can see from this competitive math benchmark, it scores like 94%. But because it's so small, it uses way less compute compared to other larger models.
Here are some other competitive math benchmarks for your reference. And as you can see, on average, it just clobbers all the other models, which are like dozens to hundreds of times larger. So, this is super impressive. Now, like I said, this is specialized to do only one thing, which is solving really hard math problems. So, it's not going to do so well in other general tasks. But if you're interested, at the bottom here, they've released the code and the models to this.
So, if you scroll down a bit here, it contains all the instructions on how to download and run this locally. And at 1 billion parameters, this is super small. So, the total size of everything is only like 2 GB in size. You can even fit this on like low-end GPUs. If you're interested in reading further, I'll link to this main page in the description below. Another open model to put on your list is Iikido Alter. This is an openweight model that specialized for cyber security.
So, they fine-tuned it on one of the best open models out there, GLM 5.3. Now, the full model is like 1.5 TB in size. It's pretty huge, but with some clever engineering, they basically compressed it all the way down to just 328 GB while preserving most of the reasoning quality of the full model. So, you can see the reduced size of this new Alter model compared to the full model. And out of 32 vulnerabilities from their test set, Alter was able to discover 23.
So, its performance is not that far from the full model which was able to discover 25. And this is open weight. So down here, if you click on download alter, it'll take you to their hugging face page, which contains the models and instructions on how to use it. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Google releases some new details about their plan to send data centers to space.
So, this is called Project Suncatcher, and they're basically exploring whether we could eventually build massive AI computing systems in space. The rationale behind this is that data centers need a ton of electricity while satellites in low Earth orbit can get almost constant sunlight and potentially generate up to eight times more solar power than systems on Earth. But before building anything close to a space data center, they first need to answer a much simpler question, which is whether AI chips can actually survive in space.
So Google is preparing a prototype satellite carrying its trillium TPUs or tensor processing unit to test exactly that. Now one problem is getting these chips into orbit. During the roughly 10-minute ride into space, the spacecraft can generate forces of up to 10G and you know these hardware components can experience up to 100g. So Google basically strapped these AI chips onto a giant vibration machine and shook it like crazy in all directions to simulate a rocket launch.
And surprisingly, the hardware survived. And then there's another issue, which is radiation. Once you're outside the Earth's protective atmosphere, solar particles and cosmic rays can mess up electronics and even flip individual bits of data. So, Google tested its TPUs inside a proton beam facility while the chips were actually running AI workloads. And early results showed that these trillium TPUs could actually withstand more radiation than expected.
Now, another challenge is cooling. You see, on Earth, AI data centers can use air or water to carry heat away from the chips. But in space, there's no air flow. So, Google is experimenting with heat pipes and radiators that move heat away from the processors and then radiate it into space. Another difficult part about this is that these AI chips need to exchange tons of information incredibly quickly. It's not just a ton of AI chips working separately, right?
So Google also needs these satellites to communicate using very high bandwidth laser links. And keeping these lasers connected requires extreme precision. Google compares this to like hitting a coinsized target from miles away while both sides are moving. So they also need to figure this part out. Anyway, Google says they plan to test the system with two satellites next year. So currently Google's first mission in this project Suncatcher isn't about sending a full data center into space yet.
It's about proving that the basic pieces can survive. Can the chips survive through launch? Can they handle years of radiation? Can they stay cool in a vacuum? And can they communicate quickly with each other in outer space like a huge AI computer? If those things can be solved, then space could eventually become a new place to build AI infrastructure powered by the sun. Anyway, if you're interested in reading further, I'll link to this main page in the description below.
Also this week we have a super tiny image model that can even run on your phone. Now this is called Supra image and this is only 100 million parameters. Keep in mind the other image models out there are like at least billions of parameters. For example, Z image is like 6 billion parameters. So this is like 60 times smaller. Now given its size, the images that it can generate are pretty low quality as you can see from these examples.
But that's expected because it's really hard to pack so much intelligence into just a 100 million parameters. But the nice trade-off here is that this is small enough to run on just your phone or other smaller devices. At the bottom of this page, it contains all the instructions on how to run this. Note that the entire model is only like 417 megabytes in size. If you're interested in a super tiny image model that can even run well on your phone or other smaller devices, this might be a good option.
If you're interested in reading further, I'll link to this main page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite? And which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content.
Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up to date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.