Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
5,719
Runtime
32:11
Speaking pace
178wpm
Reading time
24min
178 words per minute, just under the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
AI never sleeps and this week has been absolutely insane. We have a ton of new robot waifuss. This AI can turn anyone into a 4D animated character. We have a new open-source image generator and editor. Deepseek drops their latest model with vision. We also have a new super tiny texttospech generator that can even fit on low-end hardware. Bite Dance drops their latest open- source video editor. This new open- source model can create interactive worlds in real time. You can even prompt it to add events
89 words, the words spoken in the first 30 seconds at 178 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 367 |
| Average words per sentence | 15.6 |
| Longest sentence | 75 words |
| Questions asked | 8 |
| Sentences containing a number | 70 |
Most used terms
Filler phrases
63 in total: like 25 · basically 23 · actually 7 · you know 5 · kind of 3.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
AI never sleeps and this week has been absolutely insane. We have a ton of new robot waifuss. This AI can turn anyone into a 4D animated character. We have a new open-source image generator and editor. Deepseek drops their latest model with vision. We also have a new super tiny texttospech generator that can even fit on low-end hardware. Bite Dance drops their latest open- source video editor. This new open- source model can create interactive worlds in real time.
You can even prompt it to add events or effects. This AI can take any video and reconstruct an entire 3D model of the scene. This robot now beats the human world record for the highest jump and the fastest sprint. This thing is insanely fast. We also have a real transformer robot and another one that can play tennis and a lot more. So, let's jump right in. Thanks to HubSpot for sponsoring this video. First up, we have a new AI video model called Evoke.
This is fully open source, and this can generate an interactive world in pretty much real time. The input is an image plus joystick movements, and it outputs video that basically responds to these joystick movements almost instantly. You can control a variety of different vehicles like this snowmobile. You can also create various events like this volcano eruption. You can also use text prompts to add events to the video.
For example, we can tell it to add balloons to the scene like this. Or here's an example where we can prompt it to add aurora lights. And as you can see here, it understands a variety of different scenarios, including like driving different vehicles or kayaking or scuba diving, rock climbing. It has a great world understanding. So, it's essentially like Genie 3 where you can prompt it to generate any scenario you imagine.
It also works with different artistic styles as you can see here. Now, because it has such good world understanding, what I think is a really useful application for this is creating videos that could help train robots. For example, we could create some simulated videos of how to move around and operate in hospitals or care centers or whatever and then use this data to train robots. Or we can also use this to create synthetic training data for like emergency rescue situations or underwater exploration or a ton of other different scenarios.
Now here they say this is only 14 billion parameters and the reason why it's so fast and essentially real time is because it can generate videos in just three steps which is really fast and sessions are designed to continue for hours. The awesome thing is this is out already. So at the top of the page if you click on this code button and you scroll down a bit here it contains all the instructions on how to download and run this locally on your computer.
And the nice thing is this is fully open source and they also released the models for every stage of training. So you can see the final model is around 57 GB in size. So you do need a decent high-end GPU to run this. Now excluding other thirdparty code, this is under the Apache 2 license which has very minimal restrictions. If you're interested in reading further, I'll link to this main page in the description below.
Also, this week we have a new AI called 4D anyone. And here's how it works. You just need to input a single video of any person moving and it can turn that person into a 4D reconstruction that you can view from any angle. It's basically like a moving 3D model. Specifically, this creates a 4D gausian splat that represents the character and their movements. And as you can see, this works across a variety of different characters and outfits.
Now, how this works is quite interesting. So it basically takes the video and then it extracts a 3D skeleton from the source video. Then it uses that skeleton to guide the generation of many new camera views of this character. And then from all these views, it basically reconstructs everything into a 40 gausian splat model. And if you compare this new 40 anyone with other 4D generators like Recamp Master or Trajectory Crafter, which I featured on my channel before, you can see that this new one is a lot more detailed and consistent. definitely the current state-of-the-art in terms of generating 40 characters.
Now, at the top of the page, they've released everything already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. The model is only 12 GB in size, so you should be able to fit this on most mid to high-end GPUs. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a new open- source image generator and editor called Sense Nova U 1.58B.
And this is a huge deal. Here are some example generations for your reference. As you can see, it can generate some super realistic photos. It's also great at designing posters and infographics with a ton of different elements. It has no problem handling all these elements. Here's another example. And like Nano Banana and GBT Image, this one can also edit images using natural language. For example, we can change the text of this poster or here we can selectively change certain elements in this poster.
Here's a cool example where we can just directly label our instructions on the photo and it can follow all these instructions. This can also take in multiple inputs as references. So here's an example. The really special thing about this is this can generate native 4K resolution and this generates images end to end in pixel space. In simple terms, for most image and video generators, they actually generate images in a more simplified dimension called latent space.
This makes it more efficient to process. But then afterwards, you also need to decode the image from latent space back into pixel space, which you and I can see. In fact, you might be familiar with the term VAE, which is a model which you load, which basically decodes the image back into pixel space. Well, since Nova completely skips this step, there's no need for any encoders or decoders. Now, traditionally, we didn't do this because it was not efficient.
It took a lot more time and compute, but they basically optimized it, so we don't even need this step, and it can still generate 4K resolution, which is pretty cool. At the bottom here, it contains all the instructions on how to download and run this locally on your computer. Now, the model is quite huge at around 50 GB in size. So, you'll need a high-end GPU to run this, but hopefully there will be more quantized or compressed versions in the future that can fit on lower VRAM.
Another awesome thing about this is it's under the Apache 2 license, which has very minimal restrictions. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Bite Dance releases their new open-source omnimodal video editor. It's called Bernini version two. And I mentioned version one a few weeks ago. You can basically take any existing video and edit it with natural language.
So for example, we can add two characters to the scene and change this to a warmer tone or we can also remove a specific character or object in a video. We can also change the camera perspective of an existing video or change the background. And we can also add in reference images of things we want to insert into the video. So for example, we can insert this helicopter into the scene. Well, that was version one and this week they released version two.
Again, this is a huge model at 180 GB in size just for the model. Not to mention, you also need the VAE text encoder, etc. So, good luck running this locally on your device. That's why even the first version one hasn't really gotten much attention from the open source community. I think this might be a bit dead on arrival since we already have Miniax H3 which can also do reference to video, but in case you're interested in trying out version 2, I will link to this page in the description below.
Also, this week we have a new family of open models called Ornith 1.5. And this is a big deal. It's built around a pretty ambitious idea which is instead of having humans constantly creating new training tasks and data for the AI here they basically created a self-improvement loop where the system proposes new problems and builds the tools or scaffolds needed to solve and verify those problems. It also generates the solutions and then uses those results as reinforcement learning data to train the AI models.
It's basically trying to create a closed loop where a stronger AI can invent harder and harder challenges for itself. So they took this framework to train this new family of open-source models called Ornith 1.5. And this consists of 9 billion parameter dense model, a 35 billion parameter mixture of experts model, and the largest one is a 397 billion parameter mixture of experts model. And you can see for the largest one it performs incredibly well across all these agent coding and knowledge work benchmarks such as terminal bench bench deep frontier bench etc.
What's really impressive is that you know orange 1.5 is only less than 400 billion parameters but it even beats GLM 5.2 which is twice as large across all these different benchmarks. It also edges very close to opus 4.8 8, which is closed source, so we don't know the exact size, but I would assume it's over a trillion parameters. So very impressive results from this new model. Now, if you look at the medium-sized 35 billion parameter model, again, it's pretty state-of-the-art.
It even beats Quen 3.6 across most of these benchmarks. But note that they're kind of cherry-picking here because we already have Quen 3.8. The nice thing is they've released this already. So if you click on this hugging face page, it contains all these models. So the largest one is 794 GB in size. you'll need to stack like multiple accelerators to run this. But the smallest 9B version is only 18.8 gigabytes, so this should be able to fit on most like mid to high-end GPUs.
Plus, they also released GGFs for this. So, the smallest 4-bit version is less than 6 GB in size. So, this can even fit on low-end GPUs, making it very accessible. So, if you're interested in running LLMs locally in addition to Quen 3.8, here's another decent option to add to your list. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have a new super tiny texttospech generator called audio8TS.
As the name implies, this is very small at only 0.1 billion parameters. And like other texttospech generators, this basically allows you to take a few seconds of someone's voice and it can clone them to say anything you want. Here are some demos. So, first, here's the original voice. >> The room is quiet now. Take a slow breath, close your eyes, and let the day fade gently away. >> And let's get this voice to speak out this transcript. >> Before you leave this morning, remember to close the windows, check the stove, and open them again when you get home. >> As you can hear, it sounds very similar to the original voice.
Or here's another example. Here's the reference voice. >> You are really the the sunlight of my day. you. When I'm with you, I feel so much warmer and and better and and safer, and it's just the best feeling. >> And let's get that voice to read out this transcript. >> I'm getting off work a little early today, so I'll pick up some groceries and we can cook dinner together. And this is multilingual, so here are some examples with different languages.
Internationalour. Jack Buenos. All right. So, if this is of interest to you, they've released the code and the models already. Note that the total size of everything is only 1.7 GB. So, this is very tiny. This can fit on most consumer devices. And if you scroll down the HuggingFace page here, it contains all the instructions on how to download and run this locally on your computer. If you're interested in reading further, I'll link to this main page in the description below.
You've probably heard people talking about OpenAI Codeex, but if you're not quite sure what it actually does or how it could save you time at work, then this free resource from HubSpot called Codeex Prompts that replace busy work will save you hours every week. Codeex is different from the regular ChatgPT. Instead of just answering questions in a browser, it can work directly with files on your computer, complete tasks, and save the finished outputs back to your machine.
And you don't need to be a developer to use it. This guide includes five copy and paste prompts designed around some of the most repetitive and time-conuming business tasks. The first prompt creates a daily work brief by reviewing things like your calendar, emails, messages, and open follow-ups. It then organizes everything into your priorities, meeting prep, messages that need replies, and decisions that need your attention.
There's also a weekly summary prompt that turns a full week of scattered meetings, documents, messages, and product updates into a clean manager ready report. My favorite is the skill creation prompt. This lets you take a workflow that worked well once and turn it into a reusable skill, so you can run the same process again without having to explain everything from scratch. You can grab all five prompts for free using the link in the description below.
This resource was made by HubSpot, the sponsor of this video. Also, this week we have a pretty interesting AI called Geo Weaver. This can turn an ordinary video into a coherent 3D reconstruction of the whole scene. The challenge is that doing this over a long sequence is much harder than reconstructing it from just a few frames. That's because the AI can slowly lose track of the scale and camera position, causing the reconstructed 3D world to drift apart.
But Geo Weaver tries to fix this. It first breaks a long video into manageable chunks and predicts things like depth and camera position for each one. Then during inference, it gradually stitches those chunks together and adjusts them until the entire sequence agrees on one consistent 3D world. It uses nearby frames, overlapping views, and even long range matches between distant parts of the video to keep everything aligned.
And the result is more accurate camera trajectories, better global consistency, and cleaner point clouds. You can see across all these benchmark scores, it has the lowest average error rate compared to other competitor models. Now, if you scroll to the top of the page, currently they've only released a technical paper, but for now, if you're interested in reading further, I'll link to this main page in the description below.
Also, this week, we have a very interesting AI called Quen VideoEedit. And like the name implies, this can edit video based on a text prompt. Now, this isn't a completely new video model from Quinn. Instead, what it does is it uses an existing image editor called Quinn imageedit and it plugs it through a video generation workflow. Specifically, they used Alibaba's wand for this. And it turns out that you can use this to edit existing videos.
It's basically using Quinn imageed to edit the video frame by frame. And so from this, you can take an existing video and use natural language to edit any part of the video. The nice thing is they've released the code to this already. So if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this. They also released the training script for this as well.
Now the model is quite huge, so the total size of everything is 41 GB. You'll need a high-end GPU to run this. And honestly, Miniax H3 can already edit existing videos, so I don't really see a point in using this. But if you are interested in trying this out, I'll link to this main page in the description below. Also this week in robotics news, Passini Tech released their next generation data collection glove called the PX Cap Pro and this is built for one clear job, which is collecting highquality hand data that robots can actually learn from.
So you just put this lightweight glove on and it captures everything that you do while your hand is moving. It features some tactile sensors that cover the fingertips and the palm, and these are really sensitive, so it can feel forced down to just 0.1 Newtons. It also has a very wide-angle camera on the wrist, which records the full scene of your hand moving and interacting with objects. It also has precision angular encoders, which track the joint angles extremely accurately, even with magnetic interference.
So, if you need to get a robot to automate a certain dextrous task, you can just wear this glove or get employees to wear these gloves while they do the task to collect data to eventually train the robots. This can handle super delicate tasks like tying ribbons, working with balloons, packing boxes, or even lab work. Also, this week, we have a real life transformer. Not this LLM transformer, but an actual robot transformer.
So, this Chinese company called Arc Shell Robotics just showed a robot that can basically switch between four completely different forms. So, in one form, it's a bipedal humanoid robot which can stand upright. But in another form, it can also transform into this quadriped and basically walk on four legs. And then in its third form, it can also attach to this flying drone to basically be transported elsewhere by air. And for the fourth form, which is coming soon, apparently it also becomes wheeled.
Now, while this does look pretty cute and interesting, I'm not sure how practical or effective this is. If you design something with all these different forms, then you're also going to get more failure points. So, is this a really good trade-off? We're not really sure yet. They haven't released this, but if you're interested, this is the MXD1 robot from Arc Shell Robotics. Now, this week, we have the World Robot Conference in China.
And you know what the best part is? We have a ton of waifu demos. So, first up, we have this robot called the Annie Wit Annie. And here she's programmed to sing a song. You can see her lips are synced to the song, plus she can kind of move her body around. But this one doesn't look too realistic. So, we also have this demo from Ubitech. And this robot looks extremely realistic. As you can see, she can blink and look around.
Her head moves supernaturally. It's getting more and more real. Now, that's just a head and torso variant. They also have a full body variant. And again, this looks extremely lifelike. So, those are some demos from Ubtech. Now, on my channel, I mention cat girls a lot, but if that's not your thing, maybe you could also try elves. So, in this world robot conference, another company called a headform also featured their new elf bionic robot, which is called Elf Schwan 2.0.
Now, this robot features pointed elf ears with some ornate floral accessories and an elegant floral dress. In terms of realism and talking and having subtle expressions, I would say a headform's robots are currently the best. And you know, the best part about this is previously a headforms demos were only head and shoulders, but now with this elf Schwen 2.0, they actually gave her a body. So, it's a new articulated body with expanded degrees of freedom.
Right now, she can only do some basic actions, but it's only a matter of time before we can get it to walk or dance or do some other things. Also, at the robot conference, we have this giant robot horse from DAX AI. It's called the Chi, and this is a quadriped cyber horse built for rough terrain. You can climb on and ride it. The X1 version costs about only 40,000 USD. They also have a wheel liged XS version, which costs around $53,000.
And this is designed to transport a human across really rough terrain like steep slopes or gravel, mud, snow, ice, etc. Now, this thing has a 300 kg payload plus a 40 km range and a top speed of around 10 km per hour. They also have a wheeled version which is not demonstrated here, but this can cruise even faster at up to 40 km per hour. So instead of humanoid robots, now we also have potentially a robot horse which you can sit on and ride around.
Now that was the World Robotics Conference, but we also have the World Humanoid Robot Games taking place in Beijing very soon. So here we can see an open ceremony rehearsal. And we can see a ton of these Tienong robots marching across this track. Here we can see the booster robot and also the Galbot robot and a ton of other robots also marching into the scene. It's pretty crazy. It's like the Olympics, but for robots.
This event is starting soon, so I'll update you on the results next week. Also, this week, Uni Tree just previewed a new humanoid robot, which they call Superman. And this is crazy. This can do a standing high jump of about 2 m, which is already higher than the human standing high jump world record, which is only 1.8 m. It also has a top speed of almost 12.7 m/s, which edges past the fastest human sprint speed, which is only 12.4 m/s.
So, this robot can already outrun the fastest human sprinter in the world. The whole machine has only been in development for a little over 3 months, so there's still plenty of room to improve. And speaking of speed, here's another demo of this Superman robot sprinting across a track. And here's where you can see how ridiculously fast this is. Plus, you know, the funny thing about this is they designed it to run so fast that it can't really stop.
Here you can see it's really struggling to slow down and eventually crashes into the wall. Or here's another example where again, you can't really get this thing to slow down. So, it just crashes into this barrier in order to well, stop running. Pretty cute if you ask me. By the way, this isn't the only robot that's now able to outrun humans. So, we also have another demo from Honor. Here it's showing their lightning robot which just pulls ahead and completely outruns this human.
And not only that, but we also have yet another robot. This time it's called Tien Gong. And again, it's just incredibly fast. By the way, these are just warm-up videos. Like at the time of this recording, the humanoid robot games have not even started yet. But you can see it's just a massive improvement in the speed of these robots compared to just last year. Now, in addition to jumping and running, we also have a few demos of humanoid robots now being able to play tennis, which is pretty crazy.
So, the first demo is from this team, which built a system called Adapt. It basically takes in data from real tennis matches and transfer the players moves onto real humanoid robots. In this case, the Uni Tree G1. And not only can this robot rally, but it can also serve. And it's doing all of this autonomously. If you've played tennis before, you'll know that it's actually extremely hard. It ain't as simple as just hitting the ball with a racket.
You need to hit it with just the right amount of force. Plus, you also need to slice the ball, whether it's top spin or backspin. There's a lot of physics that needs to be decided and executed in real time. So, it's really impressive how they were able to program a robot to autonomously do all of this. Now, that's just one demo, but another I would say even more impressive demo is from another robot called Galbot. And here you can see a demo of it playing tennis in the world humanoid robot games which is happening right now.
Again, this is fully autonomous. You can see the robot has to run toward the ball and also hit it with the right force and spin. And it has to do all of this in just a split second. Pretty impressive. Also, this week, Deepseek releases yet another new model called Deepseek V4 Flash Vision Experimental. The awesome thing about this is it matches the performance of the regular textbased V4 flash, but now it includes vision capabilities, meaning it can analyze images, videos, and documents.
So, here are some benchmark scores comparing this with the previous V4 version without vision, and also Opus 4.8, which is closed source and probably many times larger. But, as you can see, across all these agentic coding benchmarks, this new V4 flash vision even matches the performance of Opus 4.8. And also check out the deep suite score. It improved by almost five points in less than a month. Pretty crazy. Now, currently this is only available via API.
So, if you're interested in trying this out, I'll link to this page which contains all the instructions on how to use it. Also, this week we have a new music generator which I believe is by the same lab that created Happy Horse. So, here it's called Happy Shrimp. And like most music generators, you simply describe the style of the song. Plus, you can also enter lyrics, or you can also just toggle this to instrumental.
Here are some trending examples for your reference. >> Windows like a crown of gold. You build your walls just to watch me bend. But I roots deep inside the clay. You think I will break in the shadow you make but the fire is awake. I am standing tall though. No more chains on my soul. No more games to be played. I will light the single spark. You can throw your stone, but I hold my throne. We sat on the porch while the autumn wind blew the leaves away.
I traced the lines on your palm just holding you close. Noticing the dust on your jacket from a long drive down the interstate and your shoes covered in mud from the pouring rain. But the silence in the room hits me harder than the stretch between us on the high. And I never wanted much, but I wish this sleepy town would just vanish in a heartbeat. Oh, I can't keep on waiting for you. Waking in a hustle, lonely, wondering if my heart should leave, searching blindly in a dark.
Darling, tell me, are you fine staying here? Are you fine to stay >> and wait? Will you pack and go? Will you be here when I wake up? I always hated the quiet at 4 a.m. when it's too late for sleep. >> This sounds super clean and dynamic. This is definitely one of the best music generators you can use right now. And at least at the time of this recording, you can use it for free. If you're interested, I'll link to Happy Shrimp in the description below.
Also this week, Comfy UI has open-sourced their agentic connector called Comfy MCP. This is basically like an API where you can connect an AI agent directly to your Comfy UI installation and have it understand all your workflows, all your models, etc. So instead of like manually dragging and dropping all these nodes and noodles onto your interface, you can just take an AI agent like GPT on codecs or GLM on Zcode and just prompt it in natural language to, you know, generate a video for you with Miniax H3 and it can just automatically spin up a Miniax workflow and generate the video for you without you having to actually touch the Comfy UI interface.
The awesome thing is this has direct awareness of your GPU, your hardware, all your installed models and custom nodes. So over here if you click on this link it takes you to their GitHub repo and here it contains all the instructions on how to install Comfy MCP on your computer. If you're interested in reading further I'll link to this main page in the description below. Also this week we have a really interesting robot foundation model which can kind of generalize or learn new things.
So the model is called Gen 1.5 and the real breakthrough here is that you can show it how to do something just once and it can sometimes attempt the same task without any additional training. So this is a huge deal. Instead of programming the robot step by step on how to do it or instead of training it on data for multiple rounds, this robot foundation model can just learn a task from potentially just one demonstration.
So here on the left you can see a real demo of a human doing the action and then on the right is the robot attempting to repeat the action that it just saw from one demo. So for those of you who are new to this, a robot foundation model is basically the brain that controls a robot. It takes in video from its eyes, sensor information, and it can also understand natural language, for example, from a human's instructions, and then it can output movement trajectories at 100 times per second.
And in this oneshot setup, the model can basically learn from just 3 to 12 seconds of a demonstration. Now, the success rate isn't huge. So, across 10 diverse tasks, it shows a 59% success rate from a single demonstration. And with a small amount of additional training, using about 5 minutes of data per task, you can get that success rate to rise to 83%. still not close to perfect and these tasks are relatively simple and short, but it's still a big deal that they could train this model to now learn from demonstrations.
It's a small but important step toward generalist robots where a human just has to show them once on how to do a certain action and then they at least have a chance of figuring it out. A pretty interesting project. If you're interested in reading further, I'll link to this main page in the description below. Also this week, Nvidia releases a pretty interesting framework called AO. This is basically an agentic framework or harness that allows clot opus 5 to achieve a score of 100% on this Arc AGI3 benchmark.
If you're not familiar with Arc AGI 3, this is basically a benchmark where AI models are placed into these new video game environments with zero instructions. They have to figure out the rules and goals by themselves through trial and error and basically get to the next level or win the game. Now, humans could solve these pretty easily, but it turns out that even the top AI models perform pretty badly on this benchmark.
You can see most of them perform under 10% and Claude Opus 5 only gets 30%. And that's because AI models technically cannot learn new things or patterns after training. So this doesn't just test an AI model's ability to play video games, but instead it tests their emergent ability to actually learn and apply new patterns on the fly. Well, what Nvidia found was that with just a simple agentic harness called AO, which is basically this pipeline, that alone can increase the score of Opus 5 from 30% all the way to 100%.
It basically aced the test, scoring 100 across all 25 environments, completing all 183 levels. Now, they did evaluate this on the public data set. So, the score might be inflated. But nevertheless, this is additional proof that you don't have to optimize the models themselves. You can also optimize the harness or the system around these models to unlock even more performance or intelligence. If you're interested in reading further, I'll link to this main page in the description below.
And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite? and which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week.
I can't possibly cover everything on my YouTube channel. So, to really stay uptodate with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.