Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
7,743
Runtime
40:29
Speaking pace
191wpm
Reading time
32min
191 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
AI never sleeps and this week has been absolutely insane. Anthropic releases their latest clawed model which you can use for free. OpenAI held their annual dev day event this week with some exciting new announcements. Google finally unleashes their Frontier model Gemini 4 Argon and it's an absolute beast. We have some super tiny speechto text generators that are only megabytes in size. So, this can easily fit on your phone or other edge devices. We have not one but two new state-of-the-art image generators and editors. This AI lets you take an image and
96 words, the words spoken in the first 30 seconds at 191 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 492 |
| Average words per sentence | 15.7 |
| Longest sentence | 50 words |
| Questions asked | 9 |
| Sentences containing a number | 97 |
Most used terms
Filler phrases
101 in total: like 53 · basically 21 · actually 16 · you know 4 · kind of 2 · right? 2 · I mean 1 · literally 1 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
AI never sleeps and this week has been absolutely insane. Anthropic releases their latest clawed model which you can use for free. OpenAI held their annual dev day event this week with some exciting new announcements. Google finally unleashes their Frontier model Gemini 4 Argon and it's an absolute beast. We have some super tiny speechto text generators that are only megabytes in size. So, this can easily fit on your phone or other edge devices.
We have not one but two new state-of-the-art image generators and editors. This AI lets you take an image and control the camera path in this scene. Humanoid robots can now autonomously play badminton. We have a ton of new open-source models and a lot more. So, let's jump right in. First up, we have a new AI called in spatial world 1.5. This can basically take any image plus a camera path and it'll generate a video walking in the scene through that camera path.
Now, I featured version one a few weeks ago, but version five is even better quality. And this can start with just a single image or multiple images to give it even more consistency and accuracy. Or you can even input a panorama or a full video. And if you input a video, then it also is able to capture the motion of the original video as you walk across the scene. So, this gives you ultimate control of how you want to view the scene based on a certain camera trajectory.
And of course, you can also like freeze the frame at any time to do some really cool bullet time effects. Now, at the top, they've released the code to this already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this. Note that this is based off of one 2.1. I'm surprised they haven't used Miniax H3 for this, but anyway, this is based off of 1 1.3b.
And this is only less than 6 GB in size, so you should be able to fit this on most consumer GPUs. If you're interested, I will link to this main page in the description below. Also, this week, we have a new AI called Soul Refiner by Nvidia, and this takes lowresolution video and upscales it into a sharper highresolution video up to 4K in a single step. Note that here, if you zoom in on specific sections, it's a lot more detailed than the original video.
The nice thing is this is model agnostic so you can use any video model including the best open one Miniax H3. So here are some examples. Note that with Soul Refiner it is a lot more sharper but it does tend to change some of the details quite a bit. Or here's another example for your reference. Note that the details of this barn are slightly different after going through the refiner. So this is not really faithful. Here's another example for your reference.
Again not entirely faithful to the original video. So, that's the main problem I have with this. Instead of Miniax H3, you can also use another video model like Nvidia's Cosmos or Alibaba 1. The nice thing is they've released this already. So, at the top of the page, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to run this locally on your computer. Note that this is borrowed from the LTX Refiner, but they fine-tuned it to just do the upscaling in one step.
If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a super tiny speech to text model. In other words, an audio transcription model. I think it's the smallest one I've seen so far. So, this is only like 16.9 megabytes in size. Keep in mind, the other ones are at least several gigabytes in size. So, this is really small. Now, because of its size, this can run on just a CPU with no dependencies.
You don't even need a GPU to run this. In fact, let's test this out right now. So, I'm going to speak and let's see it transcribe this. So this can transcribe up to 30 seconds of audio and it can automatically detect the language and also give you the timestamps for each word. All right. So afterwards here is the transcription of what I just said. So like I said you can insert up to 30 seconds in one pass and it supports all these languages which are automatically detected.
It's not shown in this demo but you can also get it to give you the timestamps for the start and end of each word. And look at the insane speed of this. Because it's so small, it's also way faster at transcription compared to Whisper Bass. It also decodes like six times faster. Now, here's the error rate between Whistle and Whisper Bass and another tiny transcription model. You can see, at least for some of these benchmarks, it does have a lower error rate compared to Whisper Base.
Now, at the bottom of this page, it contains the instructions on how to install and run this locally on your computer. It's just a really simple pip install command. Plus, the models are on hugging face. This is like the smallest transcription model I've seen so far. So, you can even use this on like really small devices like phones or other edge devices that don't have GPUs. If you're interested, I will link to this page in the description below.
Now, that wasn't the only tiny texttospech model we have this week. So, here's another one called Phon 2. This is a bit larger, so it's under 900 megabytes, which is still really small compared to other TTS models that are at least gigabytes in size. So, here it says this can turn an hour of audio into text in about only 20 seconds. The nice thing is this is open source. Plus, there are also some free hugging face spaces for you to try this out online.
So, here's an example of me using this right now. And as you can see, here's the transcript. Now, if you look at these benchmarks, it still performs very well compared to other models which are way larger. On average, it gets the second lowest error rate. So, at the bottom, they released the model for you to download. If you're interested in reading further, I'll link to this main page in the description below. Now, earlier this week, Anthropic dropped their latest model, Claude Sonnet 5.5.
Now, Sonnet is their smaller model, whereas Opus is the larger, more intelligent model, but this new Sonnet 5.5 punches well above its weight. If you compare this with the previous version of Sonnet, it has massive improvements over a ton of different agentic coding and knowledge work tasks, and it even scores pretty close to Opus 5.5. Here's its terminal bench score for your reference. Now, even though Sonnet 5.5 is a smaller model, it's not necessarily cheaper to run.
We'll talk more about that in a second. And then here's another chart showing Sonnet 5.5 via the blue line. And as you can see, this is slightly lower intelligence than Opus 5.5. And it's not necessarily cheaper to run. Here are some other charts. If you look at this intelligence index by artificial analysis, then the impressive thing is Sonet 5.5 is number two, even beating GPT6 Astra. But as you can see here, because it's slightly dumber than Opus 5.5, the amount of tokens or cost it takes to complete a task is actually even higher than Claude Fable or Claude Opus.
In fact, here they ranked it as the most expensive model out there to complete a certain task. And one of the reasons is that it tends to produce the most output tokens per task compared to any other model. So honestly, it's not very impressive in terms of efficiency. If you want the best intelligence, I would actually go with Opus or Fable. And if you want the best performance per cost, then I would go with the much cheaper GPT 6.1 Soul or GPT6 Astra.
But the nice thing about this new Sonet 5.5 model is that it's now the default model on the free plan. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, OpenAI held their annual Dev Day event with some noteworthy updates. So, let's go over the main ones. Now, one of the main announcements from OpenAI's Devday is their new Dots product. This is basically like a persistent AI agent that has its own computer and it can keep working even when you're not actively chatting with it.
So, unlike a normal chatbot where you just give it a prompt and wait for an answer, a dot has its own cloud computer and browser. So, it can connect to thousands of apps through plugins. It can keep working towards your goal 24/7. So you can have one continuously watching customer feedback, fixing bugs, testing the changes, and preparing pull requests while you work on something else. Your dot can just take a project and run with it even while it's working on other projects.
So you can just give it new tasks and new projects without having to spin up separate threads. And dots can work across chatgpt, slack, and teams. So you can even chat with it on your phone or any other device. And if this sounds familiar to you, well, it's basically like OpenClaw, but here Opening II just wrapped it into a much easier to use consumer product. And by the way, this is also kind of what Meta is doing with Muse and Grock with Grockbot.
Here they say that DOTs are rolling out today for pro and business premium users in eligible markets. In other words, not in the EU. So if you are on one of those plans, then in the left menu of Chat GPT, you should see your DOT over here. Note that you can call your DOT or just send it a message on Slack. And it has a cloud computer that's always on. And here they say that DOTS is powered by GPT6 Astra. Now the nice thing about the current state of AI is that you can just give it any product and get it to vibe clone the product in a matter of hours.
So literally just a few hours after OpenAI releases dots, we already have an open alternative called Open Dots. So this is an open-source alternative which you can self-host and you could link it to any model or even run the model locally if you have good enough hardware. So this is basically an alternative to OpenAI's dots or Metamuse Grockbot etc. So if you're interested in an open and local option on this page it contains all the instructions on how to download and run this locally.
Of course this is just a framework and you still need to link your own LLM to this. So that's one open- source example. We have yet another one that was also released shortly after the OpenAI's release. And this is also called Open Dots. It's the exact name as the previous tool. And like the name implies, this is also an open- source dot alternative. So you can self-host this and you can assign agents to do various things.
And just like OpenAI's dots, you can also text or call a DOT and give it instructions on what to do. Or you can also ping them on Slack. So it's basically the same thing as OpenAI's dots, but just open source. If you're interested in reading further, I'll link to this main page in the description below. Now, in addition to DOTS, OpenAI also releases GPT 6.1 Soul this week. This is an upgraded version of GPT6 Soul, which was only released last week.
So, I mean, now we're getting model upgrades in just 1 week instead of months. And here they claim that this new 6.1 soul is near Astra intelligence for a fifth of the price. If we look at various agentic coding benchmarks like deep suite, you can see that GPT 6.1 soul is a lot more efficient. So the X-axis is cost per task. Ideally, you want to be in the upper left corner. Same with automation bench. You can see that it's a lot more efficient and just slightly less intelligent than GBT6 Astro, which is the blue line.
Here's a benchmark testing how good it is at operating different interfaces. And as you can see, again, GPT6.1 is just way more efficient. And then here's a benchmark on its scientific research capabilities. Now, if you look at this independent evaluator artificial analysis, you can see that GPT 6.1 is just one point less intelligent than GPT6 Astra. Definitely among the Frontier models out there. And if you look at the cost per task, you can see that it's much cheaper than GPT6 Astra and the overpriced clawed models.
So, this is definitely the most costefficient Frontier model out there. Now, at the bottom, it says GPT 6.1 Soul is available today to all paid plans in GPT work and codecs. It's not available in chat yet. So, if you're on a paid plan, then if you click on this work tab, you should be able to access 6.1 Soul over here. Now, here's the thing. While this new GPT 6.1 Soul is a lot cheaper, one of the other updates from OpenAI that's easy to miss is this.
So basically for the $200 Pro plan, instead of having 20 times more usage compared to the cheaper plus plan, soon they're going to slash it down to just 10 times more usage. Basically, your usage is reduced by half. And by the way, they're also introducing a new plan which will cost $500, which will give you even more usage. If you're interested in reading further, I'll link to this main page in the description below.
If you want to supercharge your creative workflows, definitely check out Luma, the sponsor of this video. Normally, creating with AI means constantly jumping between models, writing new prompts, generating assets, and trying to keep everything consistent. Luma takes a different approach by giving you AI agents that can work with you across the entire project. You can start with a rough idea and have Luma help turn it into an actual creative direction.
From there, you can generate images and videos, test different concepts, make changes, and continuously iterate in the same place. More importantly, the agent remembers what you're working on, so I can give it an idea, generate something, tell it what I like or don't like, and keep developing the project from there instead of resetting the context every time. And when it's time to actually generate your assets, you get access to some of the best image and video models available.
That includes Ray 3.2, Luma's latest video model. Ray 3.2 2 is particularly good at understanding motion, physics, and 3D space. You can have characters moving through a scene, objects interacting with each other, and the camera moving at the same time, while the model keeps everything surprisingly coherent. That makes it especially useful when you're trying to create cinematic sequences rather than just isolated AI generated clips.
So, if AI is already part of your creative process, Luma brings the different pieces together into one workflow with an AI agent helping you along the way, try Luma today using the link in the description below or you can scan the QR code here. Also, this week we have a new update from Comfy UI. If you haven't heard of Comfy UI, this is one of the most popular interfaces for running image, video, and audio models locally.
But at a first glance, this interface seems super complicated, right? with a ton of nodes and noodles and it's sometimes a pain to use. Well, this week they introduced something called Comfy Agent. It's basically an agent built into Comfy UI. So, you can just prompt it on what you want and it'll plan the steps. It'll create and connect the nodes and run the workflow. So, you don't have to like manually drag and drop nodes and noodles yourself.
I think this is especially helpful for technical stuff. So if you ever get stuck on something, if you hit an error message, then you can just prompt this agent right inside Comui to hopefully autonomously fix it for you. Or it can just automatically build or edit a workflow for you and also run the workflow. So you can just ask it to for example generate a video using Miniax and it can automatically pick the right workflow and run it for you.
Now, currently this is available in Comfy Cloud, which is their paid online service, but they will also add this agent feature to Comfy Desktop in a few weeks. So, stay tuned for that. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have new AI called Pixelumm. This is a model that can both understand and generate images and videos. But here's the interesting part.
It doesn't use a decoder. It works directly in pixel space instead of using an encoder. If you have no idea what I'm talking about, basically for most image and video models, they actually generate the content in what is called latent space. This is a more condensed mathematical dimension that makes it easier to work with and then it plugs it through what's called an encoder component to basically convert that into pixel space which you and I can see.
Now the interesting thing about pixelm is that it just completely removes this step. Instead, it does the generation in pixel space, which you and I can see. So, here are some examples of its generations for your reference. I would say the quality and consistency aren't as good as the frontier models out there, but note that this is just a research preview of a new architecture. At the top, they have released the code and the model to this.
So, if you click on this link and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. And note that most of the source files are under the Apache 2 license, which is very permissive. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have a new image model called Aiogram 4.5. And this is what they call the most precise edit model ever.
And as you can see from these examples, you're able to edit existing images with natural language. and it'll only edit specific parts of the image while keeping the rest of the image the same like its pixel level consistent. In fact, if you compare audiogram 4.5 with other image models like GPT image or Nano Banana, you can see that after a ton of different edits, the other competitors just start to introduce a ton of noise in the image.
The image distorts quite a lot. But with audiogram, it stays super consistent even across all these edits. The cool thing with audiogram as with the previous version 4 is you can draw bounding boxes across the canvas to determine what goes in what region. So here's an example where we can draw a bounding box here and only edit that region while keeping the rest of the image completely the same as before. Now currently you can try audiogram on their online platform but they are also planning to open source ideoggram 4.5 so you can run this locally which is fantastic.
I think once they open source this, it will be the best local image editor. If you're interested in reading further, I'll link to this main page in the description below. Now, in addition to ideoggram 4.5, we have yet another new image model this week called Flux 3 image. Now, this also has a ton of really cool features. Just like I v4, you can draw bounding boxes across your canvas and describe what goes in each place.
So this allows you to have very granular control over the composition of the image as you can see from these examples. Now you don't have to use these bounding boxes. You can just use a regular text prompt to generate an image. And this is also super versatile. This can handle a ton of different images including realistic photos or posters with a ton of different elements and text. This can also do reference to image.
So you can plug in up to 10 reference images for it to use in its generation as you can see here. So a very flexible tool and of course this can also edit images with natural language. So you can take any photo and add or remove or replace objects or change the color of things or upscale an image, colorize a black and white photo and other use cases. Now currently this is closed source, but you can try it using their online platform.
They claim that they will open source the weights to this in the future, which is great. But note that they also said this for Flux 3 video, which was already months ago, and there's still no sign of an open model. So, it's not a good track record so far. Hopefully, they will stick to their word and open source both Flux 3 video and Flux 3 image. For now, if you're interested in reading further, I'll link to this page in the description below.
Also, this week, Google is finally back. So, they just unveiled Gemini for Argon. their best and latest model and this thing is an absolute beast. So across various benchmarks on like aentic coding and knowledge work, science and math, computer use, etc. It on average beats GP6 Astra as well as the best clawed models. That's pretty impressive. If you look at this deep benchmark on long horizon agentic coding, then you can see that Gemini 4 is actually ranked number one.
Same with Val's index which measures its performance on a diverse set of knowledge work tasks. Gemini ranks number one. Same with financial research and analysis as well as long horizon legal work. Now if you look at this independent leaderboard by artificial analysis, then you can see that Gemini 4 is not ranked number one but tied for second place with GPT6 Astra and Claude Fable. Still a few points behind Claude Opus 5.5.
However, if you look at the cost of this, then this is way cheaper than the other competitors that have roughly the same intelligence. So, this is actually very costefficient. Now, what I think is the biggest feature of Gemini 4 Argon is that it can output 1 million tokens. That is pretty insane. Like, for your reference, other Frontier models can only output like 64K to 128K tokens. This is basically like how many words it can output in its response.
And 1 million tokens is like 700,000 words in just one output. By the way, this is not the context window or how much information you can feed it, right? All the Frontier models already have a million token context window, but here we're talking about something completely different, which is its output tokens. So, imagine Gemini 4 being able to just output an entire novel or series of novels or an entire codebase in just one answer.
That's pretty crazy. Another thing that I find super impressive about Gemini 4 Argon is it has the lowest omniscience hallucination rate compared to all the other models out there. It only hallucinates 15% of the time according to this benchmark whereas the other models like GPT6 are like 50% and the claw models are all the way over here at around 60 to 70%. So this is super impressive. Now, one caveat to this benchmark is it's basically measuring when a model fails to answer a question, does it give you a wrong answer or a partial answer or just admit that it doesn't know?
All right, so it's not getting the answer correct here. It's basically measuring if it doesn't know the answer, will it just be honest and say it doesn't know rather than make stuff up. So, from this chart, you can kind of view Gemini 4 as being the most honest if it doesn't know an answer. Now, before you get too excited here, they say that they're rolling out Gemini 4 Argon only to a set of trusted cyber defenders through their private program.
So, plebs like you and I won't have access to this yet. And once they do roll it out, it's going to start with paid API customers and Google Aai Ultra subscribers. So, this will likely be paid and only available via the most expensive plan. Hopefully, they will release this soon to the public so I can test the hell out of it and see how it compares with the other Frontier models. For now, if you're interested, I'll link to this main release page in the description below.
Also, this week, Bite Dance releases not one but two new ways to speed up video and image generation. So, the first way is called DMAD, which stands for distribution matching as adversarial distillation. So how this works is you can now use a video model like Miniax to generate videos in just four steps instead of the regular 20 to 30 steps. So here are some examples of its generations with the four-step model. And then here are some other additional examples for your reference.
Now the interesting part is how they train this. So they used the base model to train a new model called the student on how to generate videos on just four steps. And it uses something called discriminator heads to compare the teacher's outputs and the students outputs. And then it uses those differences to figure out how the student should improve. And this same architecture works not only on Miniax but also across other video and image models.
Now at the top of the page, they've released this already. So if you click on this code button and you scroll down a bit here, it contains all the instructions on how to run this locally on your computer. So, you'll need to download the base miniax model, but then these luras are just 1.4 GB each. Now, that wasn't the only four-step distillation method from Bite Dance this week. So, they also released another one called PDMD.
Pretty confusing cuz the previous one is also four letters. Now, like the previous method, this one also allows you to generate videos in just four steps instead of the regular 20 to 30 steps required from Miniax H3. And if you compare this with other distillation methods, you can see that this new PDMD on the far right does perform a lot better compared to the other methods. Here are some other examples for your reference.
You can see the motion, the details, and the consistency are a lot better than the other distillation methods. Now, how this works is also quite interesting. So, for the previous method, we mentioned that we used a critic to tell the differences between the students generation and the teacher's generation and then fix those parts. Well, sometimes the critic isn't perfect, so some of its own errors get passed into the student.
And what PDMD does is it looks at that part of the training signal that comes from the critic error and projects it away. And as you can see from these video and audio quality benchmarks, on average, this PDMD does perform better than other distillation methods. Now, at the top here, they've released the code and the models to this. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this yourself.
Now, here they've actually released two different things. They've released a modified base model, so it's the full miniax model, which is 66 GB, but you can just run this with four steps now instead of 20 steps. Now, if that's too big for you, they also released a Loa for this. So, you can use a smaller quantized miniax model and then just put this Loa on top. And this Loa is only like 1.4 4 GB. If you're interested in reading further, I'll link to this main page in the description below.
Also, this week, we have a pretty wild example of AI being used in real theoretical physics. So, this has to do with plasma, which is basically an extremely hot gas made of charged particles. Now, fusion reactors need to keep this plasma trapped in place using magnetic fields. And one type of fusion reactor called a stellarator uses very complicated three-dimensional magnetic fields to do this. But there's been a mathematical question that's been unsolved in plasma physics since like 1967.
Basically, scientists wondered whether plasma could sit in a perfectly stable and smooth arrangement inside a 3D magnetic field, even when the pressure changes from place to place. And for decades, we thought that this might only be possible if the magnetic field had certain kinds of symmetry. But this week, we had not one but two papers disproving this. Apparently, they used GPT6 Astra to find two different examples showing that other unusual configurations of plasma can exist where they still remain stable and smooth.
And this isn't the first time that AI has helped solved a problem that has remained unsolved for decades. So a few weeks ago I mentioned OpenAI being able to solve this Navier Stokes Millennium Prize problem where it was able to show that a fluid under this certain condition where it spirals like this can achieve singularity. In other words, the middle part, this vortex is able to stretch into an infinite velocity at least for a short amount of time.
Well, this week AI has disproven another problem that has been hanging for decades. This time in plasma physics. And I'm sure more of these problems are going to get disproven or solved with the help of AI, especially problems that have to do with pure math. And once it finds the solution, it's really easy to verify that the solution is indeed correct. So, another pretty cool example this week of AI helping researchers discover some new mathematics in an extremely specialized field of physics.
If you're interested, I will link to the GitHubs of both papers in the description below. Also, this week, we have a new robot system called Tactile Step. This basically gives humanoid robots a better sense of touch through their feet. You see, most robot walking systems mainly care about whether the robot can successfully cross the terrain. But sometimes that can still lead to some unwanted behavior like slamming its feet into the ground or landing on edges or standing with unstable contact.
Well, for humans, we can of course feel pressure under our feet and adjust how we land. So, Tactile Step tries to give robots something similar. The researchers equipped the robot with pressure sensing insoles and then they trained the robot model further to understand where and how strongly each foot is touching the ground. And after training this unitriot, it was able to land more gently and create more stable support while landing.
So you can see in these demos, it's able to walk quite gently across stairs, slopes, and platforms and of course flat ground. In fact, here they report up to almost a 50% reduction in the peak touchdown force as well as reduction in the impact noise and up to 24% more contact area. So, this system allows robots to walk more stably and gently across terrain. If you're interested, I will link to this main page in the description below.
Also, this week, this has got to be one of the most impressive humanoid robot demos I've seen so far. So researchers from Ting Hua and other universities have now trained a humanoid robot, specifically the Unitere G1, to autonomously play Bminton. As you can see from this demo, it's able to do forehands or even backhands and jumping returns. And this can autonomously rally against the human player for multiple rounds.
Now, bminton is an extremely difficult sport for humanoid robots because, well, the shuttle moves extremely quickly and the robot only has a really small window to hit it. Like, it needs to do a ton of decision-m within a few seconds, including how to coordinate its legs, how to run towards the shuttle, and then how to also twist its torso and swing its arm and hand and the racket to hit the shuttle in the right direction.
The problem is that there isn't much highquality human motion data for training a robot on bamin. So the researchers had to take a limited number of real hitting motion and then automatically turn them into different variations of synthetic data. And then this becomes the data set that was used to train the robot. Anyway, a very impressive demo. If you're interested in the technical details, I'll link to this main page in the description below.
Also this week, 11 Labs just released their latest texttospech model, V4, and they claim this is the most emotive voice model so far. In fact, you can do a ton of things in your prompt to customize the generation. For example, you can add metatags in your prompt that dictate the direction or sound effect or tone that you want. Here are some examples. >> You hear that? No, no, no. Don't turn around. I told you not to.
You know, I've lived in these walls since before the paint, before this house, before even you. And every night while you sleep, well, I count your little breaths. One, two, three. You know, your mama would be so proud of the young man you've become. She always told me how much she admired you and ah it's just it's just I miss her so so much. Come closer son. Will you promise me promise me that you'll never forget where you came from?
You here? Swing and a foul out of play. Yeah. So, you and I were talking uh a bit during the break, and you said something that just really blew my hair back. >> Don't do it. >> Yeah. So, Jason here, ladies and gentlemen. >> Oh gosh. >> Was actually the choreographer for the Backstreet Boys, if you can believe it, back in the 90s. Is that right, Jason? >> Oh god, Parker, you're making me feel old, buddy. >> Well, I'm just giving the people the stats.
You know, that's what I'm hired to do, boss. >> Well, okay. Yeah, I can groove. I can get down. You know, I've got the RZ, as the kids like to say. >> Ah, yeah. Yeah, the Riz. Yeah, I got that for sure, bud. You wish straight up full count here. Otani with that patented stance and he cracks IT HARD TO LEFT FIELD. IT IS GOING GOING ME. OH MY OTANI with his second homer of the game. Woo. So really versatile voice model which allows you to sprinkle in metatags throughout the transcript to control the generation however you want.
Now currently this is available even on the free plan. you get like 10,000 credits to start and at least according to this leaderboard then 11 Labs V4 is the best texttospech model out there right now. So definitely give this a try. It's even available on the free plan. If you're interested I'll link to this main page in the description below. Also this week GPT6 Astra apparently helped crack a coded letter from Napoleon which looks like this which has been unsolved for like 217 years.
So this person called Carter Church started with this lowresolution image. He fed it to Astra which basically split the image into rows and it was able to recognize roughly 1,300 handwritten cipher units and then figure out which slightly different looking marks were actually the same symbol. For your reference, a previously published table only gave values for 33 letters covering about a third of the message. From there, GPT built a solver that tested possible mappings by asking whether the resulting text looked like actual French.
And after testing this for a few rounds, it eventually reconstructed the rest of the key, which looks like this. It also discovered that some letters actually represented entire words rather than just individual letters. So, the whole process took around 6 hours, but eventually it was able to decode the entire letter. And here is the answer. It turns out to be a military briefing to General Marmmont describing the troop positions just before Austria went to war with France.
So if you're interested in like history or cryp analysis, here's a pretty fascinating read on how AI can help. If you're interested in the technical details, I'll link to this main page in the description below. Also, this week we have a new AI called Pammy. So, this just takes any text prompt plus any object and it can output a 3D animation of a person interacting with that object. So, it can do a ton of things like picking up objects or moving it, spinning it, etc.
And as you can see, the person actually holds and moves the objects naturally. You can see for the other ones, the hands aren't really touching the object. They're not really connected properly. So, Pammy looks a lot more realistic. Now, here's how it works. It basically lets different body parts like the hands, feet and hips each vote on where the object should move. Then it generates a rough interaction from your text and a refiner cleans up the contact points so the hands actually properly touch the object.
Now if you scroll up they have released a GitHub to this and here they do plan to release the models so stay tuned for that. Note that this is under the MIT license which is very permissive. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have a new AI called point to part. And this takes a single image or a 3D mesh as input, and you can place a point on each part that you want to separate.
And then this AI basically outputs the whole object into these separate parts. So, here's an example of this in action. As you can see, it can take an image or a mesh and break it into clean separate parts, which you can then explode. Now, how this works is quite interesting. instead of generating each part on its own, it decomposes the entire shape based on your point prompts. So the parts never overlap and never leave gaps.
And if you place these parts together, they make up the complete object as before. And if you look at various benchmarks in regards to part quality, you can see that this new model on average even beats other similar segmentation tools. Now, at the top, they have released a GitHub to this. And here it says they do plan to release the models as well as the training code, so stay tuned for that. For now, if you're interested in reading further, I'll link to this main page in the description below.
Also, this week, we have a ton of new and noteworthy small to medium-sized open- source models. So, let me rapid fire through them. The first one is called Astro Brief by Alan AI. This is based off of Quen 38B. So, this is a pretty small 8 billion parameter model. And this is designed to turn scientific literature into properly cited research reports. So as to brief just takes in your prompt plus relevant papers and it can produce an entire cited report in just one pass.
Now as you can see from these benchmarks on report quality. You can see that it performs much better than just the Quen 3 base model and it edges pretty close to another competitor called Dr. Tulu. Here are some other scores for your reference. Now, the nice thing about the Allen Institute is that they fully open source everything. So, here they released the model and the data sets to this. The full 8B model is like 32 GB in size, so you'll still need a high-end GPU to run this, but if you do have the hardware, this is a decent local option for generating reports from scientific research.
Now, in addition to the brief, here's another model released by the Allen Institute this week. So it's called OML core 3 and this actually isn't a new chatbot. It's a new training system that they are using to build mixture of experts models. Basically a mixture of experts model is like a team of specialist AIs working together. And when you use it only a few of these experts are active at a given time. So this gives you a huge model without paying for the full compute cost of running the entire thing every time.
The problem is that once these models get really large, moving data between hundreds of GPUs becomes a major bottleneck. So, OMO Core 3 is designed to fix this problem. The main fix was to keep each expert in its own GPU and send the work to it instead of repeatedly moving the expert itself, like sending jobs to workers who stay at their desks. They also spread the model and its training data across GPUs to save memory and grouped smaller calculations together so that the chips could work more efficiently.
And together with this framework, it actually made training like 2.7 times faster compared to their previous system. Now, the nice thing is they've open sourced how they did this. So, if you're interested, they released a technical report as well as this GitHub which contains the documentation on how they ran this. So, you can inspect these links up here to learn more. If you're interested, I will link to this main page in the description below.
Also this week we have another AI called iquest Q1 and this is a new agentic model for CLI systems. So command line tools like cloud code and codecs. The goal for this model is to work through long software tasks like inspecting repos, running commands, debugging failures, using tools, etc. and keep iterating until the task is finished. Now this is quite huge at 320 billion parameters. This is a mixture of experts models.
So only 15 billion parameters are active when you use it. And as you can see from these benchmarks, it actually performs pretty well given its size. Now, at the bottom, they released the code to this already. So, if you click on this link, it contains all the instructions on how to download and run this. If you're interested, I'll link to this page in the description below. Also, this week, we have a new model called Ax 2, which is a pretty tiny 27 billion parameter model based off of Quinn, and it has a pretty interesting design.
So, instead of trying to solve a problem once, it learns to repeatedly improve its own answer. Basically, the model proposes a solution. It then tests or measures how well it worked. It looks at the feedback, figures out what went wrong, and then it tries again. So, think of it less as like asking the AI a question once, and more like giving it several rounds to think through and debug its own work. In fact, ARX 2 was trained heavily on machine learning and algorithmic programming tasks where the result can actually be checked.
And if you look at these benchmarks, it's even able to perform on par with models that are much larger. It's also super efficient. So, for example, if you look at its performance versus the number of parameters, it performs even better than like Kimmy K3 or GPT 5.6 Soul. Same with these Deep Research benchmarks. The nice thing is they've released this already. So, if you click on this models link at the top, it takes you to their hugging face where you can download the model.
Now, at 27 billion parameters, this is still quite huge at 55 GB in size. If you're interested, I'll link to this page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite? And which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you.
So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up to date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.