Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
8,182
Runtime
43:53
Speaking pace
186wpm
Reading time
34min
186 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
AI never sleeps, and this week has been absolutely insane. The mysterious stealth model under the name Aux Alpha has been revealed this week to be GLM Flash. They also released the model weights to the full GLM 5.3, which is now the best open model you can use right now. Tencent also releases their best model Hi 4, which is also open source. And then Alibaba also releases their latest model Qwen 3.8 Flash Next, which is also very close to Frontier. The best AI video generator Minimax just got better, so
93 words, the words spoken in the first 30 seconds at 186 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 465 |
| Average words per sentence | 17.6 |
| Longest sentence | 65 words |
| Questions asked | 8 |
| Sentences containing a number | 124 |
Most used terms
Filler phrases
108 in total: like 47 · actually 26 · basically 22 · I mean 6 · kind of 5 · right? 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
AI never sleeps, and this week has been absolutely insane. The mysterious stealth model under the name Aux Alpha has been revealed this week to be GLM Flash. They also released the model weights to the full GLM 5.3, which is now the best open model you can use right now. Tencent also releases their best model Hi 4, which is also open source. And then Alibaba also releases their latest model Qwen 3.8 Flash Next, which is also very close to Frontier.
The best AI video generator Minimax just got better, so you can now potentially run this in real time. We also had the world humanoid robot games this week with some insane results, which even broke human world records, and also some very cute and hilarious clips. Google DeepMind continues to cook some really useful stuff. So, this week they released an AI prediction model for mapping things like disease risk and food security.
Google also releases their latest video generator and transcription model. We have a new AI world generator, but this one can actually maintain a consistent memory of the world and the objects and the rules. We also have a ton of new open-source 3D generators and a lot more. So, let's jump right in. First up, we have a new 3D model generator called Block 3D, and this is able to turn a text prompt into 3D objects. In simple terms, how this works is it breaks a 3D object into blocks of shape tokens, and it generates these blocks one after another, but it generates all tokens within each block in parallel using diffusion.
And the speed difference here is pretty significant. So, here it says it's over five times faster than standard auto-regressive methods, and on average it only takes like 5 seconds to generate a 3D model. While it's not the best quality or the most detailed, it is among the fastest 3D model generators you can use. At the top of the page, they've released the code to this already. So, if you click on this code button, and you scroll down a bit, here Here contains all the instructions on how to download and run this locally on your computer.
Plus, they also released the code on how you can train this yourself. If you're interested in reading further, I'll link to this main page in the description below. Also this week, this AI is very useful. It's called one video one world, and this can turn a normal video into an animated 3D world with individual objects. Here's an example where on the left is our input video, and it can basically render this into a full animated scene with separate 3D meshes that are simulation ready.
Or here's another example. So, as you can see here, this AI basically separates the individual objects and reconstructs each one as its own 3D mesh, and it also has to figure out how these objects move through space. It also has to work out like the relationships of each object and how close they are to each other. The researchers here actually combine a ton of existing foundation models together, including Kwen 3VL for understanding the scene, and then SAM 3 for segmentation, Flux 2 for filling in hidden parts of objects, and then finally High 3D Gen for 3D reconstruction, and also other models for camera and pose estimation.
The nice thing is they've released everything already. So, at the top of the page, if you click on this code button, and you scroll down a bit, here it contains all the instructions on how to download and run this locally on your computer. Note that here it says over 40 GB of VRAM is recommended, plus their own code is MIT licensed, but it also uses these other platforms, which are under non-commercial licenses. If you're interested in reading further, I'll link to this main page in the description below.
Also this week, we have a new AI called fix anything. This is basically a cleanup system for broken or ugly 3D renders. You see, when you reconstruct a real environment using Gaussian splatting, nerves, meshes, or other methods, the result can look okay from familiar camera angles, but they often start falling apart when you try to view the scene somewhere else. You might see holes or floating geometry or blurry areas or other artifacts, as you can see from the video on the left.
Well, what fix anything does is it takes this degraded render and it uses a pre-trained video model to turn it into a cleaner, more realistic scene while keeping the same camera path and overall 3D structure. So, here are some examples for your reference. You can see it's able to clean up these really degraded scenes very well. The cool thing is this also works with mesh inputs as well as sparse points inputs like this.
Here's another cool aerial example. Now, at the top of the page, they've released this already, so if you click on this code button and you scroll down a bit, here it contains all the instructions on how to download and run this locally on your computer. Now, this uses 1 2.1 as the base model, which should be able to run on most consumer GPUs. I'm surprised they didn't use the more recent version 2.2 or MiniMax, but in any case, this is basically a LoRA for 1 2.1.
If you're interested in reading further, I'll link to this main page in the description below. Also this week, Google Research continues to release some very useful stuff. So, this week they released something called the Planetary Prediction Engine. This is basically like an AI data scientist for problems involving Earth. You can give it a question in natural language like predicting disease risk or identifying areas that might face food shortages, and the system automatically finds relevant geographic data, cleans and combines it, and then trains machine learning models and evaluates them to make predictions.
This is a huge deal because normally specialized teams might need to spend weeks doing this manually. Plus, you also need to have the technical expertise to pull this off. But, Google says this Planetary Prediction Engine can basically reduce all this work into just minutes. So, here are some example results. It was able to integrate a ton of different data sources to map out food security in Nigeria, and they were able to improve accuracy from 31% to 66%.
Basically more than doubling the baseline accuracy. In terms of predicting the DRC Ebola outbreak, it also was able to achieve a much better score of 88% which is over a 10% improvement over the previous state of the art baseline. So I mean across a ton of different use cases like predicting diseases, food security, etc. This new planetary prediction engine is often a lot more accurate than previous state of the art baselines.
So a super fascinating project that makes use of AI agents and the data from Google Earth. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a very interesting AI called Code World Model. And this is an interesting attempt to build AI generated worlds that actually remember what's happening. So most video world models mainly predict what the next frames should look like.
This might look impressive but they don't actually maintain the underlying rules or objects or consequences of the world. So this project actually separates these two jobs. We first have a coding agent that acts as the brain. So when something happens it actually reasons about the event, writes code, and updates a persistent world state. That state is then converted into a simplified proxy describing things like where the objects are and how the scene is laid out.
And finally it then uses a video model, in their case a fine-tuned minimax H3, that takes the proxy and renders a video from it. So think of it like a video game where the game engine is written by a coding agent while you get a video generator to actually handle the graphics. And the demos show some really long running worlds that can change styles while still preserving the underlying scene and events. You can get the characters to interact with the objects.
You can see there's a map at the top corner which keeps track of all the objects in the world so everything is is more coherent than some of the other AI world generators. And think if we do plan to create like AI-powered video games or interactive worlds, this is a necessary step. We can't just get a video generator to generate everything. We also need an agent to keep track of everything in code so that the world and all the objects and the rules stay consistent across the entire scene.
The awesome thing is they've released the code to this already. So, if you click on this GitHub repo, it contains all the instructions on how to download and run this locally on your computer. Note that this does use a fine-tuned Minimax H3. They tested it on an H800 with 80 GB of VRAM, but you can probably use a quantized version if you have lower VRAM. If you're interested in reading further, I'll link to this main page in the description below.
Also this week we have a pretty interesting AI called Voice Mem. This is basically a memory system designed specifically for voice agents. Instead of treating every conversation like it's starting from scratch, this lets the AI remember useful facts about you when you're using it. So, the system splits memory into two parts. It's left brain stores factual information like people, entities, and knowledge, while the right brain tracks things like your preferences, personality, and emotions.
And this whole thing is designed to run in real time. So, here's an example of it in action. >> What kind of person do you think I am? >> Jati, you crave real connection. Like how you always pick those quiet one-on-one talks over big crowds. And you're thoughtful. Remember that time you were wrestling with dinner ideas, really putting your heart into making it just right? That's the kind of care and attention that shows who you are. >> I'm so tired lately. >> Is it about the algorithms midterm back in April?
That C really hit you hard. Feeling stuck from that? >> So, as you can see, it can keep a persistent memory of you. And this is designed to be very efficient. So, instead of just stuffing thousands of old conversations back into the model every time, it only searches its memory graph and fetches only the few memories that matter. So, as you can see, in terms of memory accuracy, this new voice mem tool is even better than other competitor methods.
Plus, it's also way more efficient. So, look how much fewer tokens it uses. The latency is also incredibly fast. This is like sub-second. So, a very useful tool in designing like personalized voice assistants or companion agents where about past conversations are useful. The awesome thing is they've released everything already. So, at the bottom here, if you click on start building, it actually takes you to their GitHub repo.
And if you scroll down a bit here, it contains all the instructions on how to download and run this. If you're interested in reading further, I'll link to this main page in the description below. Also this week, the best open-source video generator out there, Minimax H3, just got way faster. So, the How AI Lab just released fast video for Minimax. And this can make video generation as much as 14 times faster on select GPUs.
Now, if you happen to have 8 B200 GPUs, then you can even run this faster than real time. So, it can generate a 15-second clip in less than 30 seconds. Even if you don't have such high-end GPUs, they are working on variants that can support consumer GPUs as well. Now, the way they achieved this speed up is pretty clever. So, they first use something called DMD2 distillation, which basically teaches a smaller, faster generation process to imitate the original model while using much fewer steps.
Then, they also combine that with video sparse attention, so the model only pays attention to the most useful parts of the video data instead of comparing everything with everything. In fact, their recommended version keeps only around 10% of the video attention data, which cuts a huge amount of computation. The most impressive part is that this is data free. So, they only use the already trained base MiniMax model as the teacher, and the teacher generated examples to train this fast H3.
They didn't need to go out and collect additional training data from scratch, nor did they need the original data set that trained the base MiniMax model. The awesome thing is they've released this already. So, at the top of the page, if you click on this GitHub repo, it contains all the instructions on how to download and run this locally on your computer. They also have an Apple Silicon guide as well. And on their Hugging Face, you can see they've already released the model.
The transformer itself is quite huge at 70 GB in size, but the Goat Kijai has already released a much smaller fast video checkpoint that works for ComfyUI. So, this is only 22.9 GB in size. And at least at the time of this recording, Kijai has created a PR to merge this into ComfyUI. So, right now it might not work yet, but give it a few days for them to merge this PR. Anyways, props to this team for making MiniMax so much faster.
Because for me, speed was the main issue. I usually have to wait a long time to generate videos with MiniMax. Anyway, if you're interested in reading further, I'll link to this main page in the description below. In fact, that wasn't the only MiniMax speedup that we got this week. So, Stability AI also released their own MiniMax fine-tune called H3Max, which is also insanely fast. You can see that this new H3Max basically finishes within seconds.
Here's a table on its performance. This design looks horrible, by the way, as if I'm like accidentally highlighting the numbers. But anyways, you can see that H3Max not only gets the best quality score, but also the lowest average latency. Now, here's the thing. While it is impressive that they were able to speed up MiniMax by a lot, currently it is behind a paid and closed API. Not sure how I feel about taking an open-source model, MiniMax H3, and then just optimizing a bit, and then putting it behind a paywall for profit.
If that's the case, I hope the MiniMax team at least get some rev share from this. Now, there are some comments from the file team saying they will open source this. So, hopefully they will stick to their word. For now, if you're interested in reading further, I'll link to this main page in the description below. If you want to streamline your entire creative workflow, definitely check out Luma, the sponsor of this video.
Instead of jumping between a bunch of different AI tools, Luma agents gives you one agentic workspace that can work with you across an entire project from the first idea all the way to the finished result. This is much more than a one-off prompt tool. You can give an agent a real creative task, share your assets and references, and continue working with it as the project evolves. It understands broader context, helps you make changes, and keeps everything organized inside the same workspace.
And now, Luma hosts one of the best video models, Sea Dance 2.5, along with Kling 3 and its Wraith 3.2 model. So, you can choose the best model for your projects. But, the feature I'm most excited about is Luma skills. A skill is basically a reusable AI workflow that you only have to create once. You define the instructions, save them, and then run the exact same process on any new image or video. For example, let's create a skill that takes any basic product photo and drops it in water on a white background with a nice splash effect.
Once the skill is built, I can drop in a completely new product image and run the entire workflow automatically. The composition and visual style stay consistent without me having to repeat all the instructions again. You could build skills for applying a brand kit, generating influencer videos, changing the weather in a scene, swapping outfits, and much more. Whether you're creating marketing campaigns, branded content, product visuals, or social media content, Luma agents allows you to do everything efficiently in one intelligent workspace.
And with skills, you can turn your best creative processes into repeatable workflows that work on any asset every time. Try Luma using the link in the description below or scan the QR code here. Also this week, the mysterious model under the stealth name Ox Alpha, which first showed up on Open Router where you can use it for free. They claim that they have capacity for 100 trillion tokens per day, which is crazy. Well, it turns out that Ox Alpha is actually JLM 5.3 Flash.
Even though this is just a flash model, it's actually a huge deal. Now, the best part about this is they've added vision to the model. Whereas the previous JLM models did not have vision. So this makes it like way better for front-end development or you can give it images or documents or even videos to analyze and understand. This is fairly tiny with only 320 billion total parameters and this is a mixture of experts models.
So when you use it, only 18 billion parameters are active making it super efficient. Note that other frontier models are usually over a trillion or even 2 trillion parameters. So not only is this smaller, but if you check out these benchmarks then not only does it beat JLM 5.2 and Opus 4.8, which are much larger, but it's as good as even the medium-sized GPT 5.6 Terra. Now, as a flash model, I don't expect it to be like frontier intelligence, but the most impressive thing about this is if you look at the price of this, it's dozens of times cheaper than any other competitor.
So in terms of intelligence versus cost, this is by far the best option you can use right now. If you look at this leaderboard by Artificial Analysis, note that JLM 5.3 Flash is edging very close to the top frontier models, which are like over a trillion parameters in size. And the cost of this is just absolutely absurd. So it only costs like 9 cents per task, which is way cheaper than Kimmy K3 or Grok 4.6 or the best GPT.
Here's another chart comparing intelligence versus price and as you can see, JLM 5.3 is by far the most efficient option out there. The thing I really like about all the GLM models is that it has among the lowest hallucination rate. So, as you can see, GLM 5.3 flash only hallucinated 20% of this benchmark, which is really good. If you compare this to like Opus 5, it's 60%. So, Opus hallucinates three times more. And then GPT 5.6 is even worse, like 80 to 90%.
And this thing is incredibly efficient. They basically combined a linear attention with sparse attention. So, instead of constantly comparing every token against every other token, it compresses some information and selectively focuses on parts that matter. And by doing this, they basically cut attention computation by roughly three times, and it also shrinks the model's working memory, also called the KV cache, by 4.4 times compared to the full GLM 5.3.
Now, probably the most mind-blowing thing about this is they were able to serve GLM 5.3 flash completely with Chinese AI chips. It's probably related to this piece of news a few months ago where ZAI buys a massive 1 gigawatt data center powered completely by Chinese chips. And I mean, this is a big deal. Chinese AI labs are no longer limited by access to Nvidia chips. Now, as with the previous GLM models, they've released this for you to download for free locally, which is fantastic.
So, if you click on this Hugging Face link, note that the total size of this flash model is fairly tiny at only 328 GB in size. So, it should be able to fit on some really high-end consumer hardware. Now, because this is open source, the community is quick to build a ton of different variants of this. For example, Unsloth has already released GGUF versions of this, and as you can see, the one-bit version is only 93 GB in size.
Isn't that crazy that we now have Opus 4.8 level intelligence, which you can run locally for free? Anyway, my favorite lab, ZAI, has done it again. They managed to squeeze frontier level intelligence into something that's ridiculously cheap to run. So, if you're deploying agents and you need to process millions of tokens every day, then this is by far the best option to use. If you're interested in learning more, I'll link to this main page in the description below.
Now, in addition to GLM 5.3 flash, as promised, Z AI also releases the model weights to the full GLM 5.3 this week. Now, this does not have vision capabilities, unlike the flash version, but this is the most intelligent and performant open-source model you can use right now. As you can see, it's pretty much tied with Qwen K3, but this has way fewer parameters. So, on this page, if you click on this hugging face link, here is the full model.
Now, this is 753 billion parameters, and the total size of this is 756 GB in size. However, because this is open-source, the community has already created some quantized or compressed versions of this. For example, Unsloth has released some GGUFs, and the 1-bit version is only 217 GB in size. So, you just need to stack like two or three high-end GPUs to run this. I mean, if you look at this intelligence index, this scores really close to the best GPT and the best Claude.
There's only like a one or two point difference. And the crazy thing is, you can now run this level of intelligence locally on your computer. This is currently the best open model you can run locally. So, if you're interested, I'll link to this main page in the description below. Also this week, Alibaba also releases their latest flash model, Qwen 3.8 flash next. And this is actually a huge deal. First of all, this is an early preview of the architecture that's used in their upcoming Qwen 4.
Heck, they just released 3.8, but they're already teasing Qwen 4. So, I hope you can feel the acceleration. Now, compared with GLM 5.3 flash, Qwen 3.8 flash is much smaller at only 125 billion parameters, and this is also a mixture of experts. So, when you use it, only 6 billion parameters are active, making it very efficient. Here's a table of its performance across various agentic, coding, and knowledge work benchmarks.
And as you can see, Qwen 3.8 Flash next even beats DeepSeek v4 Flash, as well as Claude Opus 4.6 Max. Although it's strange why they put Opus 4.6 and not 4.8 or a newer version. But anyways, across most of these benchmarks, you can see that Qwen 3.8 on average scores the highest. Now, on this leaderboard by Arena, where people can blind test different models, you can see that at least in terms of web dev, Qwen 3.8 Flash next actually does better than GLM 5.3 Max, and it's pretty much tied with GPT 5.6 Soul, which is pretty crazy.
However, it does lose to the latest GLM 5.3 Flash, which is just insanely good. Here, if you look at this leaderboard by Artificial Analysis, then you can see that Qwen 3.8 Flash next is tied with Gemini 3.7 Flash. It does score one point below GLM 5.3 Flash, but again, keep in mind this Qwen Flash model is actually way smaller. But it's even able to beat other models like DeepSeek v4 Pro, as well as GPT 5.6 Luna Max.
So, still very impressive. Now, this is open weights, but if you choose to use it through their API, then this is also insanely cheap. It only costs like 10 cents per task, whereas GPT costs nine times more. And let's not even talk about the Claude models, which are just ridiculously overpriced. Now, how they designed it is also very interesting. Like, how are they able to squeeze so much intelligence from just a hundred billion parameter model?
Well, first of all, they implemented with this additional 51 billion parameter N-gram embeddings, which you can think of as like a giant lookup memory for local patterns. The biggest architecture change, though, is how it handles long context. It turns out that three out of every four layers in its architecture uses something called gated Delta net. This continuously compresses everything the model has seen into a much smaller memory state.
So, this makes it much more efficient at remembering things. But then the remaining layers use something called Qwen sparse attention to go back and retrieve specific information when needed. So, think of it like reading a giant book while keeping compact notes or summaries, but then it's also able to jump back to the exact paragraph and find a specific detail when it needs to. Now, as with the other Qwen models, Alibaba has also released this Qwen Flash Next model for you to download locally for free.
If you click on this Hugging Face link, here is where you can download the models. So, the full Flash Next model is 360 GB in size. They also released an FP8 version, which is only 186 GB. Unsloth also released some GGUFs of this, and the smallest 1-bit version is only 72.5 GB in size. So, you can even potentially fit this on just one RTX 6000. Pretty crazy how we now have Opus-level intelligence, which you can just run on some high-end consumer hardware.
If you're interested in reading further, I'll link to this main page in the description below. Now, in addition to Qwen Flash Next and also GLM 5.3 Flash, Tencent also releases their best model yet, Hi 4 Preview. And as with most of the Chinese labs, they've also open-sourced this. First of all, here are the specs. So, this is a fairly large model at 770 billion parameters. This ain't a flash-size model. It's also a mixture of experts, so when you use it, only 49 billion parameters are active.
To make it so efficient, they also took some inspiration from Deep Seek and GLM. So, the architecture uses gated Deep Seek sparse attention and also index cache, and it also uses these identity hyperconnections to optimize information flow between layers. And check out the performance of this. I did not expect a 10 cent model to be at the frontier, but it looks like they've caught up as well. So, basically the dark blue bar is high four and the light blue bar is high three.
So, notice the insane improvement over the previous generation. Also note that this is currently just a preview model. As you can see, this already beats Qwen 3.8 max as well as DeepSeek v4 Pro and it scores pretty close to the leading open model GLM 5.3 and Kimiko 3, which is very impressive. Now, at the time of this recording, it doesn't look like high four is added to this artificial analysis leaderboard yet, but if we look at this web dev leaderboard by Arena, you can see that high four preview is actually tied with GLM 5.3 flash.
So, this is one of the best open source models you can use, at least for web dev. Now, here they've released the model to this already, so at over 700 billion parameters, this is quite huge at 1.56 terabytes in size, but the open source community is quick to act and this person, which I believe is also from Tencent, they've already released a GGML version of this where the one-bit version is only 229 gigabytes in size, which is way smaller.
Anyway, on this page it contains all the instructions on how to download and run this locally on your computer. And this is under the Apache 2 license, which is very permissive. If you're interested in trying this out and you do have the hardware to run this, I will link to this page in the description below. Also this week, we had the 2026 World Humanoid Robot Games in Beijing. And this was way different from what we had last year.
You see, last year the robot were still pretty clunky, they moved pretty slow, and there were a ton of bloopers. But this year, everything improved significantly, to the point where we had several human world records broken. I mean, the speed of improvement of humanoid robots, especially for these Chinese labs, are just absolutely insane. Now, one of the first highlights is this 100 meter race. Among the fastest robots were this red one by Honor and the other one which took first place and this was the Tiangong robot.
It reportedly finished 100 m in only 9.39 seconds which beats Usain Bolt's world record of 9.38 seconds. So, I mean, we already have humanoid robots that can outrun humans. Now, one of the cutest clips from the robot games is this 400-m race which involved mostly this Tiangong Omni robot which runs like a shy anime girl. But, apparently, this posture allows it to run really quickly. It seems to be a speed hack where it keeps its arms close to its torso and it leans forward which apparently makes it run faster.
Now, it learned this posture through reinforcement learning so it probably tried a ton of different variants of running styles and this one worked the best. So, maybe we've been running wrong our whole lives. Maybe this is actually the correct way to run. So, anyways, this shy embarrassed robot is the Tiangong Omni which won the 400-m race. Now, in terms of the running long jump event, here is the winning clip from team Tian Zuo's Tiangong Ultra robot.
In fact, this robot was able to jump 7.9 m which is still under the human world record of 8.95 m. But, instead of the running high jump, if you look at the standing high jump event, then again, this Tiangong Ultra took the gold medal. And here it jumped for 4.8 m which does beat the human world record of only 3.73 m. So, that's another human world record broken. Now, instead of the Tiangong Ultra, we also have this cute little robot called the Mini Pi trying its best to do a long jump.
It's nowhere close to smashing a world record, but I would give it an award for being super cute. And while the Tiangong Ultra is super impressive, not all jumps were successful. So, here's an example of a face plant where the robot was not able to maintain balance after the jump. So, this is actually really challenging. Not only does it have to jump with explosive power, but it also has to land and maintain balance.
In addition to long jumps, we also have the high jump event. And again, the Tiangong Ultra took gold for this event. So, here is the winning clip, and it was able to clear 2.88 m, which again beat the human world record of only 2.45 m. I mean, this thing has some explosive power to be able to launch itself so high and so far. A super impressive design. Now, in addition to jumping, we also have humanoid robot soccer. And here we can see some humanoid robot teams playing soccer against each other.
It's pretty slow and chaotic. Sometimes they bump into each other. There's not really much coordination, but this is really complicated. Not only does it have to track the ball, but also the positions of the teammates and opponents. And then it also needs to decide whether to pass the ball or shoot or defend or some other actions. So, it's pretty sloppy and slow and full of collisions, but still pretty entertaining to watch.
And then, in addition to jumping and soccer and running, we also have ping pong. And here you can see we already have humanoid robots that can autonomously play table tennis. So, this is the AGI Bot A3, and they took first place. We also have tug of war. And this is one of the best force control demos at the games. This is actually very challenging. The robots need to coordinate grip force and rope angle and balance data, and they need to work together to coordinate a pull.
As soon as the rope angle changes, notice the hips of the robots. They kind of drop into a squat, and then they lean back and also re-center their bodies. You can see them adjusting the pull and their stance in real time. And then we also have this martial arts event where a robot completes a full routine. It's able to do a full kung fu performance with a ton of really impressive moves. And instead of martial arts performances, we also have a break dancing competition, and this one is pretty cute.
As you can see, it's able to do a handstand, it can spin and kick and do some other impressive breakdancing tricks. Not only is this pretty cute, but it's also much better than what I saw from the human Olympics. And then we also have this 100-m obstacle race, and here's the winning clip from the AGI-bot X2 robot. I'm going to speed this up, but here you can see it walking through a ton of different steps and different obstacles, and needs to do this all autonomously.
But as you can see, it's able to successfully traverse this entire obstacle track. And it's actually pretty hard. A ton of other robots were not able to complete the course. So, those were some highlights from this year's humanoid robot games. This is definitely a massive improvement from just last year. I can't even fathom what next year would look like. At least to me, this is pretty entertaining. Let me know in the comments what you think.
Would you actually be interested in watching this? Also this week, Google releases their latest transcription model called Gemini 3.5 Transcribe. This basically takes an audio and it can turn it into a text transcript. The nice thing is, this actually has two different modes. There's a verbatim mode that basically records everything exactly as spoken. And then there's also a smart transcription mode, which helps you get rid of errors, filler words, and basically automatically formats and cleans up your text. >> I want a large iced latte with oat milk, two pumps of hazelnut, an extra shot, and light ice.
Actually, make the hazelnut into caramel. Remove the extra shot and make it regular ice. Wait, no ice at all. Change the size from large to medium. >> This new 3.5 Transcribe model can also detect multiple speakers and word-level timestamps. It's also multilingual, so here's an example where it can take audio consisting of different languages, and it's able to automatically detect these languages and generate the correct transcript. >> It's particularly good at multilingual transcription.
So, I can start a sentence in English y terminarla en español. Or maybe Hindi may bolna shuru kar sakta hoon. And the model doesn't miss a beat when I switch between English, Spanish, and Hindi one after another. The model is also great at alphanumeric transcription. This year, I'm planning it for the Beijing. >> This thing is incredibly fast. So, they also have a live version which enables real-time streaming. This delivers sub-second latency for interactive voice agents or real-time captions.
And this isn't just a separate standalone tool. They've already added it to the Gemini Mac app, for example, where you can use your voice to trigger Gemini models to do actions like analyze files, generate images, or control what it sees on your screen. So, instead of just typing text prompts to control your agents, you can just use your voice to dictate your commands. At the bottom here, it says that this new 3.5 transcribe model is already available via AI Studio or Anti-Gravity, as well as these enterprise platforms.
And for consumers, it's already available in the Gemini app for macOS, Rambler on Android, and it's also coming soon to Chrome. If you're interested in reading further, I'll link to this main page in the description below. Now, in addition to a transcription tool, Google also releases their latest video model called Gemini 1.1 Flash. This is just a flash model, and it's just point one, so it's not like a huge upgrade.
But, here are some notable new features. It can now take an existing video and keep extending it while looking at up to 10 seconds of previous footage, rather than relying on just the last frame. So, this would keep the characters and environments much more consistent. You can also give it a first frame and last frame, and the model generates the movement between them, which is useful for things like camera orbits or zooms or seamless loops.
Another practical addition is cheaper 360p draft generation. So, you would generate a draft first in low resolution, which Google says would be up to 60 times faster and would cost cheaper, and then you can test several ideas before upscaling the final version to 1080p or 4K. Now, currently this is already out, so you can try Omni 1.1 directly in Google's AI Studio. Over here, simply click on video, and here's where you can try out the new Gemini 1.1 Flash.
If you are on a paid plan, then this is also available in Google's Flow. If you're interested in reading further, I'll link to this main page in the description below. Also this week, Google DeepMind releases a pretty interesting benchmark called Orbit Plus Plus. In a simple sense, this is basically a new torture test for AI systems to try to understand how a camera is moving through the real world. And you know, this is an important application for things like 3D reconstruction or robotics and computer vision.
Now, what Orbit does is it starts with real 360° videos from the internet. It figures out the camera's motion using full panoramic view, and then it crops out normal perspective videos containing much harder situations. So, they basically used AI to make a much trickier video that looks like this to test other AI models at reconstructing this 3D scene and predicting the camera movement. And if you feed these clips to existing prediction systems like COLMAP or Mega SAM, etc., you can see a lot of them fail.
So, that's the point here. Orbit Plus Plus is basically a dataset of these much harder videos designed to test 3D vision and 3D reconstruction AI models. So, if you're building related 3D reconstruction models and you really want to stress test your model, then this is a new and fascinating way to do so. If you're interested in reading further, I'll link to this main page in the description below. Also this week, Xiaomi shows off something called the AI Cube.
And this is basically a tiny desktop computer similar to the DGX Spark, which is designed to run huge AI models completely locally. So, instead of sending your prompts to the cloud servers, you can just have a decent-sized model over 100 billion parameters sitting right there on this machine, which you can run offline. This means more privacy, no subscriptions, no dependence on any internet connection. So, Xiaomi built this whole thing around three of its X ring chips, the O3, the O100, and the D100.
The O3 is like the main processor, so it includes a 10-core CPU as well as a 16-core GPU and a neural processor. And then the O100 is specifically designed to move data extremely quickly for AI with up to 1.22 terabytes per second of memory bandwidth. And then we have the D100, which is Xiaomi's chip that is specialized for AI. And if you put all three together, you get this beast of a device, the AI Cube, which can run both a massive 120 billion parameter model and a much smaller 3 billion parameter model locally and at once.
This is really important because you can switch between them whether you need a quick response from a smaller model or much heavier reasoning from the larger model. The prototype reportedly has 80 GB of unified memory, but it can also support up to 160 GB of unified memory, which in theory can fit models of up to 200 billion parameters. And honestly, I think these devices will be the future. Instead of having an ugly and bulky GPU rack to fit local models that are hundreds of parameters, instead it's just going to be a tiny device that sits on your desk, which is optimized for AI.
So, this could be what a future personal AI computer could look like. For now, it's still a prototype, so Xiaomi hasn't announced a price or release date, but rumors say that it's going to be sometime next year. Also this week, ByteDance releases a very interesting training method for image models, which helps them improve more efficiently. So, let me try to explain this in really simple terms. Normally, when you train an image generator, you have a reward model that tells whether the final image is good or bad.
And then somehow the image generator has to figure out which of its many intermediate steps caused the problem. And that's a very vague signal, so it's not really helpful. What Diffusion OPSTD does instead is it turns that final reward into much more direct instructions about which intermediate steps actually caused the problem and how to fix them. Basically, the reward system figures out which direction would make those outputs better and worse.
So, it creates a nearby positive target and a negative target, as you can see here. The model then trains directly toward that better target and away from the worse one. And the results are pretty strong. So, here is the prompt and then on the left we have the base model without any fine-tuning, and then the middle three are some other training methods, and then on the far right is this new Diffusion OPSTD training method.
As you can see, it produces the best-looking image out of the four. Or here's another example where you can see the base model is not able to render the text correctly, and the other training methods also look pretty awful, but Diffusion OPSTD was able to train it to generate the best result. Here's another example for your reference. The funny thing is if you write "bicycle kick" in your prompt, the base model actually generates a bicycle, which is not correct.
But after plugging it through Diffusion OPSTD, it's able to understand the context a lot better and generate a legit bicycle kick without a bicycle. Here are some more quantitative results for your reference. So, it's able to achieve a 44% better quality gain over the previous strongest competitor. But this is also extremely efficient, so it took 40% fewer training hours compared to another method for Stable Diffusion 3.5 and 63% fewer hours for Z image turbo.
And this is model agnostic, so you can also use this to fine-tune other diffusion models as well. The awesome thing is they've released the code to this already, so if you click on this GitHub repo and you scroll down a bit, here it contains all the instructions on how to download and run this locally on your computer. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a new open-source image generator and editor called Fibo 1.5 by Bre AI.
Not only can this take in a normal text prompt, but you can also feed it a really detailed JSON prompt like this, where you can break down different aspects of the photo using different JSON fields. There's one for description, another one for location, relative size, shape and color, texture, etc. And it's able to generate an image that follows all this information quite well. Now, you don't have to use such long and detailed JSON prompts.
You can just input a normal text prompt, and here are some examples for your reference. I wouldn't say it's the best quality image generator out there. My current favorite is still Create 2, but the nice thing about this is you can also edit existing images with natural language. For example, here we can add a blanket draped over the armrest like this. And then we can iterate further. So, for the next prompt, we can also add some hardcover books on the floor.
And then for the next prompt, we can also add a cat on the chair. Honestly, not the best quality image model out there, but if you are interested in trying this out, I'll link to this main page in the description below. Now, last week I featured this robotics foundation model called Gen 1.5, which is able to kind of learn new things. So, it can just watch a demo of a task it has never seen before, and it could successfully carry out that task most of the time.
Well, this week we have another similar model called S1. You see, before in order to teach a robot a completely new task, it often means collecting hours of demonstrations, and then fine-tuning or training the model further specifically for that job. But with S1, you can just show the robot one video demonstration of what you want it to do, and it can attempt that task without any additional training. So, here they demonstrated this on completely new tasks, which the model has never seen before, including repotting a plant, making pour-over coffee, assembling a kit, and even cooking and flipping a pancake.
Some of these tasks last up to 10 minutes and involve dozens of individual actions, and these are tasks that the robot has never seen before, but the robot was able to carry out the task after watching just one demonstration. And this is actually a huge deal. It's really hard to get an AI model to actually learn new things after training because the parameters of its model are fixed. But here we're kind of seeing this emergent ability for this robotics model to generalize and actually apply new things it has never seen before.
And this is really useful because eventually, when robots get more and more common, especially in households, then you'd want to just demonstrate to the robot once how you want a specific task to be done, and it should be able to learn and apply that task in the future. So, a very fascinating read. If you're interested in learning more, I will link to this main page in the description below. And that sums up all the highlights in AI this week.
Let me know in the comments what you think of all of this. Which piece of news was your favorite, and which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week, I can't possibly cover everything on my YouTube channel.
So, to really stay up-to-date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching, and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.