Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
6,477
Runtime
35:03
Speaking pace
185wpm
Reading time
27min
185 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
AI never sleeps and this week has been absolutely insane. First, Deepseek releases their latest model with vision and then Alibaba also releases their latest and best Quen model. Shortly after that, Anthropic releases Claude Fable 5.1, which is at the time the best model in the world. And then shortly afterwards, Google drops their best and latest model, Gemini 3.8 Flash, which is also incredibly good. And then literally a few hours after that, Meta drops their latest and best model, Muse Spark 1.3. But all of this is dwarfed by OpenAI
93 words, the words spoken in the first 30 seconds at 185 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 374 |
| Average words per sentence | 17.3 |
| Longest sentence | 56 words |
| Questions asked | 3 |
| Sentences containing a number | 114 |
Most used terms
Filler phrases
80 in total: like 42 · basically 16 · actually 14 · kind of 3 · literally 2 · you know 2 · I mean 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
AI never sleeps and this week has been absolutely insane. First, Deepseek releases their latest model with vision and then Alibaba also releases their latest and best Quen model. Shortly after that, Anthropic releases Claude Fable 5.1, which is at the time the best model in the world. And then shortly afterwards, Google drops their best and latest model, Gemini 3.8 Flash, which is also incredibly good. And then literally a few hours after that, Meta drops their latest and best model, Muse Spark 1.3.
But all of this is dwarfed by OpenAI releasing their long-awaited GPT6 Astra, which just destroys all the other models. We also have several new open-source image generators and editors and some new ways to make the best open video model miniacs run in real time. Google DeepMind releases some really useful and practical AI models. We also have a really impressive pixel perfect world model and a lot more. So, let's jump right in.
First up, we have a really interesting project called H3 World. This basically takes the best open-source video generator, Miniax H3, and turns it into an interactive video game engine. So, here are some examples of this in action where you can basically give it a prompt plus some key presses, and it can generate a video that follows those key presses. It's as if you're interacting or moving around this world. And the clever part is that it doesn't actually build an entirely new control system.
It just translates the key presses into short English descriptions and feeds those through Miniaax's existing language understanding layer. It then connects each command to the exact part of the video where the action should happen. The awesome thing is they've released this already. So, if you click on this GitHub link at the top and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer.
They also released the training script to this as well, which is fantastic. If you're interested in reading further, I'll link to this main page in the description below. And you know, the crazy thing is progress is moving so fast. We have yet another very similar tool which also does the same thing. So, this is called Solar WM. And this also converts video models into real-time interactive worlds. And the nice thing about this one is it can keep generating worlds consistently for over an hour long, whereas some other world models start to fall apart after a few minutes.
And the really awesome thing about this is that this solar WM framework works with different video models, including 5B, 114B, LTX 2.5, and of course, the current best video model, Miniax H3. So this approach can be potentially applied across very different video generation systems. It's also worth mentioning the data set they used to train this. So the team assembled 1.3 million video clips totaling 25 terb of data which is huge and they're releasing everything for free which is fantastic.
You can click on these buttons to actually explore and download the entire data set to train world models yourself. Plus, they've also released the solar WM models for all of these video generators. So, at the top here, you can click on this checkpoint, which takes you to this page, and here it contains all the models. The awesome thing is they've completely opensourced this project. This includes everything from the training data set, the data processing pipeline, the training recipes, the models, etc.
So, if you're looking to train a world model yourself, this project can give you a ton of insights. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Google research releases an open-source time series foundation model. It's called Times FM3, and this is a zeros foundation model for predicting what happens next in numerical time series data. So, things like retail, healthcare, weather, and of course, stock charts.
The big upgrade here is that this is multivariate which means it can look at many related signals together to make its prediction. This is also super tiny at only 330 million parameters. But this is pre-trained on a massive data set of more than a trillion time points. So there's a ton of knowledge packed in. But what I think is the most powerful feature of this is this is zero shot which means you can give it a new forecasting problem it has never seen before and you don't need to like retrain the model and it's still able to help you predict future data.
So Google evaluated this new times FM3 model on various benchmarks like gifty valvbench and time and across the board it ranked first among other time series foundation models. The awesome thing is this is open source. So, if you click on this GitHub link and you scroll down a bit, here it contains all the instructions on how to download and run this locally on your computer. You can access the model on hugging face.
And at 330 million parameters, this is fairly tiny at only 1.32 GB in size. So, this should be able to fit on most consumer devices. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Bite Dance releases a pretty interesting AI called Lucida. And this basically takes a messy real room and turns it into a 3D simulation which you can edit further. And here's how it works.
Instead of creating one giant fused 3D scan, Lucida just looks at images of the room, identifies the individual objects and then builds complete 3D assets for each one and then puts those objects back into the correct positions. So this is a complete 3D model of the scene. And the advantage of this is that each object is its own separate mesh. You can move or resize or edit each object further. Now, this system works in three steps.
It first takes in these images of the room and then it parses the scene for different objects and then it proceeds to generate complete versions of these objects even if they were partially hidden in the original images. It then uses this gizmo act placement to basically predict their original locations and place them accurately back into the scene. Now, currently they've only released a technical paper on this, but if you're interested in reading further, I'll link to this main page in the description below.
Also, this week we have yet another open- source way to speed up Miniax H3 the best open video model in the world and potentially make it run faster than real time. So, this is called video deltaet and this is attempt to make minax way faster without destroying its video quality. You see, the problem here is attention. In the full Miniax H3, attention accounts for like 85% of the model's runtime. But what Video Deltaet does is it replaces this with a hybrid system.
In simple terms, nearby video frames still use the expensive but accurate attention data because that's where the fine details and motion matter. But information from far away frames is handled using much cheaper linear attention. Think of it as like spending expensive compute on what's close and using compressed memory for everything that's far away. And then combined with other optimization methods like optimized kernels, parallelism, and eight-step distillation, you actually get a super fast model with very minimal quality loss.
As you can see from these comparisons over here, so on the left is the full miniax H3 model. In the center is another open real-time method called fast H3. But as you can see here, it changes the quality quite a bit. But if you note this new VDNH3 on the right, it's incredibly similar to the full Miniax model. In fact, here it says if you happen to have eight Nvidia B200's, then this can produce a 14-second clip in only 11 seconds, making it faster than real time.
Now, I'm sure most of you don't have eight B200's lying in your basement, but even for consumer GPUs, this still offers a decent speed up with minimal quality loss. Now, at the top of the page, they've already released the model and the code to this. So, if you click on this link and you scroll down a bit here, it contains all the instructions on how to run this locally yourself. And the nice thing is there's Comfy UI support for this already.
So this user Saga Naki22 has built some custom nodes which allow you to run VDN H3 directly in your Miniaxcomy workflow. Now on a consumer GPU, this ain't going to run in real time, but it's around the same speed up as like a Turbo Laura, but the quality is a lot better. Anyway, if you're interested in reading further, I'll link to this main page as well as this Comfy UI custom node in the description below. Also, this week we have a new open-source image generator and editor.
It's called Lada image and here are some examples of its textto image generation. As you can see, this can produce some incredibly realistic photos. In addition to just generating realistic photos, this can also do posters and infographics with a ton of text and different elements. So, here are some examples for your reference. Now, we already have some really capable open textto image models like creatu and audiogram.
But the nice thing about this one is that it can also edit images. So here are some examples where we can take an input image and for example change the expression or pose of the character or change or replace certain elements in the image. Here's an example where we change the text or we can also colorize photos like this. So a super flexible tool. Now this is a fairly tiny 6 billion parameter model which is similar to Z image.
Now they've released two different versions of this. There's the base model which is mostly meant for training or generating higher quality images. This requires 50 sampling steps, but there's also a turbo model which generates images way faster. It only requires four steps. And then for each one, they've also released a BF-16 variant and an FP8 variant. For example, for the Turbo FP8 variant, you can see that the Transformer model is only like 6.7 GB in size.
So, this should be able to fit on most consumer GPUs. And then on this page, it contains all the instructions on how to download and run this locally on your computer. If you're interested in reading further, I'll link to this main page in the description below. Now, earlier this week, Deepseek releases their latest model with Vision, which is called Deepseek V4 Flash Vision Experimental. In most cases, it even beats the previous nonvision V4 Flash, and it's kind of on par with Cloud Opus 4.8, even though this is just a Flash model.
Now, they released this via the API a week ago, but this week, as with the previous DeepSeek models, they have open sourced this. So you can now download this and run it locally for free on HuggingFace. Note that this is 305 billion parameters and the total size is roughly 168 GB. And then on this page, if you scroll down a bit here, it contains all the instructions on how to download and run this. If you're interested in reading further, I'll link to this main page in the description below.
Then also this week, Alibaba upgraded their latest and best model when 3.8 Max. This latest version is called max0902 which stands for September 2nd. It's the same architecture as before. So with 2.4 trillion parameters and a million token context window, but here they basically post-trained it further on coding and co-work. Keep in mind they released Quen 3.8 Max just a few weeks ago, but this latest version significantly beats the old version across all these benchmarks.
So here are some agentic coding benchmarks for your comparison. This isn't just a marginal improvement, but for some of these it's like over 10 percentage points. Here are some other agentic and knowledge work benchmarks. Across the board, this is incredibly performant. For some of these benchmarks, it even beats Opus 5 as well as GPT 5.6 Soul, which is really impressive. Now, this latest version is not open sourced yet.
Currently, you can try it via their API in Quinn Cloud. If you're interested in reading further, I'll link to this main page in the description below. Now earlier this week, Anthropic releases their best and latest model Fable 5.1 and they claim that this model is especially good in terms of agentic scientific research as well as agentic coding, knowledge work, etc. In fact, according to this table, it not only significantly beats the previous Fable 5, but also OpenAI's best model at that time, GPT 5.6 sold.
And if you look at this independent leaderboard, then you can see that Claude Fable 5.1 Max does score the highest even above Opus 5 and GPT 5.6. However, it is by far the most expensive. So, here is the cost per task. You can see it's like more than 3.7 times more expensive than GPT 5.6 Max. There's also a ton of red flags with this. For example, even on the Max plan, after just one or two simple prompts, it already used up my 5h hour limit.
So, this is pretty much unusable. Also, even though they claim that it's really good for scientific research, I tried giving it a ton of different deep research and medical prompts, but it refused to use Fable 5.1. It just reverted to a dumber Opus 5, so it's pretty much unusable for scientific research. Anyway, I already posted a full review video on Claude Fable 5.1, including a ton of diverse tests, so I'm not going to repeat too much here.
I'll link to this video in the description below if you're interested in learning more. If you're doing any type of content creation, definitely check out Higsfield, the sponsor of this video. Higfield is an all-in-one AI creation platform built specifically for creators. Instead of jumping between a bunch of different tools, Higsfield gives you access to some of the world's leading models in one place, including the best video model, Cedance 2.5, in 1080p, as well as the best image model, GBT image 2.
The nice thing is Higsfield also has its own tools built for creative workflows. For example, there's Higsfield Supercomputer, which is basically a general purpose AI agent for content creation. You give it a prompt and it can help with the entire process. Finding an idea, creating the product concept, writing the brand direction, generating the visuals, and making the launch video. There's also Higsfield MCP, so you can connect the content generation abilities of Higsfield with an AI agent like GBT or Claude.
They do the thinking and planning while Higsfield executes the actual generation. They also have marketing studio which is really helpful if you're making marketing content. You can paste a product link or upload a product image and it can generate multiple ad formats in one workflow like UGC videos, tutorials, unboxings, product reviews, and more. So, one product can instantly turn into a whole campaign. They also have Cinema Studio which is built as a full end toend AI filmmaking pipeline.
This is especially useful if you want more cinematic control. Instead of typing a prompt and hoping the video looks good, Cinematic Studio lets you plan the scenes, control the camera, add specific characters, reuse locations, and keep everything consistent across the entire project. From idea to final output, Higsfield gives you way more control over the creative process, making it one of the easiest platforms to start creating with AI.
Try Higsfield today using the link in the description below. Also, this week, Google releases their latest and best model, Gemini 3.8 Flash and 3.8 on a flash cyber. Now, this ain't a pro model, but it actually punches well above its weight. So, you can see especially for finance and legal and agentic terminal coding, it even beats claude opus 5 as well as GPT 5.6 Soul, which is pretty impressive for a flash model. Now, the Gemini models are known to be the best in terms of multimodal capabilities.
So, they're really good at understanding images, video, and audio. And that's why you can see here this new 3.8 8 flash is also state-of-the-art in terms of understanding scientific figures and charts and also long video understanding for this humanity's last exam. This test and AI models knowledge on some really obscure domains and it seems like this new 3.8 flash also has a ton of knowledge packed in. It's also great at bioinformatics and biology research tasks.
What's really impressive is that if you look at this Deep Suite benchmark, which looks at how good a model is at Long Horizon Agentic Software Engineering, surprisingly, Gemini 3.8 Flash is now ranked number one. Isn't that crazy? A flash model actually is now ranked number one. And look at the average cost of 3.8. It's way lower than the other Frontier models. So, if you look at the Deep SW score versus the average cost, then it does look like Gemini 3.8 is the most efficient.
However, there's one caveat to this, which is if you look at the output tokens, then Gemini 3.8 Flash does use by far the most tokens compared to the other models. If you look at this leaderboard by artificial analysis, then you can see Gemini 3.8 Flash is ranked all the way here. So, one point below the open- source GLM 5.3 and Kim 3 and a few points below the Frontier Claw and GPT models. However, its speed is insane.
So, these Gemini Flash models are known to be the fastest models out there. This can reach like 348 tokens per second, which is far higher than the other competitors. In terms of cost per task, this is quite reasonably priced. If you look at this leaderboard by arena, you can see that Gemini 3.8 Flash is currently ranked number eight. Now, interestingly, if you look at this leaderboard called LiveBench, then Gemini 3.8 Flash is not even in like the top 10.
It's not even above Gemini 3.7 Flash. So they ranked it all the way down here, which is pretty crazy. So at least from this leaderboard, it seems that 3.8 flash might be a bit benchmaxed. That being said, if you are looking for a super fast model to use, especially for agentic coding, multimodal stuff, or finance, legal, or biology stuff, then Gemini 3.8 is a great option to use. So here you can see it being able to vibe code this 3D wizard game.
Or here you can see it being able to vibe code this very accurate cross-section and topographic map based on real data sets. Or here we can get it to vibe code this 3D mechanical keyboard key which can explode into separate pieces. Now they also released a cyber variant of this which is specialized for defensive security. So you can see here in terms of vulnerability discovery this even surpasses GPT 5.6 soul as well as cloud mythos.
Here's another benchmark showing how good it is at discovering vulnerabilities. Here they say that this cyber model produced 2.6 times more correct patches to vulnerabilities in Chrome compared to the best commercial models that are much larger. Now, currently both models are already out via the API as well as their coding harness called anti-gravity as well as AI Studio and Android Studio. For consumers, 3.8 8 Flash is available to paid subscribers through the Gemini app as well as AI mode in Google search and Gemini in Google Sheets.
If you're interested in reading further, I'll link to this main page in the description below. Now, literally a few hours after Google releases Gemini 3.8 Flash, Meta also drops their latest and best model, Muspark 1.3. I hope you can feel the acceleration here across various knowledge work and agent coding benchmarks. You can see that Muspark 1.3 is also near Frontier. And in some cases like Deep Suite and SUI Atlas, it even beats GPT 5.6 Soul.
The big focus here is making agents better at doing long messy jobs rather than just answering one question. So you can give it an open-ended goal with multiple files, tools, and steps. And it can work autonomously and gather information, plan out everything, correct itself, and eventually produce a final finished deliverable. As you can see, it's designed to handle workflows across a variety of different platforms. It's also designed to do much better at multitasking.
In other words, it can handle several tasks inside the same conversation without getting confused about which instruction belongs to which job. As you can see from this example here, it says Musepark 1.3 is available today in Muse Code, which is their agentic coding platform, as well as the meta model API. If you're interested in reading further, I'll link to this main page in the description below. But here's the thing, you can just forget about all the models I mentioned so far because shortly after OpenAI releases their best and latest model, GPT6 Astra, and this is an absolute beast.
This model is designed to not only answer questions, but to handle much longer and more complicated real world tasks. In fact, it can use your computer and cursor directly. Here are some examples of some ridiculous things it can do. Here, it's able to turn a circuit schematic into a manufacturable printable circuit board. Or here's an even crazier example where it can automatically use Excel to do a ton of stuff. For example, it made all of this in Excel.
And of course, it's really good at vibe coding games and other 3D environments. So, for example, here it's using Unity to build a city scene completely from scratch. Here's another example where we can get it to create a 3D model of a five-speed car transmission in Frecad and then use Blender to animate the gears in motion. And here's an example where you can easily get it to fill out tax forms or any other forms directly on your browser.
Here are some ridiculous benchmarks for your reference. So for agents last exam, you can see this is the best model in the world. And the largest improvement is computer use. So, you can see this benchmark OS World, which measures an AI's ability to operate real computer interfaces. This scores much higher compared to the other Frontier models. Here's another benchmark called Benchcad, which tests how good an AI model is at reconstructing 3D objects from different views.
And as you can see, again, GPT6 Astra is state-of-the-art. Another cool strength of GPT6 is it's able to transcribe music very well. So, here's a benchmark on how good it is at turning an audio clip into sheet music. And as you can see, GPT6 Astra performs incredibly well. It's also incredible at design, even out competing the latest Fable 5.1 as well as Opus 5. Here's another example where GPT6 can model this entire house in Blender and turn it into a walkable scene in Unreal Engine.
And of course, this is also incredibly good at agentic coding. As you can see from this new terminal bench 4, which tests an agent on complex terminal-based software engineering tasks, you can see that GPT6 Astra is again state-of-the-art. Here are some science benchmarks. So, GPQA Diamond is like its knowledge on graduate level science questions. Again, this is Frontier. Same with Healthbench, Lifai Bench, you get the point.
GPT6 Astra is just really freaking good. And in terms of cyber security, this is also incredibly good. So, for exploit bench, get this. it actually got 100% completely destroying this benchmark for exploit gym. It's also much more performant and way more efficient than the previous GPT. Now, interestingly, if you look at this artificial analysis intelligence index, then GPT6 Astra is actually second place, still behind Claude Fable 5, which is kind of interesting.
Something doesn't quite add up here because at least in my experience, GPD6 does perform a bit better than Fable 5.1 and Muse Spark should not be ranked number four. I mean it feels very benchmaxed. So maybe they need to recalibrate this intelligence index. And then for cost per task, you can see that GPT6 is quite reasonably priced much cheaper than the clawed models even though it's like similar intelligence. If you look at live bench, which is another independent leaderboard, you can see that interestingly GPT6 Astra is ranked number three even behind Fable 5.
One of the craziest achievements of Astra is its score on this Arc AGI 3 benchmark. This basically drops an AI agent in a new game environment it has never seen before. It has to figure out the rules and basically beat the game or proceed to the next level. For humans, this is quite easy, but for an AI, this is actually extremely hard because technically they can't learn after training. Their model weights are fixed.
So, it's really hard for an AI to actually take in new patterns or information and apply them. That's why you can see like for most of the Frontier models, they score under 10%. But for GPT6 Astra, if you crank it up to like the maximum thinking effort, it scores above 60%. Which is a crazy lead. And if you add an adapter or harness to this, then it can even achieve close to 100%. If you look at this Frontier Math airish benchmark, it's basically seeing if AI can solve some extremely hard and unsolved math problems, then you can see that GPT6 Astra is the only model that actually scores above 0%.
All the other Frontier models get zero. In fact, they also got GPT6 to play Pokémon Fire Red all autonomously. And at that time, you can actually watch it via a Twitch live stream. And as expected, Astra took the least amount of time to complete the game, much shorter than the previous models. Now, interestingly, Astra High actually was the fastest, even faster than Astra Max. And at least at the time of this recording, which is Saturday, GT6 Astra should be rolled out to all paid plans, including Plus and Pro plans.
If you're on the free plan, you will not have access to this yet. Now, today is reserved for my weekly news video, so I got to publish that first, but I am working on a full review of GPT6 Astra with some insane demos. So, I'll probably release it tomorrow or the day after that. For now, if you're interested in learning more, I'll link to this main release page in the description below. Also, this week, Google DeepMind continues to release some really useful stuff.
So, they just released weather next 3 and this is their newest AI system for predicting the weather. The big update here is that it can make much more detailed forecasts using live satellite data. Basically, instead of relying mainly on weather simulations that can already be several hours old, weather can look directly at recent observations of the atmosphere and produce a fresh global forecast every hour. And the detail is much higher.
You can see the previous version, weather next 2 predicts weather on a 25 km grid every 6 hours, whereas weather next 3 can predict things like surface temperature and moisture at resolutions as fine as only 5 km. So it's roughly five times sharper in resolution. This makes a big different for things like mountains, coastlines, storms, and local rain where conditions can change dramatically over a short distance. You can see for this satellite precipitation benchmark, this new weather next 3 has improvements of up to 60% compared to the previous version.
Now, this new model can also predict things specifically useful for renewable energy. Things like wind speeds at roughly the height of wind turbines or the amount of sunlight reaching solar farms. And this isn't just a research demo. Weather Next 3 is already being integrated into Google Search, the Gemini app, Google Maps, as well as their API and Google Earth Engine starting this week. So, you're going to start to see like up to 50% more accurate precipitation forecasts as well as other improvements.
If you're interested in reading further, I'll link to this main page in the description below. Now, in addition to Weather Next 3, here's another really cool project by Google Research. So apparently they now mapped out the complete brain and central nervous system of a male fruitfly which looks like this. This is what they say is the largest brain wiring map ever made. You can think of this almost like creating a circuit diagram for the entire nervous system of the fruitfly.
So this map contains more than 166,000 neurons connected through roughly 125 million synapses. And it doesn't stop at the brain. So, it also includes the vententral nerve cord, which is kind of similar to our spinal cord, meaning the scientists can begin tracing how signals come through things like vision, smell, and hearing from different parts of the fly's body, and then travel through the brain, which would then control various movements in the fly's body.
So, here you can see where this brain and central nervous system are located in the male fruitly. So, it's only like this red part over here. Here's an example of just one neuron. Now, creating this map required cutting the nervous system into enormous numbers of incredibly thin slices and then photographing those slices with electron microscopes and then using computing and AI to reconstruct those flat images into three-dimensional neurons and connections.
So, this takes a ton of time and compute and effort. Now, this fruitfly has 166,000 neurons, and this is the largest reconstruction of this nervous system we have so far. For your comparison, a human has 86 billion neurons. So, at least with our current technology, we can't really map all 86 billion neurons from a human brain at this level yet. But this is just the start. This is just the male fruitfly. The researchers have also mapped the female fruitfly.
And next, they're also looking at mapping the brains of this laral zebra fish. This is essentially one of the most complete biological wiring diagrams ever created. So, it gives neuroscientists a much more complete understanding for, you know, how brain circuits work and how they influence behavior. This might also give us insights on how to actually design AI systems and robots. If you're interested in reading further, I'll link to this main page in the description below.
Also, this week, World Labs, which is the company started by Fei Lee, who's often considered the godmother of AI, they've just released a really cool world model called Atlas. And here's how it works. You can give Atlas a ton of different inputs, including a text prompt, images, video, or 3D information, and it can use all of that together to create or reconstruct a 3D world. So, here are a few impressive examples of Atlas in action.
Here, you can input a video and then manually draw a camera path, and Atlas can generate up to a full minute of coherent 1440p video moving through the scene following that camera path with pixel perfect control. As you can see, you can like freeze the frame at any point and everything looks very coherent. This is different from other world models or 3D reconstruction models which tend to have a lot of errors, especially if you try to view the scene at different angles or especially in high action scenes like this.
But here, everything looks very coherent and faithful. Or here's another example where it can reconstruct a real place from only a handful of photos. So, this can be very useful for like real estate VR. And it doesn't have to just output a flat video. Atlas can also generate explicit 3D geometry like point clouds or gausian splats. And of course with such accurate and faithful renders then you can also use it to generate simulations or synthetic training data for robots.
Now currently this is just a preview but on this page you could request early access to Atlas by filling out a form. If you're interested in reading further I'll link to this main page in the description below. Also this week we have yet another open-source image generator and editor. In fact, this has incredibly good multimodal understanding. So this is called Intern Luminina U2. And first of all, here are some text to image generations.
You can see it can generate images in a variety of different art styles, including realistic photos and oil paintings. But it can also edit images, and this has really good prompt understanding. For example, we can get it to rotate only the orange squares, and it's able to follow this very well. or here's another example for your reference. Similar to the other image editors, this can basically edit or remove certain parts of the image.
But the really cool part is this is a diffusion large language model. So not only does it just generate images, but it also has incredible world understanding. You can plug in an image for it to analyze. For example, if you give it this medical scan, it's able to detect that it's a neuronal migration defect. Or if you give it this painting, it's able to identify the artist. Or you can even give it this incredibly complicated ER diagram and it's able to understand everything and answer your questions.
So think of this not only just as an image generator but as a pretty capable language model with vision. Now at the top if you click on this code button and you scroll down a bit, they have released the instructions on how to run this locally on your computer, but we're currently waiting for the model to be uploaded on HuggingFace. For now, if you're interested in learning more, I'll link to this main release page in the description below.
Also this week we have a new AI called Vigle Animate. This basically swaps the character in a clip using one edited frame. So you can just take a video and then for example for the first frame use an image editor to swap the character into something else and then plug it through this model and it can animate the entire video with this new character. The nice thing about this is unlike other animation transfer tools, this one does not have to do any like pose estimation or segmentation, face tracking, any of that complicated stuff.
All it needs is a frame of this new character and it'll automatically apply the original character's motion onto the new character. And this can generalize across a ton of different characters with different proportions or art styles. It doesn't have to be just a realistic human. Now, technically, this is actually a fine-tune of the best open video model out there, Miniax H3. They basically took their reference to video transformer model and trained it and distilled it further so that it can be specialized in this character replacement capability.
And on this page, they've released the models to this already. This consists of the full fine tune of the Miniax model, which is like 66 GB in size, plus they've released Allora, which is only 2.5 GB in size. And then at the bottom here, it contains all the instructions on how to download and run this. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Runway introduces GLM Worlds 2.
This can basically generate an interactive virtual world in real time where you can walk around and also do certain actions. So here are some examples of it in action. All you need to do is give it an initial image or a description of the world and it basically places you in that world. So this can be like first person or third person. It can be like any character or any environment. And then you can input commands like walk forward or like change the weather or even control certain actions of other characters or objects in the scene.
And this outputs continuously generated video at 720p resolution and 24 frames per second. Now, we already have a ton of world models out there that can do similar things, but the cool thing about this one is there's no fixed maximum duration. This can keep going on and on. Now, in technical terms, this is an auto reggressive diffusion model for both video and audio. So, in simple terms, it basically keeps looking at the world it has already generated to predict the next part continuously.
And that's why this can keep going on and on seamlessly. Now, unfortunately, this is just a research preview, and you'll need to fill out this form to get in touch about GLM Worlds 2. It's not available for use to the general public right now, but if you are interested in reading further, I'll link to this main page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this.
Which piece of news was your favorite? and which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay uptodate with all that's going on in AI, be sure to subscribe to my free weekly newsletter.
The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.