Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
6,948
Runtime
38:54
Speaking pace
179wpm
Reading time
29min
179 words per minute, just under the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
AI never sleeps and this week has been absolutely insane. Deepseek releases their latest model and it's an absolute beast. We also have a super powerful speech generator and editor. It's kind of like Nano Banana but for speech. Open AAI used a swarm of 10,000 agents to solve one of the hardest and unsolved math problems in the world. Google DeepMind also used AI to essentially map out every possible mutation in the human genome. We have a new state-of-the-art open-source model for predicting the depth and surface normals
90 words, the words spoken in the first 30 seconds at 179 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 463 |
| Average words per sentence | 15.0 |
| Longest sentence | 52 words |
| Questions asked | 14 |
| Sentences containing a number | 91 |
Most used terms
Filler phrases
96 in total: like 42 · actually 26 · basically 17 · kind of 4 · you know 3 · I mean 2 · right? 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
AI never sleeps and this week has been absolutely insane. Deepseek releases their latest model and it's an absolute beast. We also have a super powerful speech generator and editor. It's kind of like Nano Banana but for speech. Open AAI used a swarm of 10,000 agents to solve one of the hardest and unsolved math problems in the world. Google DeepMind also used AI to essentially map out every possible mutation in the human genome.
We have a new state-of-the-art open-source model for predicting the depth and surface normals of an image. So releases their latest model, but we also have a new top open-source music generator which is just as good. This AI can animate any character, including ones with unusual skeletons. We have a ton of open source AIs for reconstructing a scene in 3D. We also have a new real-time interactive world generator which is really good quality.
We have a ton of new open- source robot models and a lot more. So, let's jump right in. First up, we have this AI called Merryold V2. This basically understands 3D depth and structure inside a regular image. So what you would do is give it a regular image and it can convert that into a ton of things like a depth map, surface normals which is like the orientation of a surface in the image and then also albido which is like the raw color and other different maps.
And the awesome thing is this new method is pixel level. So this is super high resolution and much better compared to the other methods. For example, if you compare this to a previous model called Moji 3, you can see that Marold is a lot more detailed and accurate. Here's another comparison between another model, Infinidepth, and as you can see, Marold is just a lot more detailed and higher resolution. And then here's a comparison of its normal estimation capabilities.
You can see that this new Marold method is just a lot sharper and faithful. Here's another crazy comparison. And if you look at these benchmarks comparing other similar models, you can see that on average, Merryold V2 scores the best. The awesome thing is they've released this already. So at the top of the page, if you click on this code button and you scroll down a bit, here it contains all the instructions on how to download and run this locally on your computer.
Note that inference at this resolution requires around 17 GB of VRAM, whereas this resolution requires around 29 GB. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have a cute little AI called Unimate. This can basically animate completely different 3D skeleton types using just one AI model. So, here's an example where you can get it to animate a ton of different nonhuman characters like flowers or Garfield or a satellite or this dragon-like creature, and it's able to handle all these animations very well.
You see, a lot of these skeleton animation models are only specialized for like human characters, but they often struggle when you give them something unusual like a bird or snake or some other completely different objects. But here, Unimate is able to animate pretty much anything. You just need to give it a rigged 3D model plus a text instruction like a dog walks forward or the snake slithers forward and it can generate the motion automatically without any further retraining.
So, this is one of the best AI models to use to animate unusual characters or objects. Now, at the top of the page, they've released the code already. So, if you click on this link and you scroll down a bit here, it contains all the instructions on how to set this up locally. They also released the training code to this, so it's fully open source, which is fantastic. If you're interested in reading further, I'll link to this main page in the description below.
Also this week, Google DeepMind just released Alpha Genome Atlas and this is basically a giant AI generated map of human DNA. Now currently we only understand 2% of this which mostly codes for proteins but the other 98% is a lot more mysterious. So Google used Alpha Genome to basically predict the effects of every possible single letter mutation in the human genome. Now there are roughly 9 billion possible singlelet mutations and testing all of them experimentally would be pretty much impossible.
So instead of trying this out in a real lab, DeepMind just used Alpha Genome, which is an AI to predict the effect of every single one of these mutations. And the result is a one pabyte data set containing predictions for more than 9 billion genetic variants, making it over 30 times larger than the Alphafold database. So if you need to look up the effect of a mutation, instead of running the AI from scratch every time, scientists can now just look up this mutation in Alpha Genome Atlas and get predictions on how it could affect things like gene regulation or protein production.
They've also created something called the AVI score, which is basically an impact score to help researchers quickly identify which mutations are worth investigating further. And it's already showing some real results. So researchers used it to find 22% more genetic associations in the UK bioank data and it also helped support the solution of a previously unsolved rare disease. So basically this takes a huge amount of genetic information that was previously very difficult to interpret or predict and now they've just turned it into a huge database which researchers can actually search.
And the awesome thing is they've released this atlas for everyone to access. So simply click on explore alpha genome atlas at the bottom here and then afterwards you can search this database for like over 9 billion mutations. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a new realtime interactive world model called Lingbot World 2. This basically generates an interactive virtual world continuously and you can use key presses to control it and walk around.
So here are some examples of this in action. You can see that it's much better quality than the previous open- source world models. Here, everything looks a lot more detailed and coherent and higher resolution. Now, in addition to just controlling the movements via key presses, you can also enter prompts to add any event or effect you imagine. And the team says this can generate interactive worlds continuously for over an hour.
And the real-time system can reach 720p resolution at 60 frames per second, which is really impressive. They've also added an agent mechanism which means that you can now have NPCs inside the environment. In other words, you can add different characters which also move around and they can behave and respond in different ways. Now, under the hood, they've actually just used the pretty old Alibaba 1 2.2 as the video generator, but here they're generating the world chunk by chunk, so it can actually stream real time.
It also caches previous information, so it kind of retains a memory of the scene. Now, at the top here, if you click on this code button, and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. Note that there are two different variants of this. There's a larger 14b variant, which is higher quality, and a smaller 1.3 billion parameter variant, which is a lot faster.
If you're interested in trying this out, I will link to this main page in the description below. Also, this week, we have a new open-source robotics model called Isaac 0.5. This is a robot foundation model designed specifically to give robots a much more general understanding of what they're seeing and what they should do next. So you can feed it images, video, language instructions, and the robot's current state or even its previous actions.
And this model can basically take all of this input and predict the next action or how to answer questions. So it can do things like locate objects, predict what the world might look like next, or directly generate robot movements. Now this is a 36 billion parameter sparse model. So it's fairly tiny. And they trained this model on data from more than 35 robot systems. So this model can be applied or transferred to different robot types.
They also trained it on a 100,000 hours of robot experience and around a million hours of general video. One interesting feature is that Isaac doesn't just learn robot control separately. Instead, video understanding, spatial reasoning, predicting future states, and physical actions are all trained together in just one backbone. So, knowledge from all these inputs can potentially help the robot make better decisions and understand the physical world.
The awesome thing is they've actually open sourced the model. So, if you click on this button at the top and you scroll down a bit here, it contains all the instructions on how to download and run this. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a new AI called world sculpt and this turns a scene into individual 3D objects which you can edit further.
So how this works is you can input multiple images of a scene or even a video and it'll turn that into a complete 3D scene with separate individual objects. You can then edit each of these objects further like resizing them or moving them around. And this is an important distinction because you know most 3D reconstruction systems can create something that looks like the original scene but everything is just fused together.
But here world sculpt is able to reconstruct each object separately and also position them correctly inside one shared 3D world. That means you can like move each object around and resize them or edit them further. The cool thing is this can even convert gausian splat worlds into actual 3D mesh scenes. So here's an example of that. The nice thing is they've released this already. So at the top of the page, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this.
If you're interested in reading further, I'll link to this main page in the description below. Also, this week we have a new AI called Fire 3D. And this can take a series of normal photos or videos and turn it into a complete 3D scene which you can edit further. So you just input one or several images or a video and it can basically reconstruct the 3D scene with separate objects. The entire scene is simulation ready and you can edit each object further.
Every object basically becomes its own complete mesh with material information. So you can move them around or add each one to a robotic simulation. And this thing is super fast. It can recreate a simulation ready scene in under a minute and processes up to 16 objects in parallel on just one GPU. Now, at the top, they released the code to this already. So, if you click on this link and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer.
Plus, they've also released the training code to this as well, which is fantastic. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Tencent releases a really powerful audio model called out. And you can think of this as like Nano Banana, but for speech. First of all, this can do regular text to speech. And you can describe exactly the voice you want and the dialogue.
For example, here's the prompt and here's the result. >> He always said he'd come back. He always kept his word until now. >> As you can hear, this voice does sound pretty sad and full of grief. And this can also do zeroot voice cloning. So, you just need to upload a few seconds of someone's voice and then plug it through this AI to get that voice to speak out something completely different. So, here's an example. First, I'll play you the input reference voice and then the output generation. >> I survived being swallowed by the dark heart of the universe.
So, you lost seem a wee bit quaint to me. If not, even light can escape the event horizon. How do you think you'll escape me? >> But this can do much more than that. You can microedit existing speech by like adding or deleting words. For example, let me play you the original clip first. What can I say? Mamba out. >> All right, so that was the original clip where he just says mamba out. Well, we can add the word never here.
And here's what it sounds like. >> What can I say? Mamba never out. >> As you can hear, it can seamlessly add a word into the audio. Or here's another example where we can just completely delete this sentence. And here's the result. >> Yo, thank you guys for all your sacrifice. And Vanessa, you holding down the family the way that you have. As you can hear, it can seamlessly get rid of this sentence. Or here's another example.
Let's say we want to add this phrase. Let me play you the original first without this addition. >> Equipped to answer that question. [music] When I was in my 20, change is going to be a constant. That's such a great qu. >> And then here's the edited version where we add this phrase >> question. [music] When I was in my 20, change is going to be a constant. And the most important lesson is to accept changes. That's such a great >> As you can hear, this is very seamless.
And this can do even more. So you can even plug in a really noisy and messy clip and get it to remove the noise or enhance the voice. So here's the input and the output. >> She had your duck suit and greasy wash all year. She had your duck suit and greasy wash water all year. >> This can also increase the resolution and quality of an audio clip. So again, I'm going to play you the input and then the output. Another crazy thing you can do is change the emotion of an existing clip.
So, for example, let's turn this clip into a sad tone. I'm going to play you the original first and then the output. >> Every day we were asked to a social occasion and this was very very pleasant for us. >> Every day we were asked to the social occasion and this was very very pleasant for us. Or instead of sad, let's change this to angry and hear what it sounds like. >> Every day we were asked to the social occasion.
And this was a very very pleasant for us. >> The crazy thing is you can even change the tamber to something else. So for example, the original audio is a male voice. We can change this into a woman with a clear and bright voice. Let me play you the original and then the output. Lo was convicted of harassing their children at Flender Street Station and fined $700. >> Lo was convicted of harassing their children at Flender Street Station and fined $700. >> Or you can also take an existing audio clip and turn it into a whisper. >> The Eastern Coast is a place for pure pleasure and excitement.
The Eastern Coast is a place for pure pleasure and excitement. There's also a ton of other stuff you can do like removing size or adding or removing laughter, adding or removing breaths, etc. So, this is an incredibly flexible speech editing tool. The awesome thing is they released this already. So, if you click on this GitHub repo and you scroll down a bit here, it contains all the instructions on how to run this locally on your computer.
And the model is actually surprisingly tiny at only like 6.12 GB in size. So, this should be able to fit on most consumer GPUs. Definitely one of the most flexible and powerful speech editing tools I've seen so far. If you're interested in reading further, I'll link to this main page in the description below. If you're doing any type of content creation, definitely check out Higsfield, the sponsor of this video. Think of this as an entire AI creative team right at your fingertips.
In fact, you can connect the best AI model, GPT6 Astra, to Higsfield, right inside ChatGpt to give it the ability to create images, videos, and other content. Astra handles the planning and writes the prompts while Higsfield generates the visuals. You can run the whole workflow right inside your conversation. Higsfield also has a new AI motion designer that connects Astra directly to After Effects. So you can describe an animation and have it build an actual editable project complete with layers, shapes, and expressions that control the movement.
And afterwards, you can then open the project in After Effects and edit the timing, colors, and individual elements yourself, or just prompt the agent to edit it for you. You can also give it an animation reference to recreate. Automate repetitive work across hundreds of shapes, or apply your brand's fonts and colors through a composition. And if there's an effect you use regularly, you can ask Astra to turn that workflow into a reusable plug-in.
Whether you want to create product ads or cinematic videos or custom motion graphics, Higsfield gives GPT6 Astra the creative tools to turn your ideas into actual finished content. Try it today using the link in the description below. Also this week, DeepSeek releases their latest model, Deepseek 4.1 Flash. Even though this is just a.1 upgrade, it's actually completely different from the previous V4. First of all, here are some benchmarks for your reference.
Even though this is a flash model, get this, it even scores 74.2 on Deep Suite 1.1, which actually puts it as the number one model on this leaderboard, even beating GPT6 Astra, which is pretty crazy. Now, the Deep Suite team hasn't added this officially to the leaderboard yet. So, let's wait and see where they place this. Anyway, in terms of CyberJ and Automation Bench, this is also state-of-the-art. In fact, if you look at Automation Bench, this flash model even beats GPT6 Astra Max.
Now, here are some specs for your reference. So, this is a 552 billion parameter mixture of experts model. Think of this as a team of experts working together to help you solve a problem. So, when you use it, only around 8 to 16 billion parameters are active, making this super efficient. And you know, the strangest thing about this is that they used a completely new architecture for this new version 4.1. There's so many new things that we've never seen before, such as a separate encoder and decoder component.
For your reference, most modern language models only have a decoder component. They also added this new engram feature, plus a sliding window attention, plus this new CSA 2 attention. Anyways, it's quite complicated and technical, but I'll try to make a full explainer video on this next week. Basically, this is a significantly different design from the other mainstream Frontier models. If you look at LiveBench by Abacus AI, then you can see that this new Deepseek is indeed the number one open-source model just a few points below GPT6 Astra and Claude Fable.
If you look at this Val's index, which measures an AI model's performance across various knowledge work tasks, then as you can see, this new DeepS is indeed the number one open model, just slightly above Kimik 3 and GLM 5.3. But notice that this isn't really a significant difference. So they're kind of all tied for first place. If you look at artificial analysis, then according to their latest intelligence index, this new Deepseek is still behind Kimik 3 and GLM 5.3.
However, at least if you use it through their API, then this thing is blazing fast at 217 output tokens per second, which is like over three times faster than GLM and like seven times faster than Kim K3. The cost of this is also absurd. So way cheaper than the other leading open models and of course even cheaper than the closed models like GPT6 and Fable. Now as with the previous DeepSeek models, they've already open source this.
So if you click into this folder, note that it's around 510 GB in size. So this is still a massive model. However, because this is open source, the community has already released some customized or quantized versions of this. For example, this user has released GGUFS. And the smallest Q1 is only 106 GB in size. Now, if you don't have the hardware to run this locally, of course, you can also use it via their API, which is insanely cheap.
Definitely the most costefficient Frontier model out there. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, OpenAI just solved one of the hardest math problems in the world called Navier Stokes. And this is actually a huge deal. Let's go over this in simple terms. First of all, the Navier Stokes equations basically describe how water and liquids and gases move.
So things from water swirling in a cup or how air moves through a room. And you know, this equation is pretty good. It's able to describe how liquids and gases move in everyday life. But here is a famous unanswered question about this. Is there any instance where given the right conditions, this equation would break down? And this is one of the Millennium Prize problems which are like the deepest hardest unsolved problems in mathematics.
In fact, this problem has remained unsolved for roughly 90 years. But this week, OpenAI used its internal model which they claim is significantly more capable than GPT6 Astra to solve this problem. Here's the solution. So it turns out that in a whirlpool with some very specific conditions, it would generate this vortex which continues stretching and stretching into infinity. In other words, the speed of this would keep growing without limit in a finite time.
This is what mathematicians would call a singularity or blowup where it just keeps accelerating. And this is proof that indeed under certain conditions this Navier Stokes equation could be false. And the crazy thing is, OpenAI actually used an army of 10,000 concurrent agents that worked for 88 hours. Again, this is an internal model that's supposed to be way better than GPT6 Astra. So, you can see that the performance of this new internal model in white is significantly better at solving open math problems.
This is a huge deal because this problem has never been solved before by any human for around 90 years. But with just 88 hours, these AI agents were able to solve it. I mean, we are going to get some significant technological acceleration in our lifetime. This is a genuine breakthrough and you can expect even more of these in the near future. Now, there are a few caveats to this discovery. First of all, there's actually different like levels of solving this problem.
So, first of all, oiler is like the easiest level to solve. This basically removes viscosity from the equation. Now, viscosity is like the thickness of a liquid, right? So honey is more viscous than water. Now if you remove viscosity from the equation then it becomes much easier to solve. Another term you need to understand here is force. So force just means you are applying an external force to the situation to guide everything.
For example, stirring the water. And of course this also makes it easier to achieve the ideal conditions to break the Navier Stokes equation. So it turns out that these mathematicians, one from Anthropic, also coincidentally solved at least the easiest forced oiler problem around the same time. However, the solution from OpenAI was quite different. So they solved a forced non- oiler problem. In other words, their proposed solution includes viscosity in the equation, which makes it even harder.
However, it's important to note that we still haven't found a solution to the unforced Navier Stokes problem. In other words, can there be any condition where there's no external force? Like there's no spoon stirring the water. No one is guiding the liquid into this vortex. Can water or gas still naturally achieve singularity under any condition? Nobody has solved that yet. So those are some noteworthy caveats to this announcement, but still this is a pretty big deal.
I mean AI has solved something that no human was able to solve in over 90 years. This is one of the hardest math problems ever. So, this is a genuine breakthrough in mathematics. Anyway, if you're interested in reading further, I'll link to this main page in the description below. Also, this week we have a very interesting project called Show Harness. And this asks a really simple question. Can we just use vision models to control robots?
In other words, AI models that can understand and analyze images. What if we apply them to robots? So what they did here is they took an existing vision language model and they gave it a list of movements like move left, rotate, move forward or manipulate something. And the vision language model basically has to look at the scene, reason about what it should do and choose one of these actions. Now again this is just a vision model.
This is not a language action model which is what controls robots. So they also need to plug it through this robot specific action interpreter to convert these instructions into actual movements that control the robot. But it turns out that this actually works. Like if you give the vision model a list of meaningful action names, then it's actually able to output a list of actions to control the robot and success could reach up to 100%.
So this actually works. And so what they did next is they actually created a harness where you can plug in any vision language model including like Gemini, GPT, and of course open- source options like Quinn or GLM and then just use that to control a robot. So at the top here, they've released a GitHub repo to this. And if you scroll down a bit here, it contains all the instructions on how to set this up. Now, of course, you do need to have a real robot to make this work.
If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a very clever framework called Edge Zero. This is a pretty neat way to run large language models on machines that don't have enough RAM or memory. Now this is designed for mixture of experts models. And this is based on one of the best medium-sized models out there, Quen 3.535B. Now this is a mixture of experts models.
So think of this as like a team of AI experts helping you solve a problem. When you use it, it only routes to the relevant expert. For example, if you're asking it a math problem, only the math expert would be active. In fact, when you use it, only 3 billion parameters out of the total 35 billion parameters are active, making this very efficient. Now, usually you still need to have enough memory to fit the whole 35 billion parameter model.
But what edge zero does is instead of loading the whole model into memory, it only streams in the experts the model actually needs for your current question. That means with this edge zero framework, the memory usage actually depends much more on this active portion of the model rather than the full parameter count. They also compressed the Quen model down to int 4. So this is a lot more compressed and smaller, but there could be some quality loss.
But what they've done is they've introduced this recover Laura which actually helps regain much of this quality loss. So from this framework, here's the crazy thing. You can now actually take Gwen 35B and run it with as low as just 2.9 GB of memory. Or you can also take a smaller 8B variant and it can run on just 1 GB of memory. And this is based off of inclusion AI's ling 3.0 tiny. So a very clever way to run these mixture of experts model without loading everything at once into memory.
And then at the bottom here, it contains all the instructions on how to download and run this on your Mac device. If you're interested in reading further, I'll link to this main page in the description below. Also, this week we have a new and pretty interesting benchmark called Real Sweet. You see, most of the other main software engineering benchmarks like Swebench or Deep Suite are already getting saturated very quickly.
Well, this new one called Real Suite has a quite different design. So, this tests whether an AI can handle actual software engineering work and even Fable and GPT6 scored less than 40%. So, how this works is it gives agents access to private company code and asks them to make changes that matter to the business. So, things like fixing invoice taxes or migrating customer accounts. The tricky part is understanding how that particular company's software works, including rules scattered across different systems.
A common failure is simply missing a requirement. In other words, writing code is only part of the job, but the agent also has to figure out everything the change needs to accomplish. then connect it correctly to the existing product. And here's the current leaderboard. Fable 5.1 is number one, scoring only 38.8%. GPT6 Astra is 33.8% and then Gemini 3.8 Flash is number three. And then the best open model, GLM 5.3 is number four.
Actually, note that there's not actually a significant difference. So, these four are kind of statistically tied for number one. Anyway, a pretty interesting new benchmark. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Sunno introduces their best and latest music model, Sunno V6. This is designed to give you more control over both creating songs from scratch, but also microediting them afterwards.
Now, of course, you can just take a text prompt or an audio reference and turn that into new music, but the really useful part is making specific edits. You can prompt exactly how you want to edit the song. For example, changing a single lyric without affecting the rest of the song, or combining vocals from one song with an instrument from another song, or isolating a guitar rift and building a beat from that. In fact, here's their demo.
[music] >> [music] >> You'll [music] you get [music] you get down. [music] >> [music] >> Now, they've actually released three different versions of this. There's a V6 version, which is available for paid plans. This is designed for reliable, precise results. There's also a V6 Wild one which is also for paid plans. This is less predictable and more varied. And you would use this to make some more unexpected ideas. And then there's also a V6 mini which is available to everyone even on the free plan.
This is a lot faster, but it's lower quality and less precise than V6. So that's V6. Now, of course, Sunno is closed and paid. Plus, they've recently added a ton of restrictions to users, such as limits on downloads. So, instead, here's another really awesome open-source music generator that was released this week. It's called UA2. And here's the really interesting part. Before generating the final audio, it actually composes a musical plan first.
So, instead of going directly from a text prompt to the final audio, it first creates something much closer to a real musical score containing the melody, rhythm, chords, and structure. And you can actually edit the score further if you want. And then the model will turn this into a full song with vocals and accompaniment. So this means you can potentially change the melody, the lyrics, tempo, or arrangement at a structured musical level instead of just trying to prompt it further and hoping it gets it right.
Here are some examples for your reference. Hey John, can you hear it? that shuffle [music] through the door. Old shoes on a new floor. My feet itching for more. Hey John, can you hear it? Boogie woogie. [music] This rhythm [singing] turns me on. Let's go dancing soon. [music] Oh yeah. All Saints night. Two worlds [music and singing] blend. Beat the light. Trick or treat and candle [music] flame. Different paths yet both the same.
[music] >> We think of those who came before. We knock [singing and music] on every neighbor's door. Whispers of the past [music and singing] still call. Life and death can dance through all. Cities and grace, laughter [music] and tears, holding the lost through all [singing] these years. Remember the death of living [music] well. Fill your days with tales [singing] to tell. Candies [music] and grapes, sweet and true. >> That's what the devil want us [music] to do. >> Now, the awesome thing is this can also do covers.
So, for example, we can plug in Jingle Bells and get it to rerender this in a minor key. [music] >> Dashing through the snow in a one-horse open sleigh. Over the fields we go, laughing all the [singing] way. Bells on bobtails ring, making spirits [music] bright. What fun it is [singing] to ride and sing a slayighing song tonight. [music] Jingle bells, jingle bells, jingle all the way. Here they say that UA2 is not only the best open model out there, beating Miniax Music 3 and AEP 1.5, but at least according to these benchmarks, they claim that the song quality is even better than Sunno V6, which is pretty crazy.
Now, at the top of the page, they've released everything already. So, if you click on this GitHub link and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. Note that it's fairly tiny, so the model is only 7.3 GB in size, so you should be able to run this on most consumer GPUs. I'm going to do a full installation tutorial and review on this, so stay tuned for that.
For now, if you're interested in reading further, I'll link to this main page in the description below. Also, this week, OpenAI introduces Chat GPT for financial services. This basically combines GPT6 Astra with financial data sets and tools for producing actual banking research. So you can use it to investigate companies, build financial models, and turn the results into spreadsheets, documents, or presentation. It uses data from real providers like Pitchbook, Crunchb, and others.
And you can trace figures back to specific tables or passages so analysts can check where the numbers came from. Firms can also supply their own templates so the output follows their usual format. So if you're in finance, this new feature might help you automate a ton of work. Now, it's not available for everyone. You'll need to contact sales at the bottom of the page to request access. If you're interested in reading further, I'll link to this main page in the description below.
Also, this week, we have a new system called UMR, which stands for unified motion retargeting. This basically translates human movements into movements that a humanoid robot can learn from. That sounds pretty straightforward, right? But humans and robots actually have very different proportions and joints and movement limits. So copying a human's movements directly to a robot doesn't necessarily work well right out of the box.
Well, what UMR does is it represents the surfaces of both bodies as collections of 3D points and then it learns which points correspond. So think of it as like matching the shape of a human pose to a robot's body while respecting the robot's constraints. This also helps preserve contact with objects and the environment. for example, when the robot is picking up a ball, or sitting on a chair. It can even do crazier stuff like spin kicks, climbing stairs, or even playing tennis.
So, a pretty cool system that is able to take human motion data and reuse it across different robots without manually needing to redesign the motion yourself. At the top of the page, they've released the code to this. If you scroll down a bit, here it contains all the instructions on how to set this up locally on your computer. If you're interested in reading further, I'll link to this main page in the description below.
Also, this week, Unirit releases an open-source robot model, which is actually extremely powerful, but it's jam-packed into just 6 billion parameters. It's called Unio LMWLA. And basically, this can take what the robot sees plus your instruction and information about its current state and then produce the correct actions. And this one model is able to cover 64 tasks, including 54 tabletop tasks and 10 involving whole body coordination.
It also supports two finger grippers and different five-finger robot hands. So, it's not limited to just one type of robot. The key idea here is teaching the robot to predict which parts of a scene will change during the interaction. For example, when folding a towel, it learns to anticipate the relevant movement in the scene. And that prediction is connected to an action generating component that produces the robot's next movements.
Here you can see it being able to easily load laundry into a washing machine. Or here you can see it picking up this bottle and placing it in the trash, but then also taking out the trash bag and then walking towards the garbage can to dump it. So this is whole body coordination. Or here's another example of manipulating objects and walking around the kitchen. So from just one tiny 6 billion parameter robot, this can do a ton of different actions.
The awesome thing is they've actually released the training data set and the models to this. So at the top here, if you click on models, you can download this Uni full LM model which is less than 9 GB in size. So this can potentially fit locally and offline in a robot. If you're interested in reading further, I'll link to this main page in the description below. Also this week, if you're looking for a tiny model that's only a handful of parameters, so you can run this offline on edge devices, this might be the best one to use.
So, OpenBM just released their latest model, mini CPM 52B. So, this is a super tiny 2 billion parameter dense model. And compared to other models of similar sizes, this is state-of-the-art. So, you can see across all these different benchmarks in like code reasoning, math reasoning, instruction following, general knowledge, etc., On average, it's better than even Quen 3.54B as well as these other small models. The awesome thing is along with the model, they're also releasing the highquality training data set behind it.
So you can also access the data set here if you're interested in training your own small model. So on this hugging face page, note that at 2 billion parameters, this model is fairly tiny at only 5 GB in size. So this should be able to fit on most consumer devices. If you're interested in reading further, I'll link to this main page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this.
Which piece of news was your favorite? And which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week, I can't possibly cover everything on my YouTube channel. So, to really stay up to date with all that's going on in AI, be sure to subscribe to my free weekly newsletter.
The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.