Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
8,108
Runtime
45:57
Speaking pace
176wpm
Reading time
34min
176 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
AI never sleeps, and this week has been absolutely exhausting. Deep Seek releases their latest model. XAI also releases their best model Grok 4.6, and then ZAI also drops their latest and best model GLM 5.3, which is the best open-source model you can use right now. Google also releases their best and fastest model Gemini 3.7. We have not one, but two state-of-the-art open-source video models. We also have a new state-of-the-art open-source music generator, which is so tiny it can even fit on most low-end GPUs.
88 words, the words spoken in the first 30 seconds at 176 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 476 |
| Average words per sentence | 17.0 |
| Longest sentence | 61 words |
| Questions asked | 9 |
| Sentences containing a number | 135 |
Most used terms
Filler phrases
92 in total: like 56 · actually 18 · basically 9 · kind of 5 · right? 2 · you know 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
AI never sleeps, and this week has been absolutely exhausting. Deep Seek releases their latest model. XAI also releases their best model Grok 4.6, and then ZAI also drops their latest and best model GLM 5.3, which is the best open-source model you can use right now. Google also releases their best and fastest model Gemini 3.7. We have not one, but two state-of-the-art open-source video models. We also have a new state-of-the-art open-source music generator, which is so tiny it can even fit on most low-end GPUs.
Alibaba releases by far the best medium-sized model you can use offline, Qwen 3.8 27B. We have a new state-of-the-art text-to-speech generator, and a lot more. Let's jump right in. First up, we have a new open-source video editor called Joy AI Video Edit. This can edit existing videos just with a prompt in natural language, and this thing is blazing fast. First of all, here are some examples. For example, we can take this input video and transform their outfits and the setting into a castle aristocratic style.
Or we can easily specify which characters to remove in a video, and it's able to seamlessly fill in the background. Or we can turn all these dogs white and add party hats, and also turn the sunglasses from this one dog into hot pink. The awesome thing about this is this also allows you to edit videos in pretty much real time. So, here's a demo of this real-time interface where this person is like swapping the outfit, and as you can see the latency is just like around a second.
So, this is incredibly fast. And as you can see from this video editing benchmark, the quality of its edits are as good as Bernini R or even the closed-source Kling 3 Omni, but you can see this can process things way faster. In terms of the specs of this, this is a 16-billion-parameter multimodal diffusion transformer model, and this can generate 720p videos at over 30 frames per second. And this is actually an auto regressive diffusion model, which means it's designed to process the video chunk by chunk.
The nice thing is they've released the weights to this already, and it's under the Apache 2 license, which has very minimal restrictions. You can even use this for commercial purposes. So, on this page, it contains all the instructions on how to download and run this. And if you inspect the video model, this is 32.5 GB in size, so you do need a high-end GPU to run this. If you're interested in reading further, I'll link to this main page in the description below.
Also this week, Tencent releases a very useful AI called Scope. This basically gives AI video models much better control over camera movement. You see, how this works is you would give it an input image plus a camera path. And it can generate a video that follows that camera movement consistently. So, here are some examples. Here's a pullback and rise example with the same input image. And then also using the same image, here is a push sweep example.
Or instead, here's an S curve reveal. Again, you can see the camera movement reference in the bottom corner. Here's another example of a crane up. Or instead, we can take the same image and then have it dolly in. So, a very useful framework if you want to control the exact camera movement by specifying a path. At the top here, they've already released the code to this. So, if you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer.
Note that this is based off of 12.2 and Diff Singer Studio. And it's under the Apache 2 license, which has very minimal restrictions. If you're interested in reading further, I'll link to this main page in the description below. Also this week, Deep Seek releases their latest model Deep Seek V4 Pro August 13th. This is a massive 1.7 trillion parameter mixture of experts model, and this is similar to the previous Pro version, but they've also added this D-Spark speculative decoding.
If you're not familiar with D-Spark, it's a pretty big deal. I already did a full explainer video on it, so see this video to learn more. Anyway, as you can see, this latest V4 Pro version even performs as well as the leading open models, GLM, Kimiko 3, as well as some top closed models like Opus across all these different knowledge and agentic coding benchmarks, including Humanity's Last Exam, Terminal Bench, DeepSue, etc.
Now, if you look at this independent leaderboard by Artificial Analysis, then you can see that DeepSeek V4 is over here, so slightly below the open Kimiko 3 and tied with GLM 5.2. Still a few points behind the leading closed models like GPT and Grok 4.6. However, if you look at the cost of this, here is where DeepSeek V4 just destroys its competition. You can see that at just 6 cents per task, this is way cheaper than Kimiko 3 as well as GPT 5.6 and the ridiculously overpriced Claude models.
So, in terms of intelligence versus price, this is actually the best model to use. Here's a chart mapping out intelligence versus cost, and ideally you want to be in this upper left corner, and as you can see, DeepSeek V4 is like the only model that's really within this quadrant. Now, at 1.7 trillion parameters, this model is quite huge, so the total size of the raw model is 893 GB in size. You'll need to stack like multiple accelerators to run this, but because this is open source, the community is already working on creating more quantized and compressed versions of this that can run with lower memory.
So, for example, Unsloth will likely release some GGUFs of this very soon. Now, in addition to this new V4 Pro model, DeepSeek also finally releases their own harness called DeepSeek Harness. If you're not familiar with the term harness, this is basically like a framework that orchestrates an agentic model. For example, Claude Code would be the harness for Claude, or Codex would be the harness for GPT. And well, in the past, DeepSeek had lacked their own harness.
So, you had to use it in some other third-party frameworks like Hermes or Open Claw. But, this week they finally released their own harness. Now, in general, it's best to use the harness that's built by the same company as the model. So, in the future, I would expect that this would be the best harness to use for DeepSeek models. Now, currently, this is still in developer preview, and it's iterating rapidly, so expect compatibility breaking changes.
But, if you are interested in getting ahead start and trying this out, here it contains all the instructions on how to download and run this. If you're interested in reading further, I'll link to this main page in the description below. Also this week, xAI releases their top model Grok 4.6. And this is a big deal. They've actually caught up to frontier. As you can see from this artificial analysis intelligence index, Grok 4.6 is just as good as the leading GPT 5.6 Soul Max plus Claude Fable 5.
For GDP val, which measures how well an AI model performs on real professional jobs that contribute to the economy, you can see that Grok 4.6 even outperforms Fable 5 and GPT 5.6. For Deep Swe, which is a measure of agentic software engineering capabilities, you can see that it's still quite behind GPT 5.6. But, in terms of these other agentic coding benchmarks, then it does seem to do slightly better than GPT. Now, like most frontier models out there, this is trained on a wide range of agentic reinforcement learning tasks, including knowledge work, general coding, kernel optimization, web dev, etc.
And here it says that it's particularly strong at turning a vague idea into a working first version, including interactive and visual projects. It also appears to do more self-checking during longer tasks, meaning it can inspect its own work before continuing. If you look at the official artificial analysis leaderboard, then you can see that Grok 4.6 is tied with GPT 5.6 and 1 point above Chimera 3. It's speed is also quite impressive.
Not as fast as DeepSeek-V4-Pro, but still faster than GPT or the Claude models. The cost is also very impressive. This is the same as Qwen 1.5 3, which is cheaper than GPT-5.6. So, a very cost-efficient option at Frontier Intelligence. Now, this is closed and paid. Here it says Grok 4.6 is available in Cursor and Grok Build, and also available via API and other providers. If you're interested in reading further, I'll link to this main page in the description below.
Now, last week I mentioned that Alibaba announced Qwen 3.8 Max, which is their massive 2.4 trillion parameter model, and this is the first ever Max model, which they say will be open source. Well, this week they stuck to their word, and they actually released the entire model. So, this is a 2.4 trillion parameter mixture-of-experts model. When you use it, only 95 billion parameters are active, so this is fairly efficient.
And if you look at all these agentic coding benchmarks, as well as general capabilities and knowledge, for most of them, Qwen 3.8 Max does come close or even match the performance of the leading frontier closed models, which is very impressive. If you look at this leaderboard, you can see it's only two points below Qwen 1.5 3, but it does score better than DeepSeek-V4, and it's edging very close to Claude and GPT-5.6.
Now, this is open source, so you can download it and run it for free locally if you have the hardware, but if you don't, you can also use it via their API. And here's the cost per task, so it does seem to be slightly more expensive than Qwen 1.5, but still cheaper than GPT and Claude. Now, at 2.4 trillion parameters, this is a massive model at 4.8 terabytes in size. So, you're going to need to stack like a ton of accelerators to actually fit everything.
But because this is open source, I'm sure there's going to be more quantized versions of this coming very soon. Still though, props to the Alibaba team for spending millions of dollars and a ton of time and compute to train such a massive model, and then just releasing this intelligence for free. If you're interested in reading further, I'll link to this main page in the description below. Also this week, Xiaomi releases a really useful AI for creating audio called Mod A Shang LM Gen.
This allows you to generate complete audio scenes including things like speech, music, sound effects, and environmental noises. So, here's a demo of some sound effects with environmental noises. Very nice. The cool thing is this can also generate music. Here's an example of some music played by a brass quintet. >> [music] [music] >> Now this ain't a dedicated music generator, so the quality isn't as good as, you know, the best open music models out there, but pretty impressive how it's able to understand and also generate music.
Now, like I said, this can also incorporate speech in its generations. So, for example, we can have a male voice describing the vehicle's condition with engine idling in the background, and here's the transcript he should say. >> The top is in excellent shape. It does have a gray interior. >> Or here's another example with a female speaking this out with some instrumental background music. >> Get it right? If you don't have any clients, you don't have a law practice, and that's why business development really [music] is the most important skill that a private practice lawyer can teach you. >> This can also handle emotions and different languages.
For example, we can get a really angry person expressing intense frustration and rage, and let's get him to speak out this Chinese. So, a super flexible tool for generating all types of audio within the same generation. The nice thing is they've released this already, so if you click on this GitHub repo, it contains all the instructions on how to download and run this locally on your computer. The model is also fairly tiny, so the total size of everything is like less than 12 GB in size, so you can easily fit this on like a mid-end GPU.
If you're interested in reading further, I'll link to this main page in the description below. What if you never have to take notes again, but you could still record everything that happened? Well, definitely check out Second Brain by JenSpark, the sponsor of this video. They just released a personal memory system called Second Brain. This can autonomously capture everything for you, including meetings, conversations, ideas, and turn it into something you can actually use later.
It's a combination of a tiny wearable device and an AI-powered platform. The hardware part is called Second Brain Note. This is a tiny wearable recorder that's about the size of a credit card. It snaps onto your phone or slips right into your wallet, so it's always with you. It packs a 35-hour battery, a four-microphone array with bone conduction for clear audio capture, and even records directly to its own 64 GB of local storage.
To start recording, just press and hold the button for a couple of seconds until you feel a vibration. And whenever someone says something important, just tap the button once to bookmark the exact moment, so you can jump straight back to it later. And this is where the software side comes in. Second Brain connects to the tools you already use, like email, calendar, Notion, Google Workspace, and even CRM platforms like HubSpot.
With your permission, all of that becomes your personal context layer. Instead of dumping everything into one giant folder, it organizes your life for people, companies, projects, and knowledge, almost like a personal Wikipedia built from your own work. You can open up a meeting, instantly see polished AI summaries, jump directly to your bookmarked highlights, search for conversations across weeks or months, or trace how an idea evolved over time.
Unlike normal AI chats that start fresh every session, every meeting or conversation compounds over time, so your memory keeps getting smarter instead of resetting. And because it understands all this context, it can answer questions across everything you've recorded. And if you wanted to actually take action, you can switch over to GenSpark Super Agent, which uses the same context to help draft emails, create proposals, build documents, and more.
Setup is incredibly simple. Just update the GenSpark app, unlock Second Brain, and pair your Second Brain note, and you're ready to go. From then on, everything records, syncs, and organizes itself automatically. Second Brain gives you a memory that never forgets and only gets better over time. This is a limited first release, so if you want to grab it with 10% off, check out the link in the description below. Also this week, we have yet another open-source video model called LTX 2.5.
Now, this has previously been a very good model. It has audio natively built in. It can handle a ton of different scenes and artistic styles and motion. One of the biggest upgrades is this native multi-shot feature, where it can generate multiple shots of a scene in a single clip. But again, so can the other current frontier models. They say it has cleaner motions, so it's smoother and more natural results. It's also better at understanding your prompts, and this is insanely fast.
This is definitely the fastest open video model you can use. Now, of course, a ton of you are looking for a comparison between this and MiniMax H3, so here it is. I tested it on a series of diverse and tricky prompts. My first test is a fight scene, so I inputted this image plus this prompt, and here's the result from both. Both are not great, but MiniMax does look way more coherent, whereas for LTX 2.5, the motion is kind of awful.
They kind of switched sides halfway, and then the dude with the white suit somehow slipped on the floor or something and fell down. It's just very weird. Next, here's a subtle expression and emotion test. I imported this image, and the prompt is she opens a letter, spark of hope still in her expression, but after she reads it, extreme grief presses against her carefully held composure. Her eyes grow heavy and glassy, lower lids trembling before a single tear slips free and tracks slowly down her cheek.
Both are not great, but if I had to pick a winner, again, I would choose Min Max. It just looks slightly more realistic and faithful to my prompt. Next, here's our tricky continuous shot prompt. A satellite view of planet Earth, then zoom into a drone view of New York City, then zoom into an office building, then zoom into a view of a person scrolling Tik Tok. One continuous shot. Again, if I were to choose a winner, I would go with Min Max.
LTX tried so hard, but this is just a lot of places where Min Max handled this better. All right, next, here's a test on its camera movements. So, I imported this image, and then let's start with a forward dolly through this dark hallway, and then transition into a smooth orbit around a woman in the doorway before rising to a high crane shot that reveals the empty room behind her. >> I would say I prefer LTX's generation better here.
The camera actually orbited as well. Whereas for Mini Max, I didn't really see that orbit. And then next, here's an anime example. So, I inputted this image plus it's going to be a slow zoom in. The guy says this and the girl says this. Now here clearly you can see that Mini Max is better. For LTX, it kind of changed the faces of the characters. Whereas for Mini Max, it kept the consistency. Everything looks very natural.
This looks like a legit anime scene. And then here's another example testing its text rendering capabilities. So, the couple is kissing. We push in and then tilt up to show the sky with the text and they lived happily ever after. So, in terms of text rendering as you can see again, Mini Max is the clear winner. For LTX, there were some misspellings plus the text didn't look as good. And then finally, my favorite prompt which no AI model has gotten correct so far, a professor explaining the Pythagorean theorem on the whiteboard. >> The Pythagorean theorem states that in In right angle triangle the square of the hypotenuse And so, the square of the hypotenuse equals the sum of the squares of the other two sides. >> Well, both were not entirely correct.
Again, I would have to give the points to MiniMax, which could actually explain what it is. Plus, the dude actually wrote the correct equation on the whiteboard. So, that was a quick summary and comparison of LTX versus MiniMax across a series of diverse prompts. So, in almost cases, the quality of MiniMax is definitely better, but the advantage of LTX is it's extremely fast. I can generate the same video in like half the time compared to MiniMax.
And awesome thing is they've released the model to this already. So, if you click on this download open weights button, here it contains all the models you need to download. They've already added support for ComfyUI, so they have a int8 config version, which is only 22 GB in size. Plus, other platforms like Run Diffusion has also added LTX 2.5 for you to use. And the best thing about LTX 2.5 is it also works with previous LTX 2 LoRAs.
So, you don't need to retrain any existing LoRAs to make it compatible with 2.5. If you're interested in reading further, I'll link to this main page in the description below. And let me know in the comments if you want me to do a full tutorial on this. I'm thinking of skipping this because there's really no point since we already have MiniMax, but let me know if you still want a tutorial. Also this week, we have some interesting updates from Google.
Even though they're not really caught up in the AI model race, Google DeepMind continues to release some really useful AI tools. So, last week I talked about Weather Next, which is a state-of-the-art method for predicting cyclones. And this week, they released a sign language-to-text model. So, this allows a deaf or hard-of-hearing person to use sign language directly on their phone's camera, and this AI will turn that into text.
So, instead of typing a message, they can just use sign language, and it'll turn that into text. Using the live transcribe feature, you can also just use sign language during a conversation and have the system produce text in real time. Here, Google says that this model was trained on over 100,000 hours of data across more than 50 sign languages. And this is currently the most capable sign language translation model to date.
It achieved a remarkable score of 70 on this benchmark, which is significantly higher than any previously reported score. So, the model currently launches with American Sign Language, and this will be available on the Gboard keyboard and live transcribe on Pixel 11. And there's going to be support for more devices and more languages very soon. So, a very useful tool for people who are deaf or have hearing issues. If you're interested in learning more, I'll link to this main page in the description below.
Also this week, OpenAI previews their ultra-fast mode for their best model GPT-5.6 Soul. Specifically, this can generate up to 750 output tokens per second, which is 14 times faster than its standard speed. This is, of course, important in like incidents response and reliability. For example, when something critical fails, then you need to handle it immediately. Same with like financial research and security, especially with real-time quant trading, then you'll need an AI that can respond super quickly.
Same with customer support and voice and so on. Here's the insane speed of ultra-fast on the left compared to the standard speed on the right. You can see within seconds, ultra-fast has already completed this task, whereas the standard version is still working on it. The important detail here is that this is powered through their partnership with Cerebras, which enables much faster inference. To give you a sense of how crazy 750 output tokens per second is, here is a chart showing the output tokens per second of the top models.
You can see like the leading Grok or Claude and GPT models are only 60-something tokens per second. GLM 5.2 is 111. Even the fastest Gemini 1.5, which was released today, is only 340 tokens per second. So, this new ultra-fast GPT is like more than double the speed of Gemini 3.7 Flash, which is crazy. Now, before you get too excited, here it says this ultra-fast mode is available in a limited preview only to a select group of customers, so basically in private preview.
They are planning to expand access as capacity grows. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a super exciting update for local AI users. So, Alibaba just dropped their latest medium-sized model Qwen 3.8 27B. If you're not familiar with the 27B sized models, basically this is by far the most popular medium-sized model that you can run locally on a high-end device.
For example, the previous Qwen 3.6 27B has gotten over 7 million downloads. No other medium-sized model even comes close to it including Meta Muse Glimmer or Google's Gemma or or any other competitors. Qwen 3.6 was just way better. Well, this week they released an even better version Qwen 3.8 27B. And here are the results of this. You can see across most of these agentic performance and coding benchmarks as well as general knowledge, it even outperforms Opus 4.6 Max.
That's crazy. Like we now have the performance and intelligence of Opus 4.6 Max, which you can run locally on just a high-end GPU. God bless the Alibaba team. Now, here are the specs of this. First of all, this is designed to do very well in agentic coding and knowledge work tasks. Plus, this is multimodal, so it has native support for image and video understanding, which is fantastic. This is a 27 billion parameter dense model with a vision encoder, and you can extend this to up to a million token context window.
If you look at this independent leaderboard by Artificial Analysis on similar sized models, they haven't added Qwen 3.8 27B yet, but as you can see, the previous 3.6 version is already number one. So, you can expect the 3.8 to perform even better. Now, like I said, they've already open-sourced this. So, if you click on files and versions, you can see that the full model is 56 GB in size, which can easily fit on a high-end GPU.
They also released a more quantized FP8 version, which is only 30 GB in size. And on Sloth, has already released some GGUFs on this. And get this, the smallest Q2 one is only 9 GB in size. So, you can easily fit this on even a low to mid GPU. It's crazy to imagine that we basically have the intelligence of Opus 4.6 Max squished into just 9 GB. Anyway, on this Hugging Face page, it contains instructions on how to download and run this locally on your computer.
If you're interested in reading further, I'll link to this main page in the description below. Also this week, my favorite AI lab ZAI has finally dropped their latest model GLM 5.3. And this is an absolute beast. Let's jump in with the benchmark scores. So, if you compare this latest 5.3 against the previous version 5.2, which is the green bar, you can see the improvement is pretty crazy. For Criminal Bench, you can see it vastly outperforms the leading open model Kimik 3.
For Deep Seek, also a huge improvement compared to GLM 5.2. For Agent's Last Exam, it's pretty much frontier. Same with GDP Eval, which tests any AI models performance on real-world economically valuable knowledge work across various jobs. You can see that it is the world's best model. Same with Bench, it's the best model in the world. And then for Humanities Last Exam, this tests how knowledgeable an AI is on like really obscure subjects.
Again, it's edging very close to the leading Claude and GPT models. And you know, what's even more impressive about this is they didn't make a new model from scratch. Heck, they just took GLM 5.2. They didn't even make the model bigger or fundamentally change the architecture. They just took it and post-trained it harder. They gave it more environments, more diverse tasks, and basically more compute post-training this GLM 5.2.
And from that alone, it already led to major gains, especially in complex coding and long-horizon tasks and cybersecurity. In fact, this thing is an absolute beast in terms of cybersecurity. You can see in terms of Cyber Gym, it's pretty much the best model in the world, even outperforming Fable 5 and GPT 5.6. And in terms of Exploit Bench and Exploit Gym, again, note the insane improvement over the previous GLM 5.2 as well as the current open leader Kimmy K3.
This is currently the best open model you can use for cybersecurity tasks. ZAI reports that this GLM 5.3 has already found thousands of vulnerabilities across hundreds of open-source projects. Now, the same capabilities that make a model useful for finding and fixing security problems can also make it more capable of carrying out attacks, right? So, that's why ZAI is doing some additional safety testing before they publicly release the weights to this.
So, it says that they're going to open source this in around 2 weeks. But currently, you can subscribe to the Z Code plan and then try out GLM 5.3 in whatever coding agent you want. Now, today is reserved for my weekly news video, but I will definitely post a full review on GLM 5.3 including all the impressive things that it can do probably tomorrow. So, stay tuned for that. For now, if you're interested in reading further, I'll link to this main page in the description below.
Also this week, Google releases their best and fastest model Gemini 3.5 Flash. Now, as a flash model, this is not meant to be the most intelligent and performant model out there, but if you look at its abilities compared with other small or flash models such as Sonnet 5 or GPT 5.6 Tera, then you can see that at least according to this like frontier code benchmark, Gemini 3.7 Flash is number one. Here's its performance on Deep Sweep, not as good as GPT 5.6 Turbo, but it is pretty close.
For web dev, Gemini 3.7 Flash is pretty good. Same with PDF document comprehension. The nice thing about Gemini models is they're multimodal. So, they can not only take in text, but also images, video, and audio, and documents. It's still one of my favorite models to use when I need to upload a video for it to analyze or upload some audio for it to transcribe. Here's an example where we can get it to create a 3D game using assets from Nano Banana.
And as you can see, Gemini 3.7 Flash is able to execute this very well. Or here's an example where we can get it to create this interactive parallax website where if you scroll down the website, it changes the view of the scene. And here's our final result. Note that these images were generated with Nano Banana. Or here's another example where we can just enter a PDF and get it to create a website from the information in the document.
Here's a chart on its Deep Sweep performance. It's quite messy, but as you can see, Gemini 3.7 Flash is all the way over here. At least compared with the other Gemini models in blue, it is the cheapest and the most performant. However, if you compare this with this green bar over here, which I believe is GPT 5.6 Luna, then as you can see from this green dot here, it's still more performant than Gemini 3.7 and it's like three times cheaper.
If you look at this leaderboard by Artificial Analysis, then 3.7 Flash does perform better than the max version of GPT Luna. And as you can see here, the strength of Gemini 3.7 Flash is that it's incredibly fast. It has 340 output tokens per second, which is way faster than any of the other competitors. But the thing is, it is more expensive than the other small models. For example, here you can see that Gemini 3.7 is 40 cents per task, whereas for GPT 5.6 Luna Max, it only costs 5 cents per task.
So, if you want to use Gemini 3.7, it's basically a trade-off between speed versus cost. It's more expensive, but it's a lot faster than GPT Luna. Now, here it says that 3.7 flash is available in Google's anti-gravity, which is like their agent decoding platform, as well as Google's AI studio and Android studio. And then for individuals, this is available in the Gemini app, but only for pro and ultra subscribers in supported countries.
Anyway, if you're interested in reading further, I'll link to this main page in the description below. Also this week, Minimax continues to cook. If you haven't been up to date, basically last week Minimax dropped by far the best open-source video model out there, Minimax H3. Well, this week they cooked again by dropping the best open-source music generator out there, Minimax music 3. This can generate a variety of songs in different styles, and like other music generators, you can basically enter a prompt describing the genre, the speed, even the key, the instruments to use, the overall vibe, etc.
And then you would specify lyrics, so you can include meta tags like intro, outro, verse, bridge, chorus, etc. And it can generate some super clean and professional-sounding songs if you give it the right prompt. The best thing about this is it's super tiny. So, the full model is only 9.8 GB in size, which should already fit on like mid-end GPUs, but there's also an int8 in size. So, this can even fit on like most low-end GPUs, making this super accessible.
Now, I already did a full tutorial and review on this, so I'm not going to repeat too much here. See this video if you're interested in learning more. Also this week, we have a new state-of-the-art text-to-speech generator called Index TTS 2.5. Now, previously I did a tutorial on version 2, which was already state-of-the-art back then. This version is even better. So, how this works is this can just take a few seconds of a reference voice and make them say anything.
For example, here's our input voice. >> Caring for others, never let your bags out of your sight, especially when you are crossing international borders. >> And then let's get the voice to speak this out. Here's the generation from Index TTS. >> Animal Liberation and the Royal Society for the Prevention of Cruelty to Animals, RSPCA, are again calling for the mandatory installation of CCTV cameras in all Australian abattoirs. >> As you can hear, it sounds extremely similar to the input voice.
This can also handle emotions, so here's an angry example. Here's the input voice first. >> In what bucolic school offense he had been taught was beyond imagining. >> All right, so let's get that voice to speak this out. >> If you'd ever treated me like a human being, would they have dared to do this? >> And this can also handle multiple languages. So, here's a surprise example with Spanish. Here's the input voice first.
All right, and let's get that voice to read out this transcript. So, a super flexible tool, one of the best use cases for this is you can clone the reference voice and get them to speak out a different language. So, here's an example where we can take this scene and use Index TTS to dub it from Chinese to Spanish. This is definitely one of the most flexible and highest quality text-to-speech generators you can use right now.
The nice thing is, as with the previous models, they've already open-sourced Index TTS 2.5. So, on this page, it contains all the instructions on how to download and run this. If you check out their Hugging Face, note that the total size of everything is only like 5.5 GB in size. So, this should be able to fit on most consumer devices. If you're interested in reading further, I'll link to this main page in the description below.
Now, in addition to LCM 2.5, we have yet another open-source video generator this week called Magai 2. Now, while LCM 2.5 is small and fast, this one is like the complete opposite. This thing is a huge model at 114 billion parameters, but it's a mixture of experts, so when it only 6 parameters are active. And like the other frontier models out there, this also has audio baked in. Now, there are limitations to this. Currently, this only supports 10-second generations.
This also has a refiner component, which can generate results up to 1080p. But, the other frontier models out there can actually generate even higher resolution. Now, like I said, this thing is massive. The model itself is like 228 GB. You definitely won't be able to run this on consumer hardware. In fact, here it says you're going to need some Nvidia Hopper GPUs, and not just one, but eight of them. So, honestly, while it's commendable that they open-source this, I don't think it's actually usable for most consumers, at least not locally.
But, if you're still interested, especially in reading the architecture and the details behind this, I will link to this main page in the description below. Next up, we have a super tiny AI model, which is really useful for running offline on small devices. You see, for large language models, even like the smallest ones are billions of parameters in size, which often require a few gigabytes of VRAM on your GPU. Well, what if you don't even have that?
That's where this cactus needle model comes in. So, this is designed for devices that are way too small and weak to run larger models. The specs here are kind of crazy. So, this needle 2 is only 45 million parameters, not even 100 million or a billion parameters. The whole thing is packaged into a 14 MB binary that runs using about just 28 MB of RAM. So, you don't even need VRAM on a GPU to run this. Now, of course, a model that's this small would be way less intelligent than a larger model, but this is still useful for things like controlling devices, calling tools, or extracting information from documents.
Where it falls short is if you try to get it to do some really long horizon agentic or reasoning task, which requires more intellectual power, then it probably won't be able to do that very well. And this thing is insanely fast and lightweight. So, get this. This can decode at 500 tokens per second on a Raspberry Pi 5, or even up to 1,500 tokens per second on VR devices, as well as 700 tokens per second on cheap phones under 200 bucks.
So, here's a chart showing its performance versus total parameters, and as you can see, Needle 2 is all the way over here. Not the most performant, but it's way smaller. The awesome thing is they've released this already. So, if you click on this icon and you scroll down a bit, here it contains the instructions on how to download and run this locally on pretty much any device. If you're interested in reading further, I'll link to this main page in the description below.
Also this week, Tencent XuanYan continues to release some really useful tools for 3D generation. So, this week they released World Claw. How this works is instead of just making one object at a time, the goal is to generate entire open 3D worlds from an open-ended text description. So, here are some examples where you can get it to generate a snowy village, or we can generate a desert battlefield, and this contains everything from like the depth and normal information, all the way down to even generating individual objects.
Now, of course, with all these things generated for you, it's really easy to reuse this entire world downstream. For example, in video game design and production. And how this works is quite interesting. So, it uses multiple agents to first plan the world, including the layout, the materials, the terrain, etc. And then it generates all these assets sequentially from coarse to fine. And then it also inspects and refined it to make sure that everything is physically consistent.
They have released a GitHub to this. Currently, there's no indication whether they will open source this, but if you are interested in learning more and checking out more examples, I'll link to this main page in the description below. Also this week, we have a really interesting robotics model called the Dyna 2, and they're tackling one of the biggest problems in robotics, which is how do you train a robot what to do without collecting a ton of robot specific training data.
Well, the answer is surprisingly simple. Just learn from humans. So, Dyna is a world action model trained on more than 1 million hours of first-person human videos, which is roughly like 170 years of continuous experience. These include things like folding clothes, cooking, cleaning, assembling objects, etc. The model learns not just what things look like, but how the world changes when someone interacts with them. Then the researchers tested whether this knowledge can transfer to robots.
And here's the really interesting result. They found a scaling law, meaning that as they increased the amount of human data used for training, performance also improved on robot data that the model had never seen. So, this model can be adapted to different robots with only a small amount of robot data. And so, with this model as the brain inside these robots, it's now able to do a ton of different tasks from folding clothes, cleaning, and manipulating different objects.
So, a very fascinating work which shows that we don't actually need to collect a ton of robot data to train robots well. We can just plug it with a ton of videos of humans doing stuff, and this is enough to teach a robot how to interact with the world. If you're interested in reading further, I'll link to this main page in the description below. Also this week, NVIDIA releases their latest open-source model called NeMoTron 3.5 Lightning.
They also released a model router called NeMo SwitchYard, which helps you automatically pick the best model for each task. Now, this new NeMoTron 3.5 Lightning is a pretty small 30 billion parameter mixture of experts model. This has a 1 million token context window. Now, if you compare the intelligence of NeMoTron 3 Lightning against other similar-sized models, then it is considerably less intelligent than like Qwen 3.6, which is all the way over here.
But, it does come close to like Qwen 3.5 as well as Gemma 4. While it's not the most intelligent in its size range, the advantage of this is it's incredibly fast. So, here are its output tokens per second, and as you can see, this is way faster, pretty much double the speed of Qwen 3.6, which is over here, and even faster than Gemini 3.6 Flash, which is already one of the fastest models out there. So, if you care about speed, then this is definitely the fastest model in around the 30 billion parameter size range.
And then, they also released NeMo SwitchYard, which is an open-source router that picks the best model for each task. Think of it like a traffic controller for AI models. So, if you give it a complicated workflow, it might decide to give the task to a more intelligent agent, whereas if you give it a simpler task, then it might give it to a smaller and faster agent, which would help you optimize speed and cost. Here, they showed that if you use NeMo SwitchYard with Opus 4.8 as well as these other models, then not only does it actually complete more tasks than just using Opus 4.8, but it's like three times less expensive.
The nice thing is, both the NeMoTron Lightning model plus the SwitchYard model are open-source. So, here they've released the model on Hugging Face, and the total size of everything is only around 22 GB in size, at least for the FP4 version. So, you should be able to fit this on mid to high-end GPUs. Again, if you care about speed, this is the fastest medium-sized model you can use right now. If you're interested in reading further, I'll link to this main page in the description below.
Also this week, we have a pretty interesting project called Matrix, which is spelled like this, and this is trying to simulate the entire human population with 8.3 billion persona agents. You see, a lot of money and effort are poured into like surveying people, for example, for census reports, sentiment analysis, or getting users to test websites or apps to see the user journey and what needs to be improved. Well, instead of getting real people to do these tasks, what if we could just simulate their actions using AI agents?
In other words, how do you test a product on millions or even billions of different kinds of people without actually finding billions of these people. So, the idea is to create these simulated users or what they call persona agents representing different human characteristics, and then let those agents interact with the products. Here, they describe their system as having 8.3 billion persona agents with 1,290 persona attributes and more than 1,000 applications.
So, if you're building an app, for example, you could use this to have thousands of simulated users with different backgrounds, preferences, and behaviors interact with it. So, it's kind of like simulating thousands of different people using the app. This can also simulate like surveys or shopping sites where you could get these simulated users to try to check out. Or for an AI chatbot, they could test things like helpfulness, safety, or reliability across multiple conversations.
And in the end, you get feedback, scores, and interaction data. Now, the main pushback of this is, of course, these aren't real people. These are just simulated people, so how closely does this data actually match real humans? How closely do their activities and their preferences and personalities actually match real humans? It's too early to say for now, but this is quite a fascinating concept. And come to think of it, maybe we are also just persona agents living in a simulation.
Let me know what you think of this in the comments below. Anyway, if you're interested in reading more, I'll link to this main page in the description below. Also this week, Meta releases a medium-sized model called Muse glimmer, which is open source. So, this is a 30-billion parameter dense model, so this is not a mixture of experts. And this is under the Apache 2 license. Now, on their official release page here, they're comparing this to Gemma 4 and Qwen 3.6, which are similar sized.
And this does look misleading because the best performer is actually highlighted in black, not in blue. So, as you can see, it does beat Gemma 4, but for some of these instances, Qwen 3.6 does perform better. Now, Meta is known to post some pretty misleading and cherry-picked results. So, if you instead look at this artificial analysis leaderboard, then you can see that Muse Glimmer does not, in fact, outperform Qwen 3.6.
Note that Qwen 3.6 has even fewer parameters than Muse Glimmer. Plus, we are expecting a Qwen 3.8 version coming very soon. But if you are interested, you can click on this download model link, and here you can see the full model is around 60 GB in size. They've also released some compressed GGUF versions of this, which are much smaller. Now, I think this release is quite underwhelming compared to Qwen 3.6. But if you are interested in trying this out, I'll link to this page in the description below.
And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite, and which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week, I can't possibly cover everything on my YouTube channel.
So, to really stay up-to-date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching, and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.