Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
13:263.0x the video's typical replay level
or any other content, it gives you a lot more control while keeping everything consistent. You can try Seance 2.5 on Higsfield today using the link in the description below. Now around 2 weeks ago, Moonshot AI released Kimmy K3 and this is a massive 2.8 a trillion
Said at 13:20
Most replayed moment #2
4:422.5x the video's typical replay level
simple. This is model agnostic. So, they used Quinn, but you can also switch it up with another model. Now, if you're interested in running this on this page, if you scroll down a bit, it contains all the instructions on how to download and run this locally on your computer. If you're interested in reading further,
Said at 4:34
Most replayed moment #3
3:222.3x the video's typical replay level
computer. Now, the models to this are fairly tiny. It's only like under 5 megabytes in size, so you should be able to fit this on most consumer devices. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, if
Said at 3:16
The graph counts replays. It does not show where viewers stopped watching.
Words
6,001
Runtime
33:35
Speaking pace
179wpm
Reading time
25min
179 words per minute, just under the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
AI never sleeps and this week has been absolutely insane. We have a new open-source AI for medical research and reporting. Alibaba releases their latest Quinn model and it's an absolute beast. This AI can compose an entire symphony from scratch. We have another AI that can create super realistic singing voices. OpenAI's internal model just solved some massive math breakthroughs. We have a new open source AI for creating 3D CAD files. Google releases and open sources their state-of-the-art AI for predicting cyclones and natural disasters. We have a
90 words, the words spoken in the first 30 seconds at 179 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 347 |
| Average words per sentence | 17.3 |
| Longest sentence | 64 words |
| Questions asked | 4 |
| Sentences containing a number | 91 |
Most used terms
Filler phrases
76 in total: like 44 · basically 13 · actually 12 · I mean 3 · kind of 2 · you know 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
AI never sleeps and this week has been absolutely insane. We have a new open-source AI for medical research and reporting. Alibaba releases their latest Quinn model and it's an absolute beast. This AI can compose an entire symphony from scratch. We have another AI that can create super realistic singing voices. OpenAI's internal model just solved some massive math breakthroughs. We have a new open source AI for creating 3D CAD files.
Google releases and open sources their state-of-the-art AI for predicting cyclones and natural disasters. We have a new all-in-one 3D model generator and editor. Plus, this can also segment a model into separate parts. We have some exciting robotics demos and a lot more. So, let's jump right in. First up, we have a new AI for creating symphony music. It's called Symphony Gen, and this is designed to create full orchestral music while giving you control over the underlying harmony.
So, here are a few examples. [music] >> [music] [music] >> Now this is actually quite challenging because the model has to coordinate many instruments and notes and the overall structure of the piece at the same time. And actually what this does is it first generates a harmony skeleton and then it expands the outline into a complete orchestral arrangement. So think of it as first sketching the chords and the music direction and then filling in the notes for each instrument.
For example, you can start with a major chord and here are a few examples with this start. [music] >> [music] >> All right. So, as you can hear, they kind of start with the same major chord. The nice thing is you can also input your own harmony skeleton or it can also take the harmony skeleton from an existing piece. For example, here's an original piece and then I'm going to play you the AI generated piece which takes inspiration from this harmony skeleton.
[music] >> [music] [music] >> So, as you can hear, it has roughly the same chord pattern and tempo as the original XRP. The awesome thing is they've released the models to this already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. Now, the models to this are fairly tiny. It's only like under 5 megabytes in size, so you should be able to fit this on most consumer devices.
If you're interested in reading further, I'll link to this main page in the description below. Also, this week, if you're into 3D modeling or printing, this AI is super useful. It's called MAC, which stands for multi-agent CAD. And this is a really efficient AI that can help you create printable 3D models in CAD format. And all it takes is a text prompt. So for example, here is the text prompt on the left. And it'll generate a CAD file which you can print out in 3D as you can see on the right.
Here are some additional examples for your reference. As you can see, it can design a ton of different objects with different shapes and articulations. Here are some additional examples for your reference. And you can also see the cost per prompt listed below each example. And as you can see after adding this multi- aent CAD system, it's a lot more efficient. It's able to complete the task most of the time at like 10 times lower the cost compared with if you didn't implement this system.
In fact, if you compare this to another text to CAD model called CAD skills, you can see that this new Mac one uses 116 times fewer tokens. It costs 13 times less. Plus, the pass rate is also much higher. And getting this setup is also super simple. This is model agnostic. So, they used Quinn, but you can also switch it up with another model. Now, if you're interested in running this on this page, if you scroll down a bit, it contains all the instructions on how to download and run this locally on your computer.
If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Alibaba dropped their latest animation system, One Animate 2. This basically takes a photo of any character plus a reference video and it'll animate the character according to the reference video as you can see in this example. Note that it works with hands and fingers as well. Plus, this doesn't even have to be human characters.
So, here's an example where we can animate this teddy bear instead. And in addition to hands and fingers, this can also transfer facial expressions from a reference video. As you can see in this example, the nice thing is this can also animate multiple characters. So, for example, you can input a photo with two characters plus a video with two characters and it'll transfer their movements over very naturally like this.
Alternatively, you can just use a reference video with one person, but use that to animate multiple characters in a photo, as you can see in this example. And the nice thing about One Animate 2 is that you can also animate characters with irregular body proportions. So, it doesn't have to be just human to human. Another new feature about One Animate 2 is you can also determine the camera angle. Here's an example with the same inputs, but you can also change the view to the left or to the right.
And the nice thing is they've also released a smaller version called Wan Anime 2 Light, which even allows for real-time streaming if you have the right hardware. You can see like the latency is under a second. So, this can be really useful for streaming. Now, I featured a ton of these tools before, including One Animate and Dream Actor, but this new One Animate 2 is a lot more detailed and consistent and natural. Now, they didn't compare this with another leading animation tool called Scale 2.
I think the quality of this versus Scale 2 is very similar. The awesome thing is they've released this already. Now, the full one animate is around 33 GB, so you'll need a high-end GPU to run this. And there's also already support for Comfy UI. They've released an int 8 version which is half the size and this should be able to fit on just a mid-tier GPU. If you're interested in reading further, I'll link to this main page in the description below.
Also, this week we have a new AI called vocal render and like the name implies, this can generate singing voices for songs and it actually sounds incredibly realistic and expressive. So, here's what it does. This takes in lyrics and then the melody in MIDI notes and that's about it. So from this it can output a very expressive voice singing out these lyrics. Now they released two different models, Vocal Render and Vocal Render Pro.
The Pro version just sounds a bit better. Let's listen to both. [gasps] >> And then here is an example with the pro version. Now, if you compare this with other singing [snorts] voice generators like Vivo 2 or Soul X, you can hear that vocal render is a lot better. So, let me just play you the other competitors as well. [gasps] [singing] As you can hear, Vivo there didn't even follow the pitch that was specified. And how this works is actually quite interesting.
It reads the lyrics and the musical notes together and predicts how the performance should flow and it automatically decides the final timing and the audio length. This is especially useful when one syllable stretches across several notes. So under the hood, an auto reggressive component first builds a broad sketch of the singing style and timing and then a diffusion model fills in the finer details including pitch, vocal tone, articulation, and local audio texture.
Now, the awesome thing is they've released the model already. So, if you click on this view repository button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. Plus, they've also released the code on how to train this yourself. And that's because currently they've only trained the model on Chinese, but you can also train your own checkpoint with any language you want.
So, you just need to follow the training instructions on this page. Now, like I said, they have two different variants, a pro and a normal variant. Both of them are under 10 gigabytes in size, so you should be able to fit this on most consumer GPUs. Anyway, if you're looking to generate really realistic and expressive singing voices with AI, then this vocal render model is one of the best you can use so far. If you're interested in reading further, I'll link to this main page in the description below.
Also, this week, Tencent Hunen releases a really cool 3D editing model called Hunyan 3D Buffalo. This is basically a unified model that can generate, understand, edit, and separate 3D objects. So, here are some examples. We can simply edit 3D models with a text prompt. For example, we can input this model and then write turn the head into a bull's head. And here is the result. Or here we can write remove the sail. And it indeed removes the sail from this 3D model.
Or here's an example where we can add glasses to the frog like this. Or we can take this robot and put spiked gauntlets on both hands. And indeed, that is what it does. It's also able to generate 3D models from just a text prompt. So, here are some examples for your reference. The cool thing is this can also take in any 3D model and separate it into individual parts like this. Or here's another example where we can separate this model into these separate parts.
So, while most AI 3D model generators can only do one of these things, this new Hunyan 3D Buffalo is designed to be a unified model that can both generate and edit and segment 3D models, making it super flexible. The nice thing is at the top of the page here, it says the code is coming soon. So, it looks like they are planning to release the code and models to this, which is fantastic. For now, if you're interested in reading further, I'll link to this main page in the description below.
Next up, we have a new AI called Leap Talk. And this can generate talking avatars in real time. So, this just takes in any reference image of a person, any speech audio, and it outputs a lip-s synced talking head video in pretty much real time. >> Can the security and smooth passage of international waterways be fundamentally safeguarded? All parties should work together to deescalate the situation and prevent regional instability from having a greater impact on the global economy and energy security.
Now, this talking head does look very rigid. She doesn't move as naturally as some of the other frontier avatar generators out there, but the strength of this model is that it's incredibly fast. So, if you look at the latency of Leap Talk compared to other avatar generators like Hello 3 or Echo Mimic, which I featured on my channel before, Leap Talk is like thousands of times faster. And they claim that on an H200 GPU, this can achieve up to 200 frames per second, which is crazy.
So, if you're looking for a lightweight realtime talking head generator, this is likely the fastest option available right now. And the awesome thing is they've released this already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. If you're interested in reading further, I'll link to this main page in the description below.
If you want to supercharge your content creation, definitely check out Higsfield, the sponsor of this video. They've just added the best video generator out there, Seed Dance 2.5. The biggest improvement from this model is that you can now generate up to 30 seconds of video in a single pass with multiple shots and an actual narrative with audio built in. You can also extend an existing generation with new shots while keeping the same characters, locations, pacing, and overall consistency.
What really stands out is the reference system. You can feed it up to 50 references at once, including 30 images, 10 videos, and 10 audio files. You can provide your characters, environment, visual style, motion, and soundtrack all in one generation. It can even use a simple 3D clay render as a reference to understand the camera movement, lighting, and overall composition of the scene. You also get much more control over editing.
For example, you can specify exactly what happens during different timestamps, change just one section without affecting the rest of the video, or move the same performance into a completely different environment, or even change the camera angle while preserving the characters and action. Seance 2.5 supports text to video, image to video, video to video, and other references. So whether you want to create short films, action scenes, music videos, commercials or any other content, it gives you a lot more control while keeping everything consistent.
You can try Seance 2.5 on Higsfield today using the link in the description below. Now around 2 weeks ago, Moonshot AI released Kimmy K3 and this is a massive 2.8 a trillion parameter model which they open sourced and back then this was not only the best open- source model out there but it has even caught up to Frontier for some of these benchmarks it's as good or even better than like GPT 5.6 or Fable 5. Well, it's not over yet.
So this week Alibaba releases Quen 3.8 Max. This is also a massive 2.4 4 trillion parameter model and this is also the first time that they will open source a max class model which is super exciting. Here they say the weights will be released next week and check out the benchmark scores of this in terms of these agentic software engineering tasks. You can see that in some cases this new Quen 3.8 Max which is the dark blue even matches the performance of Fable 5 and GPT 5.6 Soul and it already beats Opus 4.8 which is the dark gray bar.
Very impressive for an open-source model. Now, like most Frontier models, this is designed to work across a ton of steps, autonomously use tools, inspect results, and continue working until it reaches your specified goal. Here in this demonstration, it was given an empty folder, and it was told to create a self-improving harness system from scratch that turns feedback into GitHub issues, automatically claims and codes them, and merges working changes, and keeps evolving the tool over time.
And the crazy thing is this worked for about 16 days without any human help. It just kept doing this until it achieved the goal and in the end it made like 265 commits and 127 pull requests. Or here's another example where it was given a recent paper for training language models and it was told to basically reproduce the paper and then beat it. So it worked autonomously for like 5 days. It wrote thousands of code from scratch, rebuilt the entire experiment, matched the paper's results, and then invented and tested 18 of its own improvement ideas.
And its final method actually beat the original paper by 2.7 points on a really hard competitive math benchmark. That's crazy. I mean, this is just a very recent paper, but you can just plug it into this AI to figure out an even better result. So, I mean, the age of autonomous scientific improvement is already here. Or here's another crazy example where it was told to design and optimize a chip. So it started from basically nothing.
It designed, coded, simulated, and repeatedly improved this cryptographic hardware accelerator. It was able to produce a physical layout that's like 12 times smaller than the baseline while meeting timing targets. Now, if you look at this leaderboard by artificial analysis, then you can see that Quinn 3.8 8 is just one point below Kimmy K3 while being like 400 billion parameters smaller. So I mean once they release the weights to this this will be like the second best open model out there and they're edging very close to GPT 5.6 and Claude Fable.
The thing I don't like about artificial analysis is they don't have any confidence intervals. So it's hard to say whether these top five models are actually significantly better in terms of performance. Now currently you can use Quinn 3.8 8 Max via their API. And as you can see, it's slightly more expensive than Kimmy, but still cheaper than GPG 51.6 and much cheaper than Claude Opus or Claude Fable, which I would not recommend.
Like I said, they are planning to release the weights next week. But for now, you can try this out via Quen Cloud. If you're interested in reading further, I'll link to this main page in the description below. Next up, Google DeepMind releases a really exciting update called Weather Next 2. This is basically an AI that can help predict hurricanes and tropical cyclones way earlier than other methods and way more accurately.
You see, the difficult thing about hurricane forecasting is you normally need one kind of model to predict where a storm will go and another highresolution model to predict how strong it'll become. But weather nex basically combines both jobs into a single model. It's able to predict the storm's track intensity and wind structure. In fact, it can generate forecasts as far as 15 days in advance, and it can run an ensemble of a thousand different possible scenarios to estimate the probability of where it's headed.
It's also way more accurate at predicting the track intensity and extent. So, this latest weather next model is the blue line here. And as you can see, its error rate is much lower than the other methods, even as you increase the lead time to like 5 days in advance. Here's another example where if you compare the error rates with other competitor cyclone prediction models, you can see that Weather Next has a much lower error rate.
In fact, Google says Weather Next provides more than 24 hours of additional forecasting lead time compared with leading systems. What's more impressive is it does it using weather data at roughly 28x 28 km resolution, which is around 100 times coarser than traditional models, which require much higher resolution data. So that means a full 15-day forecast can just be generated in under a minute on just one TPU, which is Google's tensor processing unit.
Here it says they basically trained this AI on nearly 20 terabytes of global atmospheric data, including nearly 5,000 historical storms. So the model is able to learn these complex weather patterns and how to identify extreme weather. The awesome thing is not only have they published a Nature paper on this, but they are also open sourcing the code and model to this. So anyone can just download and run this or build on top of it.
So if you click on this link, it takes you to their GitHub repo. And if you scroll down a bit here, it contains all the instructions on how to set this up. They're also releasing Weather Next 2 Mini, which is a compact version which you can run in Google Collab for free. So props to Google for open sourcing this. This AI is actually super helpful and it'll save a ton of lives by predicting, you know, these natural disasters earlier and more accurately.
If you're interested in reading further, I'll link to this main page in the description below. Also, this week, OpenAI announces something pretty remarkable. You see, there are rumors that their next model, GPT6, will be rolled out very soon. Well, this is internally codenamed Astra, and it looks like they used an internal version of Astra to solve 10 long-standing open math problems, which is pretty insane. These aren't just math questions with known answers.
They include open problems across geometry, coding theory, group theory, quantum complexity, cryptography, and combinotaurics, which no human could ever solve before. But here with this internal version of Astra, it was able to either resolve the problem or make substantial new progress. Now, this is very technical, especially if you don't have any math background. But in summary, one example establishes the existence of these non-sopic groups, which addresses a major question in group theory.
And then it was also able to resolve multiple airish problems. And then there were other examples where it's able to improve the bounds in sphere packing and coding theory. Now these problems are basically really hard to solve, but once it finds the answer, it's really easy to prove that the answer is correct, in case you're wondering. And you know, probably the craziest detail is the cost. So, OpenAI estimates that all of the model tokens used to discover these 10 solutions cost only $2,000 at their API rates.
That's crazy if you think about it. They only needed to spend $2,000 of compute to crack 10 mathematical breakthroughs which no human could ever solve. We are right in the middle of the scientific acceleration. In the past few months, we've already seen some huge breakthroughs in like medicine, physics, and of course, math. So, it's really exciting times. They've also released the reasoning walkthroughs on how the AI actually came up with the solution to each of these problems.
This is extremely technical math stuff, but if you are interested in digging through the details, I will link to this main page in the description below. Also, this week, Alibaba's Demo Academy released a really useful open-source model called Clinfusion. This is basically a model for holistic medical understanding. Basically, you can give it medical images like X-rays, scans, or even native 3D imaging together with a text prompt, and it can answer medical questions or generate clinical reports.
The big problem with medical AI models is that different imaging types are extremely different. But basically, what Clinfusion does is it uses a combined vision encoder designed to understand all this medical data inside just one system, including 2D and 3D medical data. And from this you can see the results are extremely strong. It beats leading and closed source models on most of these multimodal benchmarks even outperforming proprietary models like GPT 5.2.
Although note that this is quite an old GPT. Right now we're already at version 5.6. Now on to the specs of this. They released two different models. There's a 32 billion parameter model which is higher quality. And this is quite large at 72 GB in size. And then we have an 8 billion parameter version which is only 24 GB in size. So you should be able to fit this on like a mid to high-end GPU. This is probably the best open- source model for medical analysis.
You can use in that size range. So if you're interested in trying this out on this page, it contains all the instructions on how to download and run this locally on your computer. If you're interested in reading further, I'll link to this main page in the description below. In human robotics news, we have this new demo from Persona AI. So, here it's demonstrating its Gen 1 humanoid robot performing a real welding task through teleaoperation.
So, as you can see, the dude on the left is wearing this VR headset, and it's controlling this Gen 1 humanoid robot in real time. And it's even able to do this welding task, which requires very precise and stable movements. Now, this is quite a simple example, but this is a demonstration of how we can eventually deploy robots in some high-risk or dangerous industrial environments and then get them to work there while an expert controls them remotely somewhere else via tea operation.
In other robotics news, UB Robotics also previews their swarm intelligence model in action. Here you're seeing several of their new wheeled industrial humanoid robot called the Cruiser Y1. And they're all working in this warehouse. They're all taking stuff from pallets and then putting it to the right location. Now, not only do each of these have a brain, but there's also an overarching swarm intelligence system which coordinates all of them together.
So there's no redundancy, there's no overlap. And so this allows you to control like basically an army of robots to do a task concurrently. In other humanoid robot news, Xiaomi has released a new robot foundation model called Xiaomi Robotics 1, and it's designed to let robots handle everyday objects and practical tasks. So things like picking up objects and placing them in different places, as you can see in this example, it's even able to zip up this bag very effectively.
And it's able to pack a suitcase like this. It's able to navigate across the room to find various objects to put in the suitcase. Basically, how this works is you just give the robot an instruction using natural language and it looks at the environment through its cameras. It understands what needs to happen. It plans everything out and then it actually carries out the action. What makes this model especially interesting is how Xiaomi trained it.
So, normally collecting robot training data means having humans remotely control real robots for thousands of hours, which is expensive and difficult to scale. However, Xiaomi collected around a 100,000 hours of video using just a handheld gripper with a camera. So, the humans simply carried this device around while performing normal tasks in their homes or factories or offices. The model first learns general manipulation skills from all this human data.
Then, Xiaomi adapts it to actual robot bodies using another roughly 10,000 hours of real robot data. The awesome thing is they've actually released the models to this, including details of how they train this. So, if you're building your own robots or if you're looking for ways to train a robotics model, this would be a great resource for you. If you're interested in reading further, I'll link to this main page in the description below.
Also, this week, we have what they claim a self-evolving model called Big Bang. This is an experimental language model with a pretty wild idea. Instead of humans constantly designing better training questions to improve the AI, what if the AI could generate increasingly difficult training data for itself? So, Big Bang starts with the open- source Quinn 3.635b, but all its post-training data comes from an automated system where generator agents create and solve difficult scientific and technical problems.
And then afterwards, we also have this critic agent which tries to find mistakes and reject weak examples. And then finally, we also have a metacritic agent which checks whether these difficult problems actually make the model better on real research tasks. And the results are surprisingly strong compared to just the base model Quinn 3.635b. You can see that it scores much higher in terms of browse comp and bench which are like aentic coding benchmarks as well as Frontier Science.
You can see that this is an especially huge leap. The base model only scored like 12 points but here it scored like 46 points. Same with humanity's last exam, a massive improvement and also about mystery and also paper bench. Now calling this self-evolving is a bit misleading but I think what they meant here is this framework can be looped again and again. So you can keep getting it to create more and more challenging synthetic data to keep improving the AI model.
But of course at some point you're going to hit a wall in terms of intelligence and performance. The awesome thing is they've released this already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Meta quietly releases their latest AI model, Musepark 1.2.
This is their updated version that's now much more focused on real world coding and agentic workflows. The idea is to give it an entire software project and let it work across multiple files, call tools, and keep going through longer tasks. This supports a 1 million token context window, so you can potentially fit a ton of information into your prompt at once. It's also multimodal, so it can take in things like text, images, video, audio, and documents as input, although coding is one of its main focuses.
As you can see from their self-reported benchmarks, it is quite a big leap compared to the previous new Spark 1.1, and it's edging pretty close to the top models out there, including Opus 5. However, note that these self-reported benchmarks are very misleading. They just use GPT 5.6 Terra instead of Soul, which is the larger model. So, it looks like they're intentionally just including a dumber GPT model on their page.
So instead, if you look at this independent artificial analysis leaderboard, then you can see that Muse Spark 1.2 is actually all the way down here, way behind Quen 3.8, Kimik 3, and GPT 5.6 Soul Max. So it's not even within like the top 10. However, if you look at the cost per task, then this is quite cheap, even cheaper than Gemini 3.6 Flash and Kim K3 and much cheaper than the Opus models. Now, this is closed source, and for now, you can only use it via their API.
If you look at another leaderboard called Val's index, which looks at how good an AI model is across finance and coding tasks, then you can see that again, Muse Spark is all the way down here, even below GPT 5.6 Soul and Kimik 3. However, it does cost the least compared to the other models. So, this could be a fairly costefficient option. For now, note that you can only use Muse Spark 1.1 via their API and this is closed source.
In addition to Musepark, they also introduced something called Muse Code, which is a coding agent exactly like OpenAI's codeex or Zcode, Kimmyode, etc. Now, like I mentioned before, the best harness to use to run an AI model is the harness designed by the same company. So, if you're using GPT, then it's best to use Codeex. If you're using Kimmy, it's best to use Kimode. If you're using GLM, it's best to use Zcode. And the same goes for Muse Spark.
So, if you decide to use this, then the best way to use it is through Muse Code. Anyway, on this page, it contains all the documentation on how to download and run this. So, if you're interested, I'll link to this page in the description below. Also, this week, we have a very interesting harness framework called Long Horizon Harness. And this is designed to help AI agents complete really complex tasks that can take a ton of steps or a really long time, like hours or days.
You see, we already have a ton of agentic harnesses today like Codeex, which is now renamed into ChatgPT or also Claude Code, Zcode, Kimmy Code, Open Claw, or Hermes. But they all face one problem, which is that if you get it to do a really long task, then it has trouble fitting and remembering everything. So, as the history becomes longer and longer, the agent could forget the original goal. It could make some incorrect summaries and completely go off on a tangent.
It can repeat work or even mistakenly claim that something is finished. Well, what this long horizon does is it replaces that approach with three separate roles. There's a manager, an executor, and an auditor. The manager decides the next small task based on only verified progress. The executor receives that one task in a fresh context and performs the actual work, and then the auditor independently checks the files, the fixes, etc., and confirms what really changed.
Only stuff that was verified by the auditor are saved for the next round. So, think of this as like a construction project where the manager assigns the work, the builder completes it, and an independent inspector signs off before anything is marked finished. And the nice thing is this system works across different agentic harnesses like Claude Code, Codeex CLI, Gemini CLI, Zcode, Kimmy Code, and other compatible systems.
And here are some really impressive results. If they add this long horizon harness using Quen 3.7 in cloud code, you can see that it was able to increase its weavebench score by 28.9%. It's also able to triple the completion rate of OS World 2. And then for Terminal Bench, it's also able to increase the score by like 7.5% which is a huge deal. Here's the Terminal Bench leaderboard. And as you can see, if you add this LH harness to GPT 5.6 6 Luna using codeex, it's able to achieve a much higher score than without this harness.
It's all the way down here. Same with if you add the harness to Cloud Code and Quen 3.7, it's able to achieve a much higher score than without it, which is down here. Now, if you add this harness layer on top, of course, it is expected to use some more tokens. For example, for both Weavebench and OS World, it uses a lot more tokens than without it, but it does achieve a much higher score or success rate. So, there's a trade-off there.
Interestingly for Terminal Bench, not only did it achieve a higher score, but it also used fewer tokens. Anyway, a very fascinating framework which could potentially improve the performance of any long horizon tasks that you're running. The awesome thing is they've released this already. So, if you click on this code repo button and you scroll down a bit here, it contains all the instructions on how to download this for whichever Agentic system you're using.
If you're interested in reading further, I'll link to this main page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite? And which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content.
Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up to date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.