Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
7,706
Runtime
41:06
Speaking pace
187wpm
Reading time
32min
187 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
AI never sleeps and this week has been absolutely insane. OpenAI just got hacked with a simple trick using an image. We have a completely new type of AI model called Jev and not only is this incredibly fast, but apparently it has a 0% hallucination rate. GPT-6 Astra just cracked this encrypted German message from World War II that has remained unsolved for 80 years. We have a really powerful interactive world model which actually contains a game engine so it can keep track of everything. And this allows you to have multiple
94 words, the words spoken in the first 30 seconds at 187 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 433 |
| Average words per sentence | 17.8 |
| Longest sentence | 66 words |
| Questions asked | 9 |
| Sentences containing a number | 92 |
Most used terms
Filler phrases
109 in total: like 60 · basically 21 · actually 13 · kind of 8 · I mean 3 · you know 2 · literally 1 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
AI never sleeps and this week has been absolutely insane. OpenAI just got hacked with a simple trick using an image. We have a completely new type of AI model called Jev and not only is this incredibly fast, but apparently it has a 0% hallucination rate. GPT-6 Astra just cracked this encrypted German message from World War II that has remained unsolved for 80 years. We have a really powerful interactive world model which actually contains a game engine so it can keep track of everything.
And this allows you to have multiple players in the same environment. MiniMax open sources their harness MiniMax code. We have some super tiny open source AI models that are only megabytes in size. Google releases a framework for recursive self-improvement. We have some new models that allow robots to learn from human demonstrations on the fly without any retraining. Alibaba releases a ton of really cool Qwen models including an Omni model plus a real-time translator and voice cloner.
We have a ton of new open source models and a lot more. So let's jump right in. First up, we have a new AI called Meridian by VIGIL. Now this is built on MiniMax H3 and what this does is this can take an existing video and change the camera angle and timing of it. So you just need to input an existing video and Meridian can reconstruct a rough 3D representation of the scene and then regenerate the video from a different camera perspective or movement.
So for example, you can orbit around the subject or pan in different directions, change the distance, or even freeze a scene so you can get this bullet time effect. So here are some additional examples for your reference. Now underneath this Meridian uses a geometry system called VGGT Omega to basically estimate the depth of the scene. It uses this to reconstruct a 3D point cloud of the scene like this. And once it has this 3D reconstruction, then you can specify a completely new camera path, then the system can quickly render a rough version of what that scene would look like at this new camera perspective, and then plug it through MiniMax H3 to generate a full high-quality video.
So, if you want to like reshoot a video at a different angle, this method gives you ultimate control. Now, here they released everything already, but note that they fine-tuned it on the full MiniMax model, which is like 62 GB in size. Because this is open source, I'm sure there's going to be more quantized versions of this in the future, but if this current base model does fit for you, then down here it contains all the instructions on how to install and run this locally on your computer.
If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a new open-source real-time transcription model called R2D2. Oh, I mean R2T2. And this can basically turn live speech into text in real time. So, here's a demo of this in action. On the left is GPT Live Transcribe, and then on the right is R2T2, which is open source. >> T Transcribe and GPT Live Transcribe.
These models give you two different ways to build the transcription. GPT Transcribe takes a completed file and returns the full transcript. GPT Live Transcribe keeps a live open connection and returns text as the audio arrives. Both models support 57 languages, and they're much better at the parts of transcription that tend to break. >> So, as you can see, this is quite accurate at transcribing audio. It's the exact same transcript that you would get if you used GPT Live Transcribe.
And if you look at the word error of this R2T2 versus other transcription models, this has the lowest error rate. Plus, this also has the lowest latency. So, here's a chart showing error rate versus latency. Ideally, you want to be in the lower left corner, and as you can see, the R2T2 models in red are indeed the best in terms of both English and Chinese. Now, on this page, if you click on this GitHub link and you scroll down a bit here it contains all the instructions on how to download and run this locally on your computer.
Note that this is fairly tiny at only around 4 GB in size. So this should be able to fit on most consumer GPUs. If you're interested in reading further I'll link to this main page in the description below. Also this week we have a new generative world simulation called Jing and Dao by XGen Labs. You see most generative world models are basically video generators that only predict what the next frame should look like.
But here they're trying to sue simulate an actual persistent world behind those frames. So think of Dao as like the world engine. It stores the shared state of the environment, applies rules and updates what happens when someone takes an action. And it also lets autonomous agents continue doing things even when you're not looking at them. So things can continue moving. You can have NPCs and other moving parts. And then Jing is like the experience model.
It takes the underlying world and generates what a particular character should see and experience from their own point of view. So think of Dao as like the game engine and Jing as like the camera looking into that simulation. The key idea here is persistence. So if you leave an area, well this Dao model is able to keep track of everything. So when you go back to that area everything remains consistent. Plus the awesome thing is this is open-ended which means you can continue generating a world for an unlimited duration.
Of course subject to your compute constraints. Whereas other world models can only generate a few seconds or at most a minute. In fact if you compare these XGen models against other competitors like VO3 or Genie 3, you can see that this has the most features. And this is the only one that has a world state management system. In other words it's kind of like a game engine which is the Dao model. So it can keep track of everything including the environment and the movements or interactions of other characters or objects in the simulation.
And because we have this down model that keeps track of everything, now you can have multiple players all in the same environment. So, for example, here we have the perspective of three different characters all in the same environment, and notice how everything remains consistent. Now, currently this is just a research preview, but they have released the code to the Jenga model. So, if you click on this GitHub link, and you scroll down a bit, here it contains all the instructions on how to download and run this locally on your computer.
Note that for the video generation, this is based off of the best video model out there, MiniMax H3. It's quite large. They basically fine-tuned it on the full MiniMax model, so this is like 67 GB in size. Hopefully, there will be more quantized versions of this in the future. If you're interested in reading further, I'll link to this main page in the description below. Also this week, Google released a really interesting project called Dream RSI, Recursive Self-Improvement Through Evolving Worlds.
Now, this probably isn't what you're thinking. This isn't like a self-improving AI that just gets better and better by itself. Instead, this is a really interesting experiment that gets an AI agent to improve the way it searches for solutions without normally changing the underlying AI model at all. So, normally, if you want an AI to get better, you might retrain the model or use reinforcement learning. But, what Dream RSI does instead is it lets the agent learn from the entire history of its own previous searches.
Every time it explores a problem, it creates what the researchers call a discovery tree, which looks like this. Think of it as like a giant map showing everything it tried before, which branches worked, which ones failed, and what result it got. And once this map exists, then the system can replay thousands of these search trees without actually running them again. Basically, the AI is kind of dreaming about what would happen if it had pursued different branches.
It generates improved versions of how it would search through these trees, and then it tests them against all other existing records, and then it deploys whichever strategy works best and then the cycle repeats. You keep looping this again, it generates another discovery tree, etc. etc. If you look at these benchmarks on like mathematical optimization compared with other recursive improving systems, then you can see that Dream RSI on average performs the best.
In terms of GPU kernel engineering, you can see that Dream RSI is able to achieve comparable results while using over two times fewer generations compared to other methods. So, that's a really condensed summary of what this Dream RSI is. It's not like the underlying Gemini model magically rewriting itself, right? It's model weight and its architecture stays fixed, but what improves is the strategy surrounding the model.
How it explores, how it allocates compute, and how it decides what is worth pursuing. Now, at the top here, they've released a full technical paper on this. So, if you're interested in digging deeper, I will link to this main page in the description below. Also this week, Google releases their latest real-time AI voice model, Gemini 3.8 Live and 3.8 Live Extended Thinking. Here, they're trying to make AI voice sound less like a robotic voice interface and more like an actual agent which you can continuously talk to and work with.
So, the base Gemini 3.8 Live model is the faster and cheaper model designed for large-scale deployment, whereas this extended thinking model is a more powerful version for tasks that require deeper multi-step reasoning. Both can handle live audio and also visual stuff like images or videos. It can also call different tools like web search, all at the same time. In fact, one of the biggest changes is that the model doesn't have to stop speaking when it needs to do something in the background.
It can call tools or APIs or do other stuff while also continue talking with you and then incorporate the results from all these tool calls when they're ready. So, it's basically able to reason and do stuff and speak at the same time. And this enables a much more fluid and natural conversation. Another cool thing is instead of you needing to manually specify the language, this model can automatically detect and transition between 97 supported languages.
This also supports live visual inputs, so you can point a camera at something and ask the model things in almost real time. And if you look at this benchmark comparing other live voices like GPT live Astra and also Grok voice, then you can see that the new Gemini 3.8 live models have similar performance. But here's why it's such a big deal. If you look at the cost of this, then Gemini 3.8 live is many times cheaper than Grok voice or GPT live, making it the most cost-efficient frontier live voice model we have right now.
The awesome thing is they've already released this, so it's available in the Gemini and Google AI Studio, and it's also available for everyone in search live as well as the Gemini app. A very nice release from Google. If you're interested in reading further, I'll link to this main page in the description below. Also this week, Alibaba releases Qwen 3.8 OmniFlash. This is basically a multimodal model that not only takes in text prompts, but you can also feed it audio, images, and video for it to understand.
So it's really great for working with multimedia. For example, you can get it to analyze and identify segments in videos or create transcripts, and it can easily handle this. And because this is a flash model, this is also incredibly fast. You can even stream this in real time. For example, you can just turn on your camera and then ask it things about the scene, and as you can see, it's able to understand and render everything in real time.
Or here, you can upload a video and get it to break it down into the characters, camera movement, composition, dialogue, sound, etc. So this is a super flexible tool. And this has a massive 1 million token context window, so you can input roughly an hour of video or over 10 hours of audio. And And you look at these benchmarks measuring its multimodal capabilities, it's also incredibly impressive. So, the dark gray bar is Gemini 3.8 flash.
For most of these instances in terms of audiovisual understanding and interaction, you can see that it even beats Gemini 3.8 flash. Here are more benchmarks in terms of audio reasoning and speech recognition. Again, for most of these, it beats Gemini 3.8 flash. It's also insanely cheap. So, if you use this through their API, then the cost per hour of video or audio input is like than Gemini flash. Now, currently, this Qwen 3.8 Omni flash is available via the Qwen 1 AI platform, where you'll need to subscribe to a token plan.
This is one of the cheapest and most efficient omnimodal models out there. So, if you need an AI to work with a ton of audio, images, or video files, this might be one of the best options to use. Hopefully, they will also open source this in the future. For now, if you're interested in reading further, I'll link to this main page in the description below. And that's not the only Qwen that was released this week. So, they also released Qwen 3.8 live translate.
This is basically a real-time AI interpreter designed for actual conversations. You can feed it live audio or even video, and it can translate the conversation into a transcript while people are still talking. The cool thing is this can also do voice cloning. So, in addition to the transcript, this can also make people speak out a different language with the same voice. >> Wukong, aren't you able to summon wind and rain?
Please bring them a rain of sweet rain. Save the people of this region. >> Oh, you can make it rain? I've been praying for so many days, BUT IT HASN'T RAINED, and you think you can? >> It's able to separate different speakers and label who said what, as you can see here. Another cool feature about this is you can even feed it video or images to use as additional info. Now, this supports 60 input languages and can output 29 spoken languages.
And as you can see from these tests on faithfulness, fluency, and conciseness, then on average, this new Quan model has a much higher win rate than other competitors like Seed Live Interpret, GPT Real-time, and Gemini Live. The same goes for translation quality. This new Quan model scores the highest, and it's also very fast and accurate. So, in terms of latency and error rate, it's lowest among the models. Now, currently, they've only released this via API.
Plus, if you click on this button, you can try their online demo. If you're interested in reading further, I'll link to this main page in the description below. If you create videos, ads, or pretty much any kind of visual content, definitely check out Runway, the sponsor of this video. Think of it as an all-in-one creative AI platform where you can carry out all your creative workflows. You can access the best image generators like GPT image and Nano Banana, as well as the best video generators like SeaDance and Kling, all through a single platform.
But, what makes Runway especially interesting is that it goes way beyond just generating individual clips. For example, with Runway Agent, you can simply describe the video you want to make, and it can help develop the concept, create the scenes, generate dialogue and voice-over, add music, and turn everything into a complete video through one conversation. And if you already have footage, All of 2.0 lets you edit it just by describing what you want changed.
You can relight a scene, completely restyle it, change the environment, or add and remove objects without having to reshoot anything. You can even build reusable workflows that chain multiple AI tools together. So, once you figured out a process you like, you can automate it and generate consistent content at scale. So, whether you're making ads, social content, product videos, films, or entire marketing campaigns, Runway gives you the models, editing tools, and automation you need to go from an idea all the way to something you can actually publish.
Check out Runway using the link in the description below or by scanning the QR code and use my code for 50% off the first month on paid plans. Also this week we have a super tiny AI model called Needle 3 that's designed to run directly on tiny devices like phones, watches, or even microcontrollers. The crazy thing is this entire AI model is only 8 to 29 megabytes. Keep in mind even like the smallest language models out there are currently like dozens of gigabytes in size.
So this is incredibly tiny. And this is built on their simple attention network which is kind of different from transformer models. Now the full model ranges from 29 million to 121 million parameters, but one set of weights contains multiple smaller models inside it. Basically this model can turn on only two layers or up to 20 layers depending on how much compute your device has. So they call this intelligence laddering.
A tiny microcontroller might run just a few layers while a more powerful phone or a Raspberry Pi could use more layers without needing completely separate models for every device. And the efficiency here is crazy. Here they say that Needle 3 beats models that are 10 times its size on mobile tool calling benchmarks. Now as a super tiny model this ain't meant to be frontier intelligence, but this can do simple things like controlling small devices.
For example you can use this for smart home devices and prompt it to turn on or off lights. Or you can also add this to household robots and of course your smartphone, wearables, AR glasses, etc. So if you're interested in setting this up at the bottom of this page it contains all the instructions on how to download and run this on your device. And it supports all these devices and operating systems. If you're interested in reading further I'll link to this main page in the description below.
Also this week ZAI published a really interesting blog. So here's a really cool example of how they used AI to improve the infrastructure that actually runs their models. It's quite long and technical, but basically when they needed to deploy their latest GLM 5.3 flash, they were working with a cluster of over 100,000 Chinese-made accelerators, I suspect from Huawei. Now, GLM 5.3 flash is quite a new architecture with multimodal capabilities and around a million token context window.
So, it's not easy to deploy something like this on a cluster of accelerators at this scale. Normally, getting a completely new inference system to run efficiently on like hundreds of thousands of GPUs would require a team of specialized engineers spending weeks digging through kernels, memory bottlenecks, communication problems, and a ton of different bugs. But instead, for Z AI, a huge amount of this optimization work was just done by an AI infrastructure agent powered by GLM 5.3.
Basically, engineers gave the agent access to the code along with extremely detailed feedback from tests, logs, execution traces, benchmarks, profiling tools, etc. The trick is it has to be extremely detailed. So, instead of telling the AI the system got slower, they needed to show exactly where time was being spent, which calculation was producing the wrong results, and under what conditions a particular optimization helped or hurt.
They called this approach dense feedback, and this allowed the AI agent to operate more like a scientist. It would repeatedly form hypotheses, change the code, test the result, and then decide what to try next. One example involved a bug that became worse with extremely long contexts, but the agent was able to trace it down to lower precision math inside a kernel and help fix it. They're also able to get this agent to optimize this decode kernel to speed it up by 1.71 times.
There's a ton of other technical fixes that they documented here. The most ridiculous thing is they were able to optimize this and deploy the model into these hundreds of thousands of GPUs for users to use in less than 2 weeks. Now, it's important to note that this isn't like fully autonomous or recursive self-improvement. Humans still need to choose the goals and design the testing environment, and also define safety boundaries, and also give really detailed feedback so the infra agent can accurately pinpoint underlying issues.
But, it's just incredibly impressive that they were able to do all of this in under 2 weeks. And because they've already built this loop now, this can also be used to serve and deploy future AI system. So, it's basically a flywheel that keeps getting better and compounding over time. Anyway, this is quite a technical read, but props to the Z AI team for even releasing this information. I mean, this infrastructure stuff is usually top secret information for the closed labs out there.
If you're interested in reading further, I'll link to this main page in the description below. So, that was Z AI, but here's an even cooler insight from Xiaomi. So, they are doing something that we almost never get to see from a frontier AI lab. They're actually publicly streaming the reinforcement learning process of their upcoming model Mimo V2.6 while the models are still being trained. So, if you go to this page, which I'll link to in the description below, you can actually see the model being trained in real time.
And the scale of this run is enormous. This is pretty crazy. Currently, the cost is like over $1.6 million. The scale is massive. So, each training step processes almost 3 billion tokens. And currently, at the time of this recording, they have processed almost 48 billion tokens for their latest Pro model. And you can see the training batch size and the rollouts over here. You can see in real time how the model's performance improves on various benchmarks like Deep Sweep across time as the model undergoes reinforcement learning.
For example, when it first started, the Pro model only scored 58 for Deep Sweep. Right now, it's at 67, which is edging very close to the top open models out there. Same with Automation Bench. You can see the performance gradually improving over each training round. You can see the different types of prompts that were used per training run. So, it's broken down into coding, general, cybersecurity, etc. So, a super insightful dashboard that's actually showing the training results of their latest model in real-time.
You never get to see this type of stuff with the closed frontier models. If you're interested in digging deeper, I'll link to this page in the description below. Also this week, we have a really interesting project called in-context robot learning, and they released a framework called GPT policy. The idea here is to get a robot to learn a task just from watching a human do it without retraining the robot model. So, it's basically in-context learning for robots.
So, you can give it a video of a human doing a task or even a robot doing a task, or you can give it an image showing what the finished task should look like, or even give it feedback from previous failed attempts, and you can feed this information into a general-purpose vision-language model such as GPT-6 Astra or basically any model with vision capabilities. And with this framework, it'll get the robot model to attempt to learn and demonstrate this task.
So, here are some real-world examples. Without watching the human video, the robot model was not able to pick up this red towel. But after watching a human video, it was able to succeed in two out of three attempts. Or here's another example where without watching any videos on how to unscrew a bottle cap, it fails 100% of its attempts. But after watching a video on how to unscrew a bottle cap, it's able to succeed two out of three times.
And this is really important because with this framework, you don't need to retrain the robot model on a new task, which would require a ton of time, compute, and cost. Instead, you just need to show the robot a video demonstration of the task, and it can kind of learn on the fly or in context. At the top of the page, they've released the code to this. So, if you click on this code button, and you scroll down a bit, here it contains all the instructions on how you can download and run on framework yourself.
If you're interested in reading further, I'll link to this main page in the description below. Also this week, GPT-6 Astra just helped crack an encrypted German Enigma message from World War II that had apparently remained unsolved for more than 80 years. So here's what the original message looks like. This message was sent by the German army during World War II and it was encrypted using this Enigma machine. This is the famous encryption system used by the Nazis.
The encrypted message was only 82 characters long, but nobody had confirmed a solution for it. Well, this team decided to give this message to GPT-6 Astra and the impressive part is that they I didn't just guess the answer. It basically carried out an entire cryptoanalysis investigation. Astra searched through historical archives and compared different transcriptions where some of the original letters were unclear, looked for clues in related messages, and built its own Enigma simulator, as you can see over here.
It wrote code to search through possible machine settings, ran experiments in parallel, and after roughly 10 hours of work, it was able to decrypt the message. Now the plain text was pretty ordinary. The sender was essentially saying they are in rows now, asking for instructions on which route to go to, and requesting an immediate reply by radio. Now this wasn't completely autonomous, so the author chose the target and supplied crucial information.
The author had to also guide the overall investigation, but Astra handled a large amount of the actual technical work. Now they ain't claiming that they broke this Enigma encryption for the first time. We figured out how Enigma works for a long time already, but it's about recovering this one historical message whose correct key has been lost for 85 years. So no one was able to decrypt this exact message. Anyway, a pretty interesting read if you're into historical research and cryptoanalysis.
If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a completely new AI model called Jev, and this was founded by a former OpenAI researcher who helped with the reinforcement learning methods behind ChatGPT. But instead of another language model, he built a completely different model called Jev, and here's how it works. You see, they call this a system one model.
Now system one thinking is like fast, intuitive, and instant thinking as opposed to system two thinking, which is more like slow, step-by-step reasoning, which is what you see with the frontier large language models today. So this is meant to make decisions fast and pretty much instantly. So normally if you want to get GPT to make a decision, well, it can take a while to think through everything and generate its answer one word at a time, which of course takes time and compute.
And there's a possibility that it might hallucinate. Well, Jev skips all of that. In fact, you need to define the possible answers ahead of time. And what Jev does is it just gives you the probabilities or the confidence scores for each of these answers based on your prompt. So for example, if you give it an email and you ask it which department should this be sent to, it would output its confidence score for each department almost instantly.
And of course, the one with the highest confidence is the answer. In fact, they say that it typically responds in around 70 to 500 milliseconds. And because it can answer multiple questions about the same information in parallel, one call can actually potentially make many decisions at once. Now under the hood, they used a completely training method which they called reinforcement learning for calibrated decisions. And the goal here is to make Jev's output, in other words, its confident numbers actually mean something.
So when it gives you an 80% probability of something, you would expect it to be correct 80% of the time. And this lets programs do something really useful, which is you can get it to automatically accept decisions when Jev is highly confident, like over 95%. But for like lower probability cases you might need to get a human or a frontier model to review further. So this is very different. This isn't replacing chat GPT.
You can't really get it to write an essay or do some deep research. Instead, this is designed to function in software to make fast decisions about structured data. For example, like which option is correct or what should happen next. And if Jev can make these decisions really accurately while being super fast and cheap, then this kind of model can become a really useful middle layer sitting between AI agents and everyday software.
And the cool thing is it turns out that, you know, these structured decisions don't just occur at work. It actually occurs throughout life. For example, there's only a limited set of decisions you can make at any point in time. You can choose to, for example, move up, down, left, right, or shoot or jump. So this is also just like a set of multiple choice options that you could take next. So it turns out that Jev is also pretty good at playing video games.
For example, here's a demonstration of it playing Doom. Now here are some benchmarks for its particular use case. You can see how Jev is extremely cheap while achieving the same accuracy as the other frontier models. Now probably the most important chart about Jev is this. They claim that Jev has 0% hallucination rate. And that's because if you give it a list of multiple choice answers in advance, then it has zero chance of making an answer up.
That's what they mean by 0% hallucination. But it doesn't mean it's accurate 100% of the time. So sometimes its confidence scores for the answers could still be wrong. Now Jev is not open source. You can only use it through their API. In fact, right now you'll need to click on this join waitlist button in the top corner to get access. If you're interested in learning more, I'll link to this main page in the description below.
Now Jev is not open source, so I was looking for open source variations of this and I stumbled upon a very similar model called Leia. In fact, here the author claimed that they worked on this literally one year back and then they published a second paper in September of last year and they claimed that Jev basically proposed the exact same concept. So, maybe this Leia model is actually the original Jev, but it just kind of flew under everyone's radar last year.
So, this is also a system one model where you can give it a list of multiple choice answers for it to give confidence scores. And if you look at their self-reported benchmarks, then you can see that Leia even performs slightly better than Jev. It's also a lot faster and it supports multiple languages and it's open source under the Apache 2 license. So, here they released three different models of this. There's the base model and then there's a multilingual model and then there's another model with a higher context window, which is specialized for these use cases.
If you check out the main Leia model, this is quite small at only 2.37 GB in size, so this should be able to fit on most consumer devices. And then at the top here, if you click on this GitHub link, it contains all the instructions on how to install and run this on your computer. They even provide a script on how you can fine-tune Leia on your own specific data, which is fantastic. If you're interested in reading further, I'll link to this main page in the description below.
And by the way, that's not the only open-source Jev. So, here's another one called Bespoke Nimble. They basically tried to reverse engineer how Jev works and they built a very similar model where you can just give it a text prompt plus a set of options and this model would give you its confidence score for each option. So, what they did was they just took Qwen 3.5 9B as the base model and they trained a LoRA adapter on top of it to deal with such structured decisions.
And if you look at the results of this, then it actually performs pretty closely to the Jev model. Note that this is a LoRA adapter for Qwen 3.5 9B, so you also need to download this base Qwen model, which is roughly 20 GB in size. I'm sure there's also more quantized versions of this. And then this Bespoke Nimble adapter is only 193 megabytes in size. So, this should be able to fit on like mid to high-end GPUs. And then at the bottom here, it contains all the instructions on how to download and run this.
So, if you want to try out an open-source Jev, this might be a good option. If you're interested in reading further, I'll link to this main page in the description below. Also this week, OpenAI introduces Astra for law. This basically takes GPT-6 Astra and configures it specifically for professional legal work. Instead of relying on a general web search, it has a dedicated legal search index covering things like US case law, regulations, court rules, and administrative decisions.
So, a lawyer can just give it facts of a case and the system can research relevant authorities, explain arguments on both sides, draft evidence, or examine how a particular contract clause changes the risk of a deal. They're also connecting it to the tools that law firms already use through plugins and integrations. Now, if you're working with legal stuff, then of course this involves confidential client information.
So, here they claim that they're offering additional controls including zero data retention for eligible customers. So, a very useful application if you're in the legal space. If you're interested in reading further, I'll link to this main page in the description below. Also this week, here's a fun story. So, apparently just three of these security researchers managed to break into OpenAI's internal systems in less than 72 hours.
And one of the most interesting parts is that they used Anthropic's Claude to help them do it. So, the researchers are from a security company called Hacktron and the attack started in a surprisingly simple space, the OpenAI community forum. You see, they discovered a vulnerability in libheif, which is a library to process certain types of images. And specifically, when someone uploads an image to this forum, it eventually got passed into this decoder.
And it turns out that if you create a really specially crafted image, which looks like this, the researchers were able to trigger a memory bug in the pipeline and get remote code execution on the forum server. But this alone didn't give them access to OpenAI's most important stuff. The much bigger problem was a second vulnerability in OpenAI's single sign-on system. By combining the previous flaw with this identity flaw, the researchers say that they could take over ChatGPT and Codex accounts belonging to everyone who logged into the forum, including OpenAI employees.
This is where the potential impact became much larger. You see, these employee accounts can be connected to services like GitHub, Slack, email, and Google Drive. So, compromising one employee account could potentially provide access far beyond just ChatGPT itself. And for your information, here they say it cost less than $3,000 in tokens using Claude to find the vulnerabilities. Now, they reported this to OpenAI, OpenAI patched the issue, and they awarded them a $6,500 bounty.
Like, seriously? This was all the money they got? I mean, this vulnerability could be worth millions. And it's important to note that it wasn't completely autonomous hacking. So, these experienced researchers were still guiding the process, but the AI was able to accelerate their work that traditionally required really specialized expertise. So, I think the bigger story here is that you can just get an AI model to hack a frontier AI lab in less than 72 hours, and need to burn like a few thousand dollars of tokens.
So, it makes you wonder, how many security holes are there in everyday apps and software that still haven't been found yet? And how widespread are they? Anyway, if you're interested in reading further, I'll link to this main page in the description below. Also this week, MiniMax has open-sourced their coding harness called MiniMax Code CLI. This is essentially the agentic harness around a coding model. If you're not familiar with the term, basically the underlying AI model provides the intelligence and the thinking, but you can also put it in a harness which controls how it works.
For example, how it reads files or calls tools, manages permissions, and keeps track of progress, and decides what to do next. In fact, designing a good harness is just as important as making a good foundation model. Here you can see that for this frontier honesty evaluation, if you pair Minimax code with Kimiko 3, not only does it outperform other harnesses, but it also finishes the task way faster. And then here's another chart showing the cost of this.
Again, this was able to complete a lot more tasks than the other harnesses, plus at a much cheaper cost. So, they've released a GitHub on this, and down here it contains all the instructions on how to download and run this locally on your computer. Keep in mind, this is just a harness, so you can pair this with any model you want. If you're interested in reading further, I'll link to this main page in the description below.
All right, next up we have a ton of new open-source models. I'm just going to rapid-fire through them. The first one is Bonsai 2 27B. This is an insanely compressed model. So, basically what they did was they took Qwen 3.8 27B, and they compressed it down like crazy, so you can even fit this with much lower memory. You see, normally AI models need to work with really long decimal numbers, but here what ternary Bonsai does is it uses ternary weights, which means that each number that the AI has to work with can only be -1, 0, or +1.
And this allows them to give it a dramatically smaller memory footprint and make it way more efficient. So, for your reference, the original Qwen 3.87B is 55.6 GB in size, but here with this new ternary Bonsai method, they were able to reduce the total model footprint to only 5.9 GB, which is more than nine times smaller than the full model. Now, the awesome thing is if you compare this across various benchmarks, its performance is still very similar to the original full model, even though it's over like nine times smaller.
So, this is a huge deal. You can now fit one of the best medium-sized models, Qwen1.5-27B, on even just a mid-consumer GPU. Now, on the right sidebar here, if you click on the Hugging Face link, they released different versions of this ternary bonsai 2. There's an MLX version for Apple devices, or I'm going to click on this one, and note that they released different models of various compression sizes. The Q1 is only 5.95 GB.
You can have a slightly higher quality Q2, which is still pretty small at only 7.21 GB. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have another very efficient medium-sized model called Okami 1.0. If you look at its performance compared with other similar medium-sized models, then you can see across all these agentic benchmarks, it is on average the best.
In fact, for some instances, it's even able to compete with much larger models like GPT-5.6 Soul or Qwen1.5-27B Max and DeepSeek-V4-Pro. And because this is fairly tiny, this is extremely efficient. This has like the best intelligence-to-cost ratio. And for your information, this is actually a post-train of Qwen1.5-35B. So, it has the same architecture, it has 35 billion parameters, but when you use it, only three of them are active.
And not only have they released the model weights to this, but they also released the training recipe as well. So, this is fully open source. They also released a ton of GGUFs for this, and like the smallest one is only around 12 GB in size. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have another model called ZGCM-1. And this is a really small model at only 7 billion parameters, and this is designed for mathematical reasoning and agentic search.
Across various agentic and reasoning benchmarks, it does perform quite well. Although note that they are comparing this with older generations of models, so they're kind of cherry-picking here. In terms of its performance across competitive math benchmarks, then as you can see, on average this does rank very well. So, if you need a fairly small model that can run on a consumer GPU, which is specialized for mathematical reasoning, then this might be a decent open model to use.
In fact, not only have they released the models to this, but also the full training pipeline, which is fantastic. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a new robot model called Odyssey 3. And this is quite a big deal. You see, they are attempting to build just one foundation robot model that can understand completely different machines, including like robot arms, humanoid robots, autonomous cars, drones, or even video game agents.
The key idea is instead of training a completely separate model for every different robot or every different task, Odyssey 3 first learns a broad model of how the world works with a ton of visual data. In other words, it tries to understand things like motion, physics, cause and effect, objects, and other behaviors before it learns how to control a specific robot. So, basically this huge Odyssey 3 world model provides an understanding of what's happening and just general knowledge of movement and physics of the world, and then this action decoder translates this understanding into actual motor commands that can move all these different types of robots.
And as you can see from these demos, it's pretty impressive. This can also control humanoid robots and get them to manipulate different objects. Then here you can see this model being able to operate drones. They also tested this with autonomous driving. So, you can place the same model into an EV, and with only 20 hours of simulated driving data, Odyssey 3 could learn how to drive through the streets of India. And instead of real robots, you can even deploy this model to a world simulator and get it to control, you know, any type of character or robot like this.
So, a pretty fascinating model that can generalize across a ton of different devices, robots, and other environments. If you're interested in reading further, I'll link to this main page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite and which tool are you most looking forward to trying out?
As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up-to-date with all that's going on in AI, be sure to subscribe to my free weekly newsletter.
The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.