Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Search · @theAIsearch
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Search's most watched videos.
Most replayed moment at 10:28
3.1x that video's typical replay level
and then fix any errors that it sees. And so afterwards, it rendered the video successfully and that's pretty much it. In just one prompt, I didn't even need to prompt it further. Here's our final result. >> Four companies, one quarter, and a half-trillion-dollar bet on artificial
Said at 10:22
Most replayed moment at 15:21
4.1x that video's typical replay level
and deep agent for only $10 a month. This is way cheaper than if you paid for each tool separately. Definitely check out chat.llm that comes with deep agent in the description below. You can think of the residual connections not as a simple pipe carrying the signal forward,
Said at 15:15
Most replayed moment at 12:41
6.0x that video's typical replay level
other content, Higgs Field is a game changer that will supercharge your production workflow. Try it today using the link in the description below. Now, if we dive deeper, here's how it works in technical terms. They used something called a Markov head. In probability theory, a Markov process assumes that
Said at 12:33
The graph counts replays. It does not show where viewers stopped watching.
Words
5,814
Runtime
31:00
Speaking pace
188wpm
Reading time
24min
188 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Deepseek did it again. They've released a new model and not only is this among the frontier models out there, but it's also the most optimized, efficient, and frictionless AI model we've seen so far. As always, not only have they open- sourced the model, but they've also released a technical paper on this. And how they designed this is just really unexpected and sometimes even seems wrong. But once you understand the logic behind everything, then it suddenly becomes absolutely brilliant. In this video, we're going to do a deep dive into its
94 words, the words spoken in the first 30 seconds at 188 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 365 |
| Average words per sentence | 15.9 |
| Longest sentence | 45 words |
| Questions asked | 17 |
| Sentences containing a number | 24 |
Most used terms
Filler phrases
85 in total: like 39 · basically 25 · actually 7 · kind of 7 · right? 4 · I mean 1 · literally 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Deepseek did it again. They've released a new model and not only is this among the frontier models out there, but it's also the most optimized, efficient, and frictionless AI model we've seen so far. As always, not only have they open- sourced the model, but they've also released a technical paper on this. And how they designed this is just really unexpected and sometimes even seems wrong. But once you understand the logic behind everything, then it suddenly becomes absolutely brilliant.
In this video, we're going to do a deep dive into its design so that you can see how genius and cracked this is. Now, this is a super technical paper, but as always, I'm going to break it down into simple terms so that anyone can understand. Let's jump right in. Let's first set the stage by reviewing the situation that Deep Seek is in. Keep in mind that this is just a small Chinese lab that does not have nearly as much funding as OpenAI.
Their team is like dozens of times smaller. Plus, they don't even have access to the best Nvidia GPUs out there, nor do they have a massive data center. In fact, they are severely constrained in terms of compute and resources. Yet, they just released their latest model, Deepseek V4.1 Flash, which even matches the performance of Frontier models, even though this is just a Flash model. Plus, it's way faster and more efficient.
And get this, its memory footprint is like over 400 times smaller compared to the first generation. How on earth did they pull this off? Well, to understand the brilliance of their solution, we first need to go over the basics of what actually happens when you use an AI model. It's actually broken down into two different phases. The first phase is called prefill. This is basically when the AI reads your prompt along with any other information or documents you give it so it can understand the context of everything before it starts generating an answer.
Specifically, your text is broken down into smaller pieces called tokens which are then passed through the AI model's layers. In fact, each layer in the model calculates two important sets of numbers called keys and values for each token. These are stored in something called the KV cache. You can think of this KV cache as like the notes the model takes as it reads through your information. Now the second phase is the decode phase or the writing phase.
This is where the AI starts generating your answer. Now interestingly large language models generate its answer one word or token at a time. In order to do so, the model needs to look at everything that came before it to figure out the next most probable word that should come next. And to do that efficiently, it needs to refer back to its notes that it took during the reading phase. In other words, it needs to look at all the KV cache that it has created.
This KV cache lets the model quickly access the relevant information without having to recalculate the entire conversation from scratch. To understand this a bit better, here's a nice analogy. Think of an AI model like a student watching a really long lecture and then taking notes along the way. Later, when the student needs to answer a question about the lecture, he doesn't need to replay the entire lecture from the beginning.
Instead, he can just refer back to his notes. Well, you can think of the KV cache like these notes. It gives the AI a quick way to refer back to information it has already processed without having to recalculate everything from scratch. Now, here's the problem the industry is facing right now. You see, for a short prompt like this, everything works fine. The AI model can easily convert this into a fairly small KV cache that fits well within its active memory.
But here's the thing, the industry is now focused on making AI handle really long and complex tasks. We want to get AI agents to work autonomously for hours or even days. We want to give them a ton of different documents or a huge codebase to keep track of and keep working on for a really long time. And in these scenarios, the KV cache or the notes that the AI has to take is going to be massive. So going back to our analogy now, instead of just watching one lecture, the student has to watch weeks and weeks of online lectures and take notes on all of them.
His notes will start to pile up fast to the point where they can't even fit on his desk anymore. They're going to fill up his entire room and eventually he needs to like shove his notes into filing cabinets in another room down the hallway. And then every time the student needs to answer a question, he needs to dig through all these huge piles of notes. Well, this analogy is exactly what happens to an AI when it needs to work with a ton of information.
You see, inside a GPU, you have this high bandwidth memory or HBM. This is the incredibly fast memory that's right next to the processor chip. It's really fast and easy to access. So, it's kind of like the surface of a student's desk in our analogy. any notes that are sitting on this desk are incredibly quick and easy to access. But the desk space is limited. Well, high bandwidth memory is exactly the same. There's limited space and it's also extremely expensive.
So when the KV cache gets too big, this high bandwidth memory gets full. So the computer has to start storing this data somewhere else. For example, on solidstate drives or SSDs which are outside the GPU. So, in our analogy, it's like putting some of the students notes into filing cabinets in another room down the hallway. You have much more storage space there, but there's a trade-off. He needs to go to the other room to grab the notes from a cabinet, which takes much longer than if it was just sitting in front of him on his desk.
So, similarly, SSDs, or these solidstate drives outside the GPU, have way more storage capacity, but they're located further away from the GPU, and the latency there is devastating. If the KV cache needs to be stored there, well, every time the AI needs to predict the next word, it has to fetch data from the SSD, pull it through the motherboard up into its high bandwidth memory, and then into the GPU processor. This data transfer speed becomes the absolute bottleneck.
The processor is just essentially sitting idle waiting for the massive KV cache notes to travel through the wires. So, to sum things up, there are two major problems that current AI systems are facing. One problem is compute. The student is basically just drowning in his own notes and it's really hard for him to find things. The second problem is speed. Because there are so many notes, some of these notes are stored in filing cabinets in another room and the student needs to run back and forth to fetch these notes and put them on his desk which takes a ton of time.
So these are the limitations that DeepS is facing. How on earth did they tackle this? First of all, let's go over the architecture of a normal large language model. They use the transformer architecture which is basically made up of many layers. In fact, if you're curious about how transformers actually work under the hood, definitely see this video where I do a full explainer. But anyway, what happens is your prompt is broken down into data that flows through these layers of the transformer which ultimately outputs the next most probable word for its answer.
And then that word is appended back and then it's run through the model again to predict the next most probable word. and then this loops again until it generates your final answer. Now, each layer that it goes through generates its own notes or KV cache. And all of this must be stored somewhere. This often takes up a lot of space. So, not only does it fill up the high bandwidth memory, but it also spills over to SSDs that are outside the GPU.
Well, the architecture from this new deepse has many layers. But here's the really unusual part. They split this into two halves. There's a causal encoder component and then there's a decoder component. This already looks completely different from a standard transformer model. And here's how it works. During the prefill phase, again, this is where the AI is reading your prompt and all the information you give it. The last half is essentially turned off.
They just do nothing. And then the first half of the layers do all the heavy lifting. They read everything. They build the contextual understanding and they generate what's called the global KV cache. After reading everything, then the last half basically generates the answer. it does the writing or the decoding phase. The thing is it still needs to know the context in order to do the writing right to predict the next word.
So you might be wondering if it doesn't generate the KV cache itself, how does it understand the context? Here's why this design is so genius. Instead of having to read everything itself, it just looks at the output from the last layer of this encoder component. In other words, it borrows the completed global KV cache directly from the end of this encoder block. Now, this is quite a shocking and unexpected design because if they did this, if half its brain basically skipped the reading part, doesn't it lose some understanding of the context, which would make its answer worse?
Well, that's the exact risk of this architecture. And that's why it took incredible engineering to balance. What's fascinating here is that this last half, this decoder component, doesn't really entirely skip reading. It just skips calculating the global context. In other words, the global KV cache, but it still computes the local context for itself. Now, you might be wondering, what's the difference between global and local context here?
Well, global context is basically all the information that was given to it plus your prompt and everything else attached. It's basically all the notes that the student took after watching weeks and weeks of lectures. In contrast, the local context is just the immediate context of the sentence the AI is currently writing. In other words, what's directly relevant to the next word that it needs to output. You'll see this decoder component doesn't calculate any global KV cache, but instead it uses something called a sliding window attention to pay extremely close attention to the most recent tokens only, but not everything before it.
And it turns out that this design works pretty well. Here's a nice analogy to wrap your head around this. Imagine a company where they need to analyze thousands of pages of financial reports. Well, they would first get junior analysts to read every single page and crunch out the numbers, do the data analysis, and then write a very dense and accurate executive summary. Well, these junior analysts are basically like the first half of the model.
And then this executive summary is like the output at the end. They then hand the summary to the senior executives, which are like the decoder layers. These senior executives absolutely do not read the original thousands of pages. Instead, they just rely on the executive summary provided by the juniors. But when it comes time to sign off on something, then these senior executives put on their reading glasses and scrutinize the exact wording of the page sitting right in front of them.
That's basically the local context. The senior executives rely on this global summary for direction, but they mostly focus locally on the page in front of them for execution. And by structuring the model this way, DeepC completely bypasses the need for basically half of the model to generate its own massive KV cache notes. And this is a huge deal. It essentially slashes the compute required to read things by half. Now, reducing the compute is great, but we still have the problem of memory, right?
It's still generating these massive KV cache notes whether it's global or local and these are like overflowing on the student's desk and he's forced to like store these excess notes in filing cabinets in another room. How on earth can we reduce these massive piles of notes? And this brings us to one of the most fascinating parts of DeepS's new design. And this part is just brilliant. In fact, let me show you the results first so you can see how insane this is.
If you do any kind of content creation, definitely check out Luma, the sponsor of this video. Think of it as a creative AI agent that works alongside you through your entire creative process. Instead of just giving you the results of a single prompt, I can access the best image and video models out there. And the nice thing is instead of manually jumping between these tools, I can just get Luma agents to autonomously do entire workflows for me.
It can develop the concept, generate the visuals and shape the project all within the same workspace. For example, I can get it to generate a brand kit for me, design different products, and generate other marketing assets all inside the same project. And if I need to edit something, I can just prompt the agent to refine the results iteratively. One of the most powerful features is Luma skills. You can basically create reusable skills for workflows you use all the time.
Basically, you give Luma a set of instructions once and then you can run that same workflow on different assets whenever you want. For example, I can create a skill where I can input any product photo and it'll output some UGC videos of an influencer talking about the product. Or here's another example of a skill where I can upload a product photo and it'll generate a 360° orbit view like this. Luma basically gives you an intelligent creative co-pilot that can autonomously carry out your workflows.
Whether you're creating marketing campaigns, branded content, product visuals, or social media content, Luma is one of the best platforms you can use. Try Luma today using the link in the description below or by scanning the QR code here. If you compare the global KV cache size per token, which is basically the size of the notes the student has to take, DeepSeek V1 is almost 390,000 bytes. Now, if you fast forward just a few generations to this latest V4.1 Flash, it's only 890 bytes per token.
They basically shrunk the size of the notes down by like 437 times, which is crazy. Even if you compare this to the previous DeepSseek V4 Flash, this one still required like 3,500 bytes per token. So, this new update is like almost four times smaller than the previous generation. But here's the challenge to all of this. How can you compress these notes so much without losing the meaning? How can you still maintain the AI model's understanding of everything?
Well, Deepseek used a mechanism called compressed sparse attention 2 or CSA2. To understand this, let's first review how a normal transformer model works. Each layer in the model has to calculate its own KV cache. In other words, it has to make its own nodes. Conceptually, you can think of each layer as focusing on different things. For example, some layers might focus on certain patterns, while others focus on things like grammar or relationships between ideas or the broader meaning of the text.
Now, with all these layers each making their own notes, you can see how the total size of the KV cache could become really hard to manage. And if each layer has to calculate its own notes, you can see how things could become redundant. Well, this new CSA2 mechanism by DeepSeek completely shatters this redundancy. It introduces the concept of extreme sharing. So instead of creating new notes from scratch every time, each layer could use three different operating modes.
Full mode, reindex mode, and reuse. Let's go over each one. So if the layer is in full mode, it has to do all the hard work. It has to create brand new notes from scratch. In other words, it needs to make the full KV cache. But here's the important part. It also creates an index for future layers to search these notes. Think of the index like a guide or map or like a table of contents. Other layers can just look at this table of contents to figure out where exactly to search in the notes instead of trying to read the whole thing from start to finish.
So this makes it way faster to search for information. Another mode is called reindex. And here's where the efficiency kicks in. If the layer is in reindex mode, it just reuses the notes or in other words the KV cache from a layer in full mode. It doesn't create its own notes from scratch. But what it does do is create a new index from scratch. Again, think of this as like creating a guide or a new table of contents that searches the same notes as before, but highlights completely different parts.
For example, let's say you're giving the AI a ton of information about the history of the world. The full layer could make an index about things in chronological order, which might look like this. A reindexed layer would take the exact same notes, but give it a completely different table of contents. For example, instead of chronological events, it could be themes across time or it could be different technological breakthroughs.
It's basically like mapping different paths through the same notes. So that's the reindex mode. And then finally, we have the third mode, which is maximum efficiency. And this is the reuse mode. If the layer has this mode, it exerts almost zero memory effort. It just reuses the notes from the full layer as well as the indices from either the full layer or the reindexed layers. It doesn't write any new notes, nor does it create any new table of contents.
It just takes information that's already available to it. And with this design, with these different modes, each layer doesn't have to store as much notes. The total size of these notes is reduced significantly because some of these layers don't even need to create new notes at all. They're just reusing notes and indices from previous layers. But wait, this ain't all. Deepseek takes this one step further. They've added something called a hierarchical sparse indexer.
And here's how it works. You see, in the last half of the model, this is the decoder part. The very first layer acts kind of like a gatekeeper. It scans all the global notes from the previous layer. Remember, this is also called the global KV cache and it generates a candidate pool. Think of this like a short list. Basically, from those piles and piles of notes, it figures out just the most relevant concepts. Specifically, out of a million tokens, it only selects around 16,000 that are the most relevant.
Everything beyond it is basically ignored and then all the subsequent layers are forbidden to search for anything else in the notes. So, this drastically narrows the universe of possible answers right at the start of the writing phase. Now, obviously, as you may expect, if we restrict the AI's search space like this, it might hallucinate or miss important details, right? its response will become dumber if it doesn't look at everything.
But here's the genius behind this. Deepseek was able to make it work by training the model to build these candidates with such high accuracy that the later layers don't even notice the rest of the info is missing. The elegance of their engineering is just profound. Let's take a moment to appreciate what they did here. They don't have the best Nvidia GPUs. They don't have the biggest data center in the world. Heck, they're pretty starved for compute.
So instead of focusing on the hardware side, they completely optimized the software part so the hardware doesn't have to work so hard. And we ain't done yet. You see, all this intense optimization also led to some additional issues they had to fix. Remember this sliding window attention mechanism we talked about earlier. This is where the senior executive doesn't read everything, but when he signs off on something, he has to look really closely at all the information on the page in front of him.
Well, this sliding window attention is the AI's hyper local short-term memory. And according to the paper, this hyper local memory was actually causing a huge storage problem. You see, when a user has a conversation with the AI over multiple turns. In other words, when you say something, the AI replies and then you reply back and so on. The AI has to save this local context of every single turn into its SSD so it won't forget the flow of the conversation.
And this constant saving of short-term local memory was completely clogging the hard drives. Going back to our analogy, this is like filling up all the cabinets in the other room down the hall. Now, when you cache data, it means you save it in a temporary location so you can retrieve it quickly later, right? But if it gets too big and these notes are located in cabinets in another room, well, fetching this information gets super slow. the student needs to run down the hallway to the other room to grab the notes and then place them back on his desk.
So, DeepSeek recognized this issue and their solution was something called SWA bounded replay. And this is probably the most unexpected and shocking part of their new design. Their solution to this local memory clogging up the hard drives is to simply delete it. They literally deleted the AI's short-term memory completely. So, when the AI finishes generating its reply to you, its local memory just evaporates into thin air.
As you can imagine, if we delete its short-term memory, shouldn't it become like completely disoriented? Wouldn't it lose track of what's going on? Well, that's what we would expect. So, what Deep Seek did was they basically got the AI to recalculate the last parts of the conversation instantly, specifically the last 128 tokens of the conversation. In other words, they forced it to generate the most immediate new notes from scratch right on the spot every single time.
This sounds so counterintuitive, right? I just spent the past few minutes in this video explaining how they tried to reduce the compute and split the brain into halves to avoid generating and reading so many notes. But now, if we get this AI to recalculate this last part of the conversation every time, wouldn't this be incredibly inefficient? Doesn't this slow everything down? Well, here's where DeepSseek gives us a masterclass in efficiency.
Let's walk through the trade-off here. You kind of have two possible ways to handle this short-term memory. The first way is the normal way where we take data from the GPU, we push it through the motherboard and write it into a hard drive. And later, when it needs to use this shorter memory, it needs to search for it, pull it back up through the motherboard and into the GPU processor. This is super slow because you're physically moving data across this distance.
Now, the second way is to just use the raw power of a modern GPU to simply recalculate its short-term memory. In other words, just rewrite the most recent and relevant notes from scratch. And for a modern GPU, just crunching like 128 tokens is just a microscond operation that barely takes up any time or power. So, doing the math is actually much faster than trying to transfer this data back and forth from the SSD. The Deep Seek team realized that saving this short-term memory to the hard drive was just a massive waste of a really slow resource.
By completely removing the step and just getting the GPU to quickly recalculate everything on the fly, not only does it make things faster, but it also freed up a ton of storage space. Let's review what we've gone over so far. Deepseek used this sliding window attention to focus on the most immediate bits of information. They also split the model in half and then used hierarchical sparse indexing to significantly reduce the number of notes that the final layers need to process.
They also forced some of the layers to just reuse existing notes, which helps lower compute. And then finally, they also completely removed its short-term memory to save hard drive space. But guess what? We're not done yet. So DeepS also added some additional components that support the main architecture. One of them is called the single pass MHC. Now, to understand why this matters or what this is, you first need to know that running an AI model isn't just about doing a huge amount of math.
It's also about constantly moving data around. When one part of the model performs calculations, it often produces intermediate values that the next part of the model needs to use. Normally, you might need to store these intermediate results into the GPU's memory and then load it again when the next operation needs these values. And when the AI is processing a really long task, it's basically doing this step but billions of times.
Even though one single movement is just a split second, if you multiply this by billions, then this latency can add up. And this step of transferring data can actually be the bottleneck. Well, single path MHC is designed to eliminate some of these unnecessary trips. It's quite technical, but basically it mathematically aligns multiple operations together so that they can occur simultaneously. It's basically combining steps instead of running them one by one.
And it turns out that this significantly reduces the memory traffic inside the GPU, making it way faster. And that's not all. They also introduced another supplementary component called the engram. This is kind of like a separate memory module. Now this contains 168 billion parameters and instead of living in the expensive memory of the GPU, this is designed to live in the standard RAM of the server or the computer. So it's physically separated from the GPU.
And this part is in charge of storing static facts. So it's memorizing things like historical dates, capitals, or other fixed facts. You see, this is actually really important because we want the GPU to focus entirely on active reasoning and thinking. We want to free up as much of the GPU's expensive memory as possible to maximize its thinking capabilities. For these fixed static facts, which the AI doesn't really need to think about, we can put this in cheaper memory that lives outside the GPU.
So, this prevents the GPU's memory from being clogged with static data. Only when the model needs to access these facts would it pull from this engram module. It's kind of like a really brilliant senior lawyer working on a case and synthesizing information. He doesn't need to memorize every single clause out there. He's in charge of the strategic reasoning. Instead, he has an assistant sitting next to him so that when the lawyer needs a specific date or a precise quote, he can just get the assistant to instantly fetch it for him.
Well, this assistant is kind of like the engram. It frees up the thinking capacity for the senior lawyer so he can focus on the real strategic reasoning work. And we're not done yet. Deepseek also introduces a third supplementary component called DS-spark. In fact, I already did a full explainer video on D-Spark right when it came out. And this is quite a revolutionary breakthrough. You see, normal AI models need to output their answer one word at a time, which can be very slow.
What DSpark does is it essentially allows the model to output multiple words at a time, making it way faster. Now, this is quite technical, so if you're interested, see this video for a full deep dive on DSpark. All right, so we've covered a ton of stuff. Let's take a step back and summarize everything so far. You see, this freak of a model isn't just from one single breakthrough, but a ton of different components added together.
They've added this sliding window attention to only focus on the most immediate bits of information. They split the brain into two halves to drastically reduce the compute. We also have extreme compression using CSA 2 and with the reindex and reuse layers. This drastically reduces the amount of nodes or KV cache that it needs to generate. We also completely deleted its short-term memory with this bounded replay mechanism.
And we also added a ton of supplemental mechanisms such as this MHC pathway to make its calculations more efficient and this engram component so it doesn't waste any compute on hardfax. And finally, we also added this D-spark mechanism which helps it generate more than one word at a time. And when you combine all these parts together, it becomes an absolute Frankenstein of efficiency. They've built the most optimized, turbocharged, frictionless Frontier model we've seen so far.
But don't take my word for it. Here's a chart showing how ridiculous this is. If you look at figure two, this is one of the most incredible demonstrations of the efficiency of this new model. Here it's showing the flops on the y-axis which is like the raw computational power required to generate a response and the x-axis is the context window size or basically how much information it can store in its memory at once. When DeepSeek increases this window from a standard 4,000 tokens all the way to a massive 1 million tokens which is roughly 700,000 words or like a medium-sized code base.
You can see that for this new 4.1 flash model, the decode curve remains almost completely flat. In other words, if you feed the AI a tiny one-page document versus thousands of pages and ask it a question, this chart shows that the AI spends roughly the same amount of energy per word. This is a massive deal. It completely defies the laws of AI because normally what we would expect is if you scale the context, the compute also increases as you could see with the previous generations of deepseek models.
But here they've basically flattened the curve. And you know the ridiculous thing is not only is this super efficient and frictionless, but its performance is also state-of-the-art, pretty much on par with the Frontier models. For example, if you look at Deep Su 1.1, this scores 74.2, which not only beats the other open models out there, but if you look at the official leaderboard, then it's basically on par with GPT6 Astra, which also scores 74.
Or if you look at Cyberjim, this is pretty much state-of-the-art. Same with Automation Bench. You can see that this freak of a model even beats GPT6 Astro Max. If you look at this leaderboard by LiveBench, then you can see that this new Deepseek is the number one ranked open model. Same with Val's index, which measures an AI's performance across knowledge work tasks. You can see that Deepseek 4.1 Flash is currently ranked number one.
And look at the insane cost of this. Not only is this the most performant model, but it's also the cheapest. You can see that second and third place are over 20 times more expensive. Here's another chart showing the cost per task. You can see that this new DeepSeek model is all the way over here, which is like way cheaper than the Frontier GPT6 Astra as well as the extremely overpriced clawed models. If you look at the output speed if you use it through their API, again, this is insane.
This achieves over 200 tokens per second, which is like four times faster than GPT6. The latency, in other words, the time to its first answer is also the lowest in the industry. All right, so we've gone over a ton of stuff. The way they designed it is just so unexpected and at many times counterintuitive, but once you understand why they did it, then you'll see how brilliant and genius this is. And as always, they've open sourced the model for you to download locally, so you can do whatever you want with it.
My hat off to the Deep Seek team for pulling off the impossible. Once again, I mean, the previous DeepSseek V4 was already super efficient, but with this latest model, they've managed to completely redesign the architecture and squeeze even more juice out of it. That sums up my deep dive on this new Deep Seek 4.1 Flash. This is one of the more technical papers I reviewed so far on my channel. So hopefully I made it easy for you to digest.
In fact, the paper is jam-packed with a ton of additional technical details which I didn't have time to cover. So if you're interested in digging deeper, I'll link to this original paper in the description below as well. Let me know in the comments what you think of this. As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content.
Also, there's just so much happening in the world of AI every week, I can't possibly cover everything on my YouTube channel. So, to really stay up to date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching, and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.