Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Two Minute Papers · @TwoMinutePapers
Words
734
Runtime
5:10
Speaking pace
142wpm
Reading time
3min
142 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
The new DeepSeek is here. 4.1 flash and it is wow. Look at that. It is incredibly fast. And I was also stunned by this. Look, it can outperform Claude Opus 5 Kimi K3 on some tests. Not everything. We'll see about that. But there is a catch. I'll show you at the end. It also reliably outperforms DeepSeek 4.0 Pro, which is a huge surprise. I mean, that system costs
71 words, the words spoken in the first 30 seconds at 142 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 76 |
| Average words per sentence | 9.7 |
| Longest sentence | 30 words |
| Questions asked | 5 |
| Sentences containing a number | 13 |
Most used terms
Filler phrases
1 in total: I mean 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
The new DeepSeek is here. 4.1 flash and it is wow. Look at that. It is incredibly fast. And I was also stunned by this. Look, it can outperform Claude Opus 5 Kimi K3 on some tests. Not everything. We'll see about that. But there is a catch. I'll show you at the end. It also reliably outperforms DeepSeek 4.0 Pro, which is a huge surprise. I mean, that system costs maybe 300k dollars to run locally and this one can be run for a quarter of that.
Well, that's still a lot of money, but the tendency is undeniable. This might run in our pocket in a year's time. It also has native visual understanding. So, finally, you can give it images of an iconic game menu and have it write a game that reproduces it. I'll tell you more about my other experiments at the end. And this is huge and small at the same time. How? Well, now hold on to your papers, fellow scholars. It's small because of the KV cache.
This contains the context and is 437 times smaller than V1 was 3 years ago. But also four times smaller than the previous 4.0 flash. But that was just a few months ago. And you compressed it 4x. That is what? How did they do that? Do we know? Yes, we do. Dear fellow scholars, this is Two Minute Papers with Dr. Károly Zsolnai-Fehér. Why? Because this is an open model and there is a free research paper explaining [music] it.
I am not an expert, just a student, but I'll try to break it down. They call this technique CSA2 and they are finally onto something here. You see, a neural network is given by a lot of layers of neurons, and as information is propagated through these layers, each of them has its own KV memory. But, not here. DeepSeek 4.1 Flash finally gives us shared memory between the layers. Not everyone has to remember everything.
No. They share the memories. [clears throat] It's a nightmare to pull off well, because different layers need different views of the same history. And if we pop the hood and look closer, hohoho, look, an encoder-decoder structure. The encoder creates a shared global memory, and the decoder reads from it. And all that is what finally gives us a much smaller KV cache, which is the bane of my existence, because it needs too much video RAM.
Incredible leap forward. Now, I also said this is huge, more than 500 billion parameters. So, I have no chance to run this at home whatsoever. And when I tried to reproduce this beautiful honeycombing paper, GPT-6 Astra nailed it. Absolutely stunning. Ooh. Opus 5.1, not as stunning, but the physics is quite formidable. Rendering a bit less so. Now, this one. Well, we still need to get a couple more papers down the line to match that, but it will happen.
Of that, I am certain. And you see, I am here to show you the truth, not just believe the headlines. Now, also, since YouTube stopped recommending my physics simulation papers, which breaks my heart, at least I get to show you this, which fills me with joy. Subscribe and hit the bell to see more. Also, the catch, 4.1 flash likes to think a lot and burns a lot of tokens. Fortunately, it is not that expensive, but man, that's a lot of tokens.
But, if you have the hardware, you get to run all this for free. I don't, so I run it through the API or lambda, but given that every time Deep Seek publishes a new paper, AI gets cheaper for all of us. So, thank you so much for that. Huge contribution to humanity. This is going to help a lot of doctors and scientists do their work. Open science for the win. What a time to be alive. I use lambda to reproduce AI research papers often in minutes.
It's also great to train your own models or fine-tune an existing one. Run inference or text to image or video, easy peasy. Running a Deep Seek chatbot or agent, super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover, and moments later, results. Love it. Seriously, try it out now at lambda.ai/papers.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.