Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Hefty LLM · @HeftyLLM
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
0:076.6x the video's typical replay level
When a startup built a chip that can run AI 20 times faster than Nvidia's best GPUs, that too while using 10 times less power doing it, Nvidia didn't try to beat it. Instead, they paid $20 to make it disappear. And I'm not even joking. Nvidia controls roughly 90% of the AI
Said at 0:00
Most replayed moment #2
2:413.1x the video's typical replay level
millisecond the core asks for it. That is how the LPU is able to hit 80 terabytes per second and obliterates the GPU in raw speed. But here's where it breaks. When every single bit of memory requires six transistors instead of one, your memory explodes in size. SRAM is
Said at 2:34
Most replayed moment #3
3:293.0x the video's typical replay level
cannot fit very much of it on a single chip. While an Nvidia GPU holds an easy 80 GB, a single LPU can still only hold 230 MB, at least for now. And unlike a GPU, it cannot train a model. It can only run it. That sounds like a death sentence until you see what happens when
Said at 3:22
The graph counts replays. It does not show where viewers stopped watching.
Words
1,139
Runtime
6:13
Speaking pace
183wpm
Reading time
5min
183 words per minute, just over the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
When a startup built a chip that can run AI 20 times faster than Nvidia's best GPUs, that too while using 10 times less power doing it, Nvidia didn't try to beat it. Instead, they paid $20 to make it disappear. And I'm not even joking. Nvidia controls roughly 90% of the AI chip market. Governments are stockpiling their hardware, and every major tech company is spending billions just to get enough of them. But in December 2025, Nvidia acquired Groq in a deal specifically structured to avoid antitrust regulators. They silently
92 words, the words spoken in the first 30 seconds at 183 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 85 |
| Average words per sentence | 13.4 |
| Longest sentence | 40 words |
| Questions asked | 7 |
| Sentences containing a number | 16 |
Most used terms
Filler phrases
10 in total: actually 4 · like 3 · literally 3.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
When a startup built a chip that can run AI 20 times faster than Nvidia's best GPUs, that too while using 10 times less power doing it, Nvidia didn't try to beat it. Instead, they paid $20 to make it disappear. And I'm not even joking. Nvidia controls roughly 90% of the AI chip market. Governments are stockpiling their hardware, and every major tech company is spending billions just to get enough of them. But in December 2025, Nvidia acquired Groq in a deal specifically structured to avoid antitrust regulators.
They silently absorbed their engineers, the patents, and the entire architecture, then folded it into their next-generation hardware. That tells you exactly how dangerous this thing was for Nvidia. So, what did Groq actually build? And if this architecture is truly 20 times faster, why hasn't it already replaced the GPU? This is an LPU, and the reason it demolishes a GPU at inference comes down to a fundamental physical flaw inside every GPU that ever existed.
You see, a GPU was never built for AI. It was a graphics accelerator. It was built for gaming. And when you use it to run a large language model, you immediately run into a problem. Every single time the GPU generates one token, it has to go fetch the entire model's weights from VRAM, drag them to the cores, do the math, and then send the result back. It does this for every single token. At 40 tokens per second, that's 40 round trips to memory every second.
The actual math is almost instant. The chip is literally just waiting most of the time. But LPU's entire architecture is built around one idea: what if the model never had to leave the chip at all? And to make that happen, they had to abandon traditional GPU memory completely. Modern GPUs store massive amounts of data using DRAM. To hold a single bit of information, DRAM uses just one tiny transistor and a microscopic capacitor.
Because it's so small, Nvidia can cram 80 GB of it onto a single GPU. But there are two massive catches. First, capacitors leak electricity and have to be constantly refreshed. Second, because this memory bank is so incredibly huge, it cannot fit inside the compute core. It sits outside, centimeters away from the core. A few centimeters sounds tiny, but when a chip is doing billions of calculations per second, those centimeters turn into hundreds of miles.
Groq bypassed this entirely by using SRAM. SRAM doesn't use leaky capacitors. It uses a locked circuit built out of six transistors. And more importantly, the memory isn't sitting on a separate chip across the board. The SRAM is baked directly into the silicon die itself. The memory and the math cores are literally millimeters apart. The data has no commute. It's instantly available the exact millisecond the core asks for it.
That is how the LPU is able to hit 80 terabytes per second and obliterates the GPU in raw speed. But here's where it breaks. When every single bit of memory requires six transistors instead of one, your memory explodes in size. SRAM is incredibly fast, but it is physically massive on the silicon die. Now, you might be thinking, "Hmm, isn't that an over-glorified L1 cache? What a big deal?" But it is indeed a massive deal.
There is a reason why you don't get 200 megabytes worth of L1 cache on your GPU. You don't even get 200 kilobytes of it, actually. Printing this many six-transistor circuits onto one continuous die is not a joke. It pushes the absolute radical limit of semiconductor manufacturing. But even an engineering miracle can't break the laws of physics. Because those SRAM blocks are so physically large, you simply cannot fit very much of it on a single chip.
While an Nvidia GPU holds an easy 80 GB, a single LPU can still only hold 230 MB, at least for now. And unlike a GPU, it cannot train a model. It can only run it. That sounds like a death sentence until you see what happens when you start stacking them. But first, how did Groq manage to print 230 MB of it when nobody has dedicated this much silicon to pure on-die SRAM as a memory replacement? Simple. They deleted the rest of the chip.
Nvidia uses up their silicon space by building complex physical hardware to direct traffic and schedule data. Groq realized AI math is entirely predictable, so they ripped out the hardware schedulers, leaving massive amounts of empty space on the silicon. They filled all that empty space with SRAM, but by deleting the physical traffic routers, Groq had to move all that traffic control into the software, and that software is the actual reason Groq is so incredibly terrifying to Nvidia.
But even though it's a big deal to have this much SRAM on it, it's still nowhere close to running an actual state-of-the-art model. Real models are hundreds of gigabytes. So, how is that even competing against GPUs if it can't even run some of the smallest AI models? And that's why the chips are built from ground up to be clustered with hundreds of them together. And here's where every other chip manufacturer fails. The moment you chain that many chips together, they spend more time waiting on each other than actually doing what they're supposed to.
But Groq's compiler is perfectly deterministic. It knows everything before it even happens. Before the code ever physically touches the silicon, Groq's software maps out the exact path of every single electron. It pre-calculates the entire journey down to the exact nanosecond. The chips never have to ask each other for data, and they never wait. They just blindly execute their math, and the data arrives perfectly on time.
The compiler forces 256 independent chips to act as one flawlessly synchronized brain. And the real reason why Nvidia paid $20 isn't actually about the raw chips. Their GPUs are already the undisputed kings of raw compute. Nvidia's actual nightmare is the network. When a company like OpenAI builds a data center, they are never using just one GPU. They are wiring together tens of thousands of them. And right now, getting 10,000 GPUs to talk to each other without causing massive network bottlenecks is one of the hardest engineering problem on Earth.
Groq solved multi-chip scaling gracefully. The more chips you stack together, the better it gets. They bought Groq to rip out their deterministic networking technology, fold that software into their next generation of GPUs, and permanently neutralize the only company that figured out how to beat them at scale. If you liked this video, you will definitely like this one where the chip literally just stops using electricity to compute.
See you there.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.