Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Hefty LLM · @HeftyLLM
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
0:494.0x the video's typical replay level
current architecture incompetent with our growing AI needs. First and the most important is the memory wall. Memory wall refers to the increasing gap between how much we have improved in terms of chip performance compared to the improvements in the memory performance. This gap is worse than ever and it's only
Said at 0:41
Most replayed moment #2
5:103.7x the video's typical replay level
currently being developed in university labs, the photonic latch. Researchers are figuring out how to create optical versions of SRAM, allowing microscopic loops of light to act as ultra-fast temporary cache memory directly on the chip without ever converting back to electricity. If they can scale that up, the GPU
Said at 5:03
Most replayed moment #3
1:123.3x the video's typical replay level
consumption of a GPU is 500 watts, all 500 watts would be converted to heat and the GPU would consume almost zero power by itself. Why, you may ask? Because all a GPU or any other compute device does is calculate and produce information. Information is not a physical quantity and it doesn't possess any physical
Said at 1:05
The graph counts replays. It does not show where viewers stopped watching.
Words
1,184
Runtime
5:26
Speaking pace
218wpm
Reading time
5min
218 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
By 2026, GPUs have been used for AI more than anything else, but they are hitting a massive physical wall, power consumption and extreme heat. But a company just successfully deployed a completely different kind of processor into a German supercomputing center. The craziest part is that it doesn't even compute with electricity. And surprisingly enough, it can run on your computer just like any normal GPU. And the claims are even crazier. Qant claims that their chips provide 3,000% higher efficiency and 5,000% higher compute compared to a traditional GPU. Well, it turns out computing at the speed of light creates a completely new bizarre bottleneck that nobody
109 words, the words spoken in the first 30 seconds at 218 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 74 |
| Average words per sentence | 16.0 |
| Longest sentence | 45 words |
| Questions asked | 5 |
| Sentences containing a number | 4 |
Most used terms
Filler phrases
10 in total: like 6 · actually 2 · kind of 1 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
By 2026, GPUs have been used for AI more than anything else, but they are hitting a massive physical wall, power consumption and extreme heat. But a company just successfully deployed a completely different kind of processor into a German supercomputing center. The craziest part is that it doesn't even compute with electricity. And surprisingly enough, it can run on your computer just like any normal GPU. And the claims are even crazier.
Qant claims that their chips provide 3,000% higher efficiency and 5,000% higher compute compared to a traditional GPU. Well, it turns out computing at the speed of light creates a completely new bizarre bottleneck that nobody saw coming. But right now, they are at the edge of solving it. But before moving on to the photonic chips, we first have to understand why the current GPUs need a replacement. There are some massive bottlenecks that are slowly making our current architecture incompetent with our growing AI needs.
First and the most important is the memory wall. Memory wall refers to the increasing gap between how much we have improved in terms of chip performance compared to the improvements in the memory performance. This gap is worse than ever and it's only increasing. Considering the fact that how dependent GPUs are on the VRAM, it directly makes them a subject to this holy grail of bottleneck. The efficiency and heat generation on the other hand is a bit complex.
You see, if the total power consumption of a GPU is 500 watts, all 500 watts would be converted to heat and the GPU would consume almost zero power by itself. Why, you may ask? Because all a GPU or any other compute device does is calculate and produce information. Information is not a physical quantity and it doesn't possess any physical fundamentals such as mass, velocity or position. Therefore, it cannot contain any energy.
This is why all energy coming into the GPU would have to go out as immense amount of heat. This is why the limits of efficiency is far from our reach and there is a lot we can improve. And here's the interesting part. Qant's native processing unit doesn't even compute with electricity. It computes with light and it doesn't use standard digital zeros and ones, either. It is fundamentally an analog processor. Instead of shoving billions of transistors in a small silicon die, the core of this chip is made of a crystalline material called thin-film lithium niobate.
Instead of pushing an electrical current through logic gates, the data is converted into a laser beam. The chip then splits, bends and modulates those light waves across different frequencies. And this is where it gets interesting. The physics itself In a traditional GPU, if you want to calculate a heavy AI workload like a complex non-linear function or running a transformer model, you need millions of transistors flipping on and off across multiple clock cycles just to brute force the multiplication.
But on this chip, you shine the light through a specifically shaped optical element. As the different wavelengths of light intersect, they naturally create interference patterns. That interference pattern itself is the mathematical answer. The math literally solves itself instantly as the light passes through the chip. And because photons have no mass, they don't grind against the material create any resistance that can produce heat and waste energy.
On a GPU, the electrical resistance is the single biggest reason for the heat output. But photonic chips avoid that by using photons instead of electrons. So, the claims seem to be making much more sense than they should. But what about real-world use cases? We frequently see such revolutionary stuff that never gets to the market. But recently, they packed this crazy photonic chip into a standard 19-in rack-mountable server and deployed it to the Leibniz Supercomputing Center in Germany.
They successfully installed the world's first commercial analog photonic AI processor into a fully operational high-performance data center. So, how does a data center actually use a photonic chip? The best thing is that it's built as a standard PCIe card. This chip fits directly into the exact same server slots that traditional GPUs currently use. But the hardware is only half the battle. Remember why the GPU is practically a monopoly?
It's because of the software and infrastructure. If developers had to learn optical engineering just to train a neural network, these chips would be dead on arrival. Nobody is going to rewrite millions of lines of code for this. To solve this, Qant built a bridge called QPAL, the photonic algorithm library. It acts as an invisible translator. Developers can write their AI models in standard Python and PyTorch, the exact same way they write code for Nvidia GPUs today.
The library handles the translation from digital software to analog light waves automatically. And as for what they are actually doing with it, these light processors are being evaluated for some of the most compute-heavy non-linear workloads we have. Real-time medical imaging, complex climate modeling and running massive visual AI models with drastically fewer parameters than a traditional GPU would require. But before we declare the GPU as completely dead, there is a massive elephant in the room to address.
And it's the exact same problem we started this video with, the memory wall. Photonic chips like the ones from Qant are practically magic at calculating data, but they have one fatal flaw, they cannot store the calculated data. This means that while the actual math is being done by light, the AI model's weights, its memory cache and the final results still have to be stored in traditional electrical memory like standard VRAM.
And this introduces a new bottleneck called the conversion penalty. Every single time the photonic chip needs to pull data from memory, the system has to convert standard electricity into a laser beam. And every time the math is finished, it has to convert that laser beam back into an electrical signal to store the answer. This optical-to-electrical conversion takes time, power and introduces latency. If an AI workload requires a massive amount of back and forth reading and writing to the VRAM, which it usually does, those conversions will eat up all the latency and energy savings that the light-speed calculation just gave you.
So, how is the industry fixing this? Right now, chips like the Qant NPU are highly specialized. They are deployed as coprocessors alongside traditional hardware, specifically targeted at math-heavy non-linear workloads where the pure calculation takes significantly longer than the memory fetch. But the holy grail is currently being developed in university labs, the photonic latch. Researchers are figuring out how to create optical versions of SRAM, allowing microscopic loops of light to act as ultra-fast temporary cache memory directly on the chip without ever converting back to electricity.
If they can scale that up, the GPU monopoly is done for. Also, if you like this video, you can consider subscribing to my channel. Oh, and you can also join my channel membership, by the way.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.