Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Hefty LLM · @HeftyLLM
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Hefty LLM's most watched videos.
Most replayed moment at 0:49
4.0x that video's typical replay level
current architecture incompetent with our growing AI needs. First and the most important is the memory wall. Memory wall refers to the increasing gap between how much we have improved in terms of chip performance compared to the improvements in the memory performance. This gap is worse than ever and it's only
Said at 0:41
Most replayed moment at 0:07
6.6x that video's typical replay level
When a startup built a chip that can run AI 20 times faster than Nvidia's best GPUs, that too while using 10 times less power doing it, Nvidia didn't try to beat it. Instead, they paid $20 to make it disappear. And I'm not even joking. Nvidia controls roughly 90% of the AI
Said at 0:00
The graph counts replays. It does not show where viewers stopped watching.
Words
2,070
Runtime
12:11
Speaking pace
170wpm
Reading time
9min
170 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
In 2024, researchers at Microsoft published a paper that figured out how to strip away the heaviest mathematical operation in AI. Quite literally replacing it with elementary school addition and subtraction. The result were models that match the intelligence of full precision ones on benchmarks. And guess what? That's all while running on cheap consumer hardware. It's called BitNet. It cuts power consumption by up to 82% shrinks memory by 75% and allows models to run on a CPU at blazing fast speeds without touching
85 words, the words spoken in the first 30 seconds at 170 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 119 |
| Average words per sentence | 17.4 |
| Longest sentence | 45 words |
| Questions asked | 12 |
| Sentences containing a number | 44 |
Most used terms
Filler phrases
28 in total: actually 7 · basically 7 · like 4 · kind of 3 · literally 3 · I mean 2 · right? 1 · sort of 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
In 2024, researchers at Microsoft published a paper that figured out how to strip away the heaviest mathematical operation in AI. Quite literally replacing it with elementary school addition and subtraction. The result were models that match the intelligence of full precision ones on benchmarks. And guess what? That's all while running on cheap consumer hardware. It's called BitNet. It cuts power consumption by up to 82% shrinks memory by 75% and allows models to run on a CPU at blazing fast speeds without touching a single GPU.
Yet today, Anthropic, OpenAI, Google, and not even most of the open-source models use it. So, what went wrong? Are we simply early or is this just monopolies being monopolies? You see, everything a language model knows is stored in its weights, which are basically just billions of numbers saved at 16 bits each, so 2 bytes per weight. So, if you do the quick math, an 8 billion parameter model is 8 billion times 2 bytes, which is 16 GB, and your GPU has to drag every single one of those bytes across its memory bus for every single token it generates.
That's the expensive part, since moving a number from memory to the chip costs hundreds of times more energy than the actual math. When you run a model at home, your speed is pretty much decided by how much memory bandwidth you have. That's pretty much why everyone runs 4-bit models locally. What you're really paying Nvidia for is the memory. No wonder RAM prices are so absurdly high. But hey, what if we go beyond 4 bits?
That's what Microsoft said, too. The paper is called The Era of 1-bit LLMs, and the subtitle literally says all large language models are in 1.58 bits. In their design, every weight inside the big matrix layers can only be +1, 0, or -1 with one shared scale number for the whole matrix. Normally, each weight gets multiplied by a number flowing through the network, but if a weight is plus one, the chip just adds that number to the total.
If it's minus one, it subtracts it. And if it's a zero, it skips it completely. Attention and a few small layers still multiply, but the giant weight matrices, which is where most of the models' math happens, they turn into integer additions. The weird 1.58-bit number is just how much information of three-state weight holds, which is log base two of three. And since three to the power of five is 243, which fits inside a single byte, you can pack five weights into eight bits and get down to 1.6 bits each.
And the results were pretty wild. At three billion parameters, the ternary model matched a regular 16-bit model Microsoft trained on the same data. It was using 2.22 GB of VRAM instead of 7.89. It's 2.71 times faster. Then in October 2024, Microsoft released BitNet.cpp. It's a CPU runtime built on llama.cpp. On an Intel laptop chip, it ran 2.37 to 6.17 times faster than llama.cpp running these same models at 16 bits.
All while cutting energy per token by over 72%. A 100 billion parameter model running on a single M2 Ultra CPU at five to seven tokens per second. But if you read Microsoft's repo, it says the models they tested are dummy setups built to show off the speed. Obviously, since a real ternary model that big didn't exist. Their first real one came out in April 2025, BitNet B1.58 2B4T, with two billion parameters. Against Qwen 2.5 1.5B, a normal 16-bit model trained on over four times more data, it scored 54.2 on average to Quen's 55.2.
And it did that while generating each token in 29 milliseconds on a laptop CPU compared to Quen's 65. So, if all of this works, what went wrong? Well, the obvious move would have been to just take Llama or Quen and convert it, right? Like how the llama.cpp crowd squeezes every new model to four bits within a day of release. But, you can't just round your way to ternary. At four bits, every weight snaps to one of 16 levels, which is rough, but close enough.
While at ternary, it snaps to one of three. So, most of what each weight learned basically gets erased, and those errors pile up through every layer. In test by a startup called Prism ML, which we'll get to in a second, even a standard two-bit build of a Quen model drops from the four-bit build's average score of 85 down to 70 two. So, how do you actually train a model to work like this? It's called quantization-aware training.
The model keeps two versions of every weight while being trained. One is the full precision version, and the other is the ternary one. The ternary weights are used to make the forward pass, and the full precision ones are used for the backward pass. You see, how a normal training works is essentially the model weights making a guess, and then based on whether it was right or wrong, they update their own weights. Making a guess is called forward pass, and learning what was the right answer is called a backwards pass.
Basically, how it works for ternary models is that the rounded-off ternary weights do the forward pass and make a guess, but since they lack the precision to actually learn the right answer in the backward pass, a full precision model is used for it. Formally, it's called straight-through estimator. After billions of low-precision guesses, the model naturally learns to produce values that just so happens to clearly round on those three numbers.
The catch here is that this training still runs in 16-bit. So, ternary makes the finished model cheap for inference, but training it is still as expensive as it was before. In early 2024, the only proven way to do it was a full pre-training run from scratch, which is how Microsoft's 2B model ended up eating 4 trillion tokens. And even when huggingface converted Llama 3 8B later that year, it still took 100 billion tokens of retraining all for an architecture nobody had proven past a few billion parameters.
So, if you're a lab that already has a working 16-bit model, I mean, that's a pretty hard sell. And even if a lab does pay the cost, ternary models hit a limit that's not exactly what training can fix. It can hold at most 1.58 bits of information, which caps how much a model can memorize per parameter, and you can see it in the benchmarks. Microsoft's 2B model loses to Quen on TriviaQA 33.6 to 38.4 even while beating it on several reasoning tests.
TriviaQA is basically a test on how good the information a model can retrieve, so the model can still think fine during this, it just can't remember the facts back out when you ask for it. Hard drives, on the other hand, are pretty similar to this as well. It stores all the data in full precision, but when it comes to actually retrieving what it stored, it only gives you a few options. You either have to know the exact file name, or you have no other option than scrolling to find that file on a decently large folder.
And in order to fix this, we need to give more data that can be used to retrieve a file. So much so, that instead of typing out an exact string of the file name, you can now simply describe what the file looked like and still get the file in your hands. Let's say an image of a dog that's sitting on a couch, or maybe a video file where you vaguely remember you saw a red Ferrari. Or how about a document or an audio file where someone briefly mentioned how much they love donuts.
Software like Hefty Search allows you to do that without ever connecting to a cloud server. If you still couldn't guess what I'm talking about, Hefty is an app that I've been building for the past 9 months. I made it to solve my own problems, but I soon realized that it's not just me. It allows you to press two buttons on your keyboard and this search bar comes up. You can type out whatever file you want by a small description of it and it just finds it in under half a second.
Check out the link in the description and let me know your thoughts in the comments. So, are we simply early? Kind of, because that training wall only started cracking this year. In September, Prism ML released ternary bonsai 2. It's a version of Qwen 3.8 27 billion where every one of its language weights is in ternary and it went from 54 GB down to 5.9. On Prism ML's own benchmarks, it kept 98.2% of the original's average score with math and coding basically on the same level.
The weights are up on Hugging Face right now if you want to run the model. And the trick is that they didn't start from scratch. The hidden full precision copy starts out as Qwen's actual trained weights. So, all the language and intelligence is already there. The original model runs alongside as a teacher and the ternary version gets trained to simply match its answers. The first 27B bonsai came out in July and people are already running it on a single 16 gig 5060 Ti.
Apple actually got pretty close to this. The roughly 3 billion parameter model behind Apple Intelligence already runs at 2 bits per weight using this same kind of training. But, we're only halfway there yet. We still need to figure out how to remove the multiplier, but it's kind of messy if you think about it. You see, the tensor cores inside your GPU are basically arrays of multipliers. It can very easily do all that complex matrix multiplication, but it's gone so specialized that it literally can't do that addition and subtraction at all.
The weights have to be unpacked back into normal numbers before the multipliers can even touch them. And then those multipliers spend their time multiplying by these ternary numbers. You can also see how inefficient this is, but hey, at least the RAM consumption is lower. Despite this inefficient method, the gains are still there. The token generation gets about 2.7 times faster because of lower VRAM consumption. Mark Horowitz from Stanford put a 16-bit multiply at roughly 37 times the energy of an 8-bit addition, but on a GPU that saving is pretty pointless because the multipliers are firing just as often as before.
So, is this Monopoly's being Monopoly's? Well, sort of. Nvidia went low-bit, too, but only down to NVFP4, where its multipliers are actually being used as it's supposed to. Their Rubin GPU started shipping to OpenAI, Microsoft, Google Cloud, Meta, and CoreWeave this summer, each rated at 50 petaflops of 4-bit floating-point, which gets most of the memory savings without changing how the tensor cores work. And when you own roughly 80% of the AI hardware market, whatever format your chips runs basically becomes the format everyone else's software gets tuned for.
But even with the perfect chip, there's still the RAM consumption the weights don't even cover. The context is still stored at 16 bits by default, so for a typical 8B model, a 64,000 token of context eats over 8 GB on its own, which is more than the model itself. Speaking of long context, Bonsai 2 also drops nine points on a test about reasoning through long stories. And one more thing, at the time of recording, every Bonsai 2 score I told you in this video has been given by prism ml themselves.
We can't be so sure that these are actually going to hold up in independent tests. So are 1.58 bit llms actually better? Well, of course they are at least for users without expensive GPUs with massive vram capacity. I mean, what other option do you even have? Those brain dead models that have been compressed so hard they can't even remember their name? These models exceptionally retain almost all the intelligence and are far cheaper and superior to the models quantized below 4 bits by conventional means.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.