Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Pico AI Lab · @PicoAILab
Words
3,274
Runtime
20:24
Speaking pace
160wpm
Reading time
14min
160 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Training a big language model gets all the attention, but here's the thing almost nobody talks about. Training happens once, inference happens billions of times a day, and in 2026, for the first time ever, companies are spending more money running models than training them. That flip is a huge deal for anyone building with AI because it means the hard problem has quietly moved. The hard problem is no longer how do we make the model smart? The hard
80 words, the words spoken in the first 30 seconds at 160 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 217 |
| Average words per sentence | 15.1 |
| Longest sentence | 40 words |
| Questions asked | 3 |
| Sentences containing a number | 26 |
Most used terms
Filler phrases
23 in total: like 10 · actually 8 · kind of 2 · basically 1 · literally 1 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Training a big language model gets all the attention, but here's the thing almost nobody talks about. Training happens once, inference happens billions of times a day, and in 2026, for the first time ever, companies are spending more money running models than training them. That flip is a huge deal for anyone building with AI because it means the hard problem has quietly moved. The hard problem is no longer how do we make the model smart?
The hard problem is how do we serve it fast and cheap without setting a pile of money on fire. So, in this video, I want to walk through why inference is actually difficult because once you see what's going on under the hood, a lot of confusing stuff suddenly makes sense. Why your local model takes forever to start, why the first token feels slow, but the rest stream out fine. Why a model written mostly in Python can beat one written in C++.
All of it connects back to a few core ideas, and by the end, you'll have a mental model that actually holds up. Quick definitions first, so we're on the same page, and I'll keep this short. Training is when you show a model massive amounts of data and adjust its weights until it learns patterns. It's expensive, it happens in big bursts, and when it's done, you have a frozen set of numbers. Inference is everything after that.
Every time you send a prompt and get tokens back, that's inference. Think of it like writing a cookbook versus running a restaurant. Writing the cookbook took years and cost a fortune, but it's done. The restaurant has to serve thousands of customers every single night forever, and every plate costs real money. The cookbook is a one-time capital expense. The restaurant is an ongoing operational expense that scales with every customer who walks in.
That's exactly the relationship between training and inference, and it's why the economics of AI are really the economics of inference. Now, here's where it gets interesting. When you download a model, say an open weights model from Hugging Face, you don't get a program. You can't double-click it. What you get is a folder full of files. The biggest one is usually a safe tensors file, and that's just the weights, billions of numbers sitting there doing nothing.
Next to it, there's a config.json that describes the architecture, like how many layers, how many attention heads, the vocabulary size, that kind of thing. And there's a tokenizer file that explains how to chop text into tokens. So, what you actually downloaded is a parts list and an assembly manual. You still need an engine to assemble the thing and run it. That engine is the inference engine, and there are a bunch of them: vllm, exllama.cpp, tensorrt, tllm, tgi, ollama.
They all take the same parts and build the same model, but they make wildly different decisions about how to do it. And those decisions are where all the performance lives. And this brings up a fun puzzle. llama.cpp is written in C++. vllm is mostly Python. If you've ever seen one of those programming language speed charts, C++ crushes Python by a hundred times or more. So, llama.cpp should be way faster, right? Except in a lot of server scenarios, vllm wins.
That single fact tells you something important about inference. The bottleneck is not the language. The actual math runs on the GPU inside compiled kernels either way. Python is just the traffic cop directing work, and a traffic cop doesn't need to be fast. It needs to be smart. The engines that win are the ones with smarter scheduling, smarter memory management, and smarter batching. Keep that in mind, because everything I'm about to show you is really about one resource, and it's not the one most people think.
Let's start with just getting the model off your disk. Say the weights file is 15 GB. Before anything can happen, those weights have to move from your SSD into memory, and then usually into GPU memory. The lazy way would be to read the whole file into RAM, make a copy, and load that. But, that doubles your memory use for no reason, and it's slow. So, most engines, especially llama.cpp, use a trick called memory mapping.
Instead of eagerly loading everything, the engine tells the operating system, "Hey, here's where the weights live on disk. Just pretend they're in memory, and pull in the pieces when I actually touch them." The operating system handles it lazily. Pages get loaded on demand, and if RAM gets tight because you've got 40 browser tabs open, the system can quietly evict some pages and reload them later. The penalty for reloading is small.
Over a modern NVMe drive at around 7 GB per second, even pulling back 5% of a 15 GB model is maybe a tenth of a second. That's why llama.cpp can have a model answering questions within seconds of you hitting enter. vLLM takes the opposite bet. It can take minutes to start up, and that's not because it's badly written. It's because it's doing a ton of work up front, compiling optimized kernels, pre-allocating GPU memory, setting up its scheduler so it can juggle hundreds of simultaneous requests later. llama.cpp optimizes for you, one person on a laptop who wants the first token fast. vLLM optimizes for a server that's about to get hammered by traffic all day, where a 3-minute startup is irrelevant.
Same model files, completely different philosophy. And neither one is wrong, they're just answering different questions. Okay, model's loaded. Now the part that I think is genuinely the most important concept in this entire video. When a transformer answers your prompt, it works in two phases and they have opposite personalities. Phase one is prefill. The model reads your entire prompt at once, all tokens in parallel, and builds up its internal understanding of the context.
Because everything happens in parallel, the GPU gets to do what it loves, giant matrix multiplications at full tilt. Prefill is compute-bound, meaning the limit is raw math horsepower. Phase two is decode, and this is where the model writes the answer, and it can only do it one token at a time. Generate a token, feed it back in, generate the next one, over and over until the response is done. Each step is a tiny amount of math, but here's the killer.
For every single token, the GPU has to read essentially all of the model's weights out of memory, tens of gigabytes of reading to produce one token. Then it does it again and again. So decode is not limited by compute at all. It's limited by how fast you can move bytes from memory to the chip. That's called being memory bound, and it leads to a stat that honestly changed how I think about all of this. Over the decade from 2012 to 2022, GPU compute power grew roughly 80 times.
Memory bandwidth grew about 17 times. People call that gap the memory wall, and it keeps getting wider. Picture a chef who got 80 times faster at cooking, while the person carrying ingredients from the pantry only got 17 times faster. The chef just stands there waiting. During decode, the most expensive chips on the planet are mostly standing there waiting. On a typical setup, GPU utilization during prefill can hit 90% then crater to under 30% the moment decode starts.
You're paying for the whole chip and using a third of it for most of every request. Once you understand that, almost every inference optimization you've ever heard of snaps into focus because they're all just different ways to fight the memory wall. The first weapon is the KV cache. During attention, the model computes key and value vectors for every token it has seen. Without a cache, generating token number 500 would mean recomputing all that work for the previous 499 tokens, which would be insane.
So, the engine stores those intermediate results in GPU memory and just reuses them. The catch is the cache grows with every token in the context, and long contexts eat enormous amounts of VRAM. A long conversation can have a KV cache that rivals the size of the model itself. That's the real reason context windows are finite, by the way. It's not some arbitrary limit, it's memory pressure. VLLM's signature trick, paged attention, treats cache memory like an operating system treats pages in small blocks that can be allocated and freed without fragmentation.
That one idea is a big part of why a Python project beat the C++ crowd at scale. The second weapon is batching. Remember how decode reads all the weights to produce one token? Well, here's the beautiful part. If you have 32 users generating at the same time, you can read the weights once and produce 32 tokens, one for each user. The memory cost barely changes, but you got 32 times the output. It's like an elevator. The trip to the 10th floor costs the same whether one person is inside or 10, so you want it full every ride.
Old school static batching grouped requests together and made the whole group wait until the slowest one finished, which wasted tons of capacity because someone asking for a haiku got stuck waiting on someone generating a 2,000 word essay. Modern engines do continuous batching, where the moment any request finishes, a new one slides into its seat mid-flight. The elevator never stops moving. This is the single biggest reason cloud API providers can charge so little per token.
They keep the elevator packed all day, and a solo user on a dedicated GPU simply can't. Weapon number three is quantization, which is just compression for weights. Models are typically trained at 16-bit precision, so every weight takes 2 bytes, but it turns out you can store those numbers at 8 bits or even 4 bits, and the model barely notices, kind of like compressing a photo. Done well, your eye can't tell the difference, but the file is a quarter of the size.
And since decode is all about how many bytes you have to move per token, cutting the bytes in half literally doubles your speed ceiling. It also lets bigger models fit on smaller cards, which is the whole reason a 70 billion parameter model can run on a consumer GPU at all. There's a whole zoo of formats here. GGML for local stuff, AWQ, which uses calibration data to figure out which weights actually matter and protects them, and so on.
The honest summary is that they're like zip versus rar versus 7z. Different recipes, same goal. The detail worth knowing in 2026 is that newer chips run low precision natively in the hardware. Hopper cards do FP8, Blackwell does FP4 formats, and that's not a software hack. The silicon itself does the math at that precision. Hardware and quantization are co-evolving at this point. Weapon four is speculative decoding, and this one's clever.
The painful part of decode is the sequential loop, one token per memory pass. So you take a tiny draft model, something fast and cheap, and let it guess five or six tokens ahead. Then the big model checks all those guesses in a single parallel pass, which it's great at because verification is parallel even though generation isn't. If the guesses are right, you just got six tokens for roughly the price of 1. If a guess is wrong, you throw away everything after the mistake and continue.
For predictable text, and a lot of language is predictable, the speed ups are real. It's a junior writer drafting and a senior editor approving, and the editor reads way faster than they write. And weapon five is flash attention, which doesn't change the math of attention at all. It changes the memory choreography. Standard attention kept writing big intermediate matrices out to slow GPU memory and reading them back.
Flash attention restructures the computation so it stays inside the chip's tiny ultra-fast on-chip memory and never takes that round trip. Same answer, dramatically less data movement. Notice the pattern across all five of these. Cache, so you don't recompute. Batch, so you reuse what you read. Quantize, so you read fewer bytes. Speculate, so you read less often. Restructure, so data stays close to the compute. Every single one is an attack on data movement.
The math was never the problem. Now, there's a newer idea that takes this logic to its conclusion, and I think it's one of the most interesting shifts happening in serving right now. If prefill is compute-bound and decode is memory-bound, then running both phases on the same GPU means the hardware is wrong for half the job, no matter what. During prefill, you're wasting expensive memory bandwidth. During decode, you're wasting expensive compute.
So, big deployments have started splitting them. Disaggregated inference runs prefill on one pool of hardware and decode on another with the KV cache handed off in between and each pool gets sized for what it actually does. Nvidia is even building dedicated prefill chips now that skip the expensive HBM memory entirely and use cheaper memory because prefill doesn't saturate bandwidth anyway. Teams adopting this are reporting two to four times cost reductions.
It's a complete rethink of the idea that a GPU is a GPU. Let's talk about how you actually measure any of this because averages will lie to you. The metric users feel most is time to first token, which is basically your prefill cost plus queuing. If that's slow, the app feels broken even if everything after is fast. Then there's tokens per second for a single stream, which is your decode speed and that's what makes text feel like it's flowing or dribbling.
Throughput is the system-wide number, total tokens across all users and it's what determines your costs. And then there's tail latency, the slowest 1% of requests, which is where code starts and bad scheduling hide. The diagnosis flow is pretty clean once you separate the phases. Slow before the first token shows up means prefill trouble, so shorten prompts or add compute. Slow dribbling after it starts means decode trouble, so quantize, batch better, or get more bandwidth.
Running out of memory means KV cache trouble, so page it or trim context. Different symptoms, different organs. On engines, the practical landscape in 2026 looks roughly like this. For local single-user work, llama and llama.cpp are the easy answer with great CPU and GPU hybrid support where part of the model sits in RAM and part on the graphics card. For self-hosted production serving, vllm is the default choice. Huge community, works across hardware vendors, and its throughput is excellent.
SG Lang has been gaining serious ground, especially for agent-style workloads with lots of structured output and shared prefixes. TensorRT-LLM is the option when you're all in on Nvidia and want every last percent at the cost of painful compilation and setup. And TGI slots in nicely if you live in the Hugging Face ecosystem. The differences between them are exactly the topics we just covered. How they batch, how they manage cache memory, how they schedule.
Pick based on your workload shape, not benchmarks from someone else's workload. Money time, because this is where inference stops being academic. The good news first, GPU rental prices have collapsed. An H100 that cost $8 or more per hour at the peak of the shortage now rents for around $2 to $3 on most specialized clouds, and spot prices go even lower. Hardware got cheap, relatively speaking. But here's the trap that catches almost everyone who tries to self-host to save money.
The economics are completely dominated by utilization. A GPU sitting at 10% load makes your cost per token roughly 10 times worse, and suddenly your self-hosted setup is more expensive than just paying a premium API. The big providers win on price because they have the traffic to keep their elevators full around the clock. You probably don't. The honest math says self-hosting starts making sense when you have sustained, predictable volume or hard requirements like data residency, where legal says the traffic cannot leave your cloud, or latency targets an external API can't hit.
If you have spiky traffic and no special constraints, a managed API is genuinely the right call, and there's no shame in it. There's one more force multiplying all of this, and it's agents. A chatbot does one inference call per user message. An agent doing a real task might do 50 or more calls, reading files, calling tools, checking its own work, looping. Every one of those calls pays full prefill and decode costs. And agent contexts tend to be long, which makes the KB cache problem worse.
So, as the industry shifts toward agents, inference demand isn't growing linearly with users. It's growing multiplicatively. That's a big part of why inference spending overtook training spending this year, and why every optimization in this video went from nice to have to existential for anyone running this stuff at scale. The other direction things are moving is toward the edge. Phones and laptops now ship with NPs, dedicated neural processing units designed for exactly this workload at low power.
A quantized three or seven billion parameter model on device gives you zero network latency, total privacy, and zero marginal cost per token. The pattern I expect to win is hybrid, a small local model handling instant stuff like auto and complete and simple questions, with a big cloud model handling the heavy reasoning in the background. The same prefill and decode logic applies at every scale, from a four billion parameter model on your phone to a trillion parameter model in a data center.
The physics doesn't change, only the budget does. If you want to actually feel all of this instead of just hearing me talk about it, here's a 10-minute experiment. Install a llama, pull a model in two versions, the full precision build and a four-bit quantized build of the same model. Run the same prompt through both and watch two numbers, memory usage and tokens per second. You'll see memory drop by more than half and speed jump with answers that are nearly identical.
Then paste in a huge prompt a few thousand words and watch the pause before the first token. That pause is pre-fill and you'll feel it being compute bound. Then watch the steady drip of generation afterward. That drip rate is your memory bandwidth live in front of you. Once you've seen it on your own machine, none of this is abstract anymore. So here's the whole video in one thought. Inference is hard because language models generate sequentially while hardware got fast in parallel and the gap between how fast chips compute and how fast memory feeds them keeps widening.
Every technique that matters, caching, batching, quantization, speculative decoding, flash attention, disaggregated serving, is a different way of moving fewer bytes or reusing the bytes you already moved. The model is frozen, the weights never change. All the engineering and all the money is in the delivery. If this was useful, subscribe because I'm planning follow-ups going deeper into the serving stack and let me know in the comments which engine you're running and what tokens per second you're getting on your hardware.
I'm genuinely curious what setups people have. Thanks for watching.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.