Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

DIY Smart Code · @DIYSmartCode
Words
460
Runtime
2:20
Speaking pace
197wpm
Reading time
2min
197 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Five names keep coming up when people run AI on their own computer. llama.cpp, Ollama, LM Studio, vLLM, and Lemonade Server. Four of them share one motor, one is not built for you at all. Three words first. A model is one file, 4 to 30 GB in the GGuf format. Quantization squeezes it down, so a model that needed 40 GB fits in 12 for a small quality cost. And graphics card memory is the ceiling. If the file fits there, answers come fast. If not, the rest spills into system memory and everything crawls. llama.cpp is the
99 words, the words spoken in the first 30 seconds at 197 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 41 |
| Average words per sentence | 11.2 |
| Longest sentence | 26 words |
| Questions asked | 1 |
| Sentences containing a number | 4 |
Most used terms
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Five names keep coming up when people run AI on their own computer. llama.cpp, Ollama, LM Studio, vLLM, and Lemonade Server. Four of them share one motor, one is not built for you at all. Three words first. A model is one file, 4 to 30 GB in the GGuf format. Quantization squeezes it down, so a model that needed 40 GB fits in 12 for a small quality cost. And graphics card memory is the ceiling. If the file fits there, answers come fast.
If not, the rest spills into system memory and everything crawls. llama.cpp is the engine, a C++ program that reads the model file and does the math on your processor, your graphics card, or both. It runs on Windows, Mac, Linux, even a phone. Almost every friendly app here calls it underneath, and those GGuf files you keep downloading are its format. Ollama is that engine behind one command. You type Ollama run and a model name.
It downloads the file, loads it, and drops you into a chat, leaving a server on port 11434 that your editor and your scripts can call. The price is control. You get the defaults until you go digging. LM Studio is the same engine with a desktop app around it. You browse models inside the app, and it flags which ones fit your machine before you pull 12 GB you cannot run. The server tab speaks the same API as Ollama, and the console shows your tokens per second.
Free at home and at work now, and closed source, unlike everything else here. vLLM came out of Berkeley. It wants Linux and data center graphics cards and answers many people at once. Its trick is paged attention, many conversations packed into the same card memory instead of a reserved block each. On one laptop for one person, it buys you nothing. Lemonade Server is AMD's answer. Same shape as Ollama, but it routes the work to whichever chip on that machine is fastest.
The processor, the built-in graphics, or the NPU in Ryzen AI laptops that most tools ignore. Open source, same API. It runs on Nvidia and Intel, too. Four things go wrong whichever you pick. A model too big for your card still runs, just painfully slowly. A tiny model is immense more. Everyone pulls gigabytes per model and none of them ship a password. So check what the port is bound to. Start with LM Studio to see what your machine can handle.
Move to a llama once your scripts needed. Use it for a Ryzen AI laptop and lemonade wins instead. Skip VLLM until a second person needs your model. Which one is running on your machine right now? Drop the name, nothing else.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.