Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Tech With Tim · @TechWithTim
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Tech With Tim's most watched videos.
Most replayed moment at 3:14
2.3x that video's typical replay level
So all the model is doing is just giving text to a system that says, hey, I want to call this thing and it doesn't. Now that hand off is essentially the whole way that AI agents work at scale. The model decides what to do, but the code that's running this model is
Said at 3:10
The graph counts replays. It does not show where viewers stopped watching.
Words
5,626
Runtime
24:27
Speaking pace
230wpm
Reading time
23min
230 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Everybody is talking about running AI locally, but almost nobody explains what it is or how you actually get it to work. Now you get these videos where people throw around terms like weights or quantization your VRM graph, and it sounds like you need a PhD in a $10,000 computer just to try this out. And that's exactly why most people give up and just keep paying for subscriptions like ChatGPT. So let me give you the honest one sentence version here. That is that local AI is simply a model file that's sitting on your computer and a program that runs it. That's it. No cloud, no API keys, no internet, no subscription.
115 words, the words spoken in the first 30 seconds at 230 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 366 |
| Average words per sentence | 15.4 |
| Longest sentence | 54 words |
| Questions asked | 11 |
| Sentences containing a number | 31 |
Most used terms
Filler phrases
106 in total: like 48 · actually 30 · right? 8 · kind of 7 · literally 5 · basically 4 · you know 4.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, published by the channel, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Everybody is talking about running AI locally, but almost nobody explains what it is or how you actually get it to work. Now you get these videos where people throw around terms like weights or quantization your VRM graph, and it sounds like you need a PhD in a $10,000 computer just to try this out. And that's exactly why most people give up and just keep paying for subscriptions like ChatGPT. So let me give you the honest one sentence version here.
That is that local AI is simply a model file that's sitting on your computer and a program that runs it. That's it. No cloud, no API keys, no internet, no subscription. Everything else is really just a small detail. So in this video, I'm going to break down what's actually happening under the hood, the real building blocks and none of the fluff. And then I'm going to show you how to run a model on your own machine at four completely different ways from an app that you download all the way to pure low level code where you're running it yourself.
Now, by the end of this video, you're going to understand locally I better than most people that are posting about it. So let's dive in. So first let's clear up the most common confusion that I see. Now that's what's actually different between something like ChatGPT and a local model. Now when you use ChatGPT Claude Gemini, really any of these tools, what's happening is that you type a message. That message leaves your computer.
It travels over the internet to something like a data center, where a massive computer that you don't own runs a giant model, and then the answer gets streamed back to your screen. So your computer did basically nothing. It's really just a window into somebody else's machine. Now, locally, I completely flipped the model. So the actual file that contains all of the intelligence gets downloaded onto your computer. And when you ask you to question your own CPU or GPU is the one that's doing the work.
So nothing leaves your machine. And that gives you three big advantages. Now first, it's private. That's because your data never goes anywhere. It stays on your computer. Second, it's free because there's no subscription and there's no per token cost. And lastly, it works offline. So if you're in a plane or a coffee shop with terrible Wi-Fi, doesn't matter, you can use these local models. Now, I do want to be honest here about the trade offs, which is that the models that you're going to be able to run at home are much smaller than the frontier models that you'd use from something like Claude or OpenAI, but they've gone shockingly good over the last couple of years.
And for a huge amount of everyday tasks, even coding tasks, they're more than enough. So with that said, let's have a look at the actual pieces here. So you understand local models much deeper. Now the first part here is the model itself. And I want to be really clear about what a model actually is, because this is where some people imagine some kind of magic black box. Now a model is literally just a file. It's a really big file.
Could be hundreds of gigabytes. That's full of numbers that we call weights. Now these are billions of numbers. They got baked in when the model was trained at that file. Doesn't think it doesn't run. It just sits there on your disk like any other file that you would have. And companies like meta, Google, Alibaba, Mistral, whatever they release these files for free. Now those are your open models and things that you keep hearing about, like Lama or Gema or Quen or Deep Secret.
You get the idea. You can literally just download them because they are literally just a file full of billions of different numbers. Now, when you go to look at these models, you're going to see names like 4 billion, 8 billion, 70 billion. Now that B is billions and it stands for the number of parameters. Now that's effectively just how many numbers are inside of that file. And the rule here is pretty simple. Generally speaking more parameters means a smarter model.
But it also means a bigger file that needs more memory and compute to run. So if you look at an 8 billion model, this is maybe a few gigabytes in storage, while 70 billion model is a file so large that most laptops simply cannot even load it in, some may not even have enough storage to download it. Now, that leads directly into the biggest trick that happens with local AI. Now this is something called quantization. Now it sounds scary and really complicated, but it's really the same idea as compressing a photo.
So you can take these billions of numbers and you can store them with less precision. And that means that this model file is going to get dramatically smaller with barely any quality loss at all. So a model that would normally need 16GB of memory in its original form might only need five, six, or seven gigabytes after quantization. And when you see the term GIF floating around, that's basically just the standard file format for these compressed models.
Now this is the entire reason that normal computers can even run AI models at all, or at least some of the bigger ones. So when you hear about quantization or a quantized model, just think of that as a compressed model that's meant to make it smaller so it's easier to run. When you compress these models, you keep almost the exact same performance, but again, you just reduce the size drastically. So it's a lot easier to actually run.
Then the next piece is the inference engine. And this is the part that almost nobody explains. Remember that the model is just a file full of numbers, and a file can't run itself. So you need a program that can actually load these numbers into memory and do the actual math to make the model work and predict the next token. Now that program is called an inference engine. Now the most famous one is called llama Dhcp. And here's a secret.
Almost every tool that I'm going to show you today. So LM studio, a llama Docker model runner, they're all basically just a wrapper around engines like this. Now the engine does all of the work and the tool just kind of makes it nice to use what you're going to see later on. Now the last building block here is your hardware, because there's really only one question to decide what local models you're going to be able to run.
Now that is how much memory do you have on your computer and how fast is that memory? Now on a PC with a graphics card. So a dedicated graphics card, that number is going to be your Vram. So if you have an Nvidia 39 TI or 49, you're 59, you're I'm just naming random GPUs. You're going to look at the Vram on that device. Now, if you're running on a modern Mac computer like an M3 and M4, really any M-series MacBook, then you're just going to be looking at the amount of Ram that your computer has.
And that's because Apple shares its memory with the GPU and has something called unified memory. Now, other devices have different specifications, but generally speaking, if you have a relatively new computer, if it has a dedicated graphics card, you're looking at the Vram. That's the amount of memory you're going to have for running local models. And if you're on a mac, again, a modern one, you're looking at the amount of unified memory.
And the rule of thumb here is pretty simple. The model file needs to be able to fit inside of that memory that you have, with a little bit of room to spare. So roughly speaking, if you have eight gigabytes of Ram or unified memory, you're going to be able to run 3 to 4 billion parameter models even without being quantized. And if you have 16GB of memory, you can go up to 7 or 8 billion parameter models. And then if you go up to 32GB, you start to be able to get into the 14 to 30 billion parameter range.
And this is where things start to feel genuinely smart, especially for local models. But keep in mind, you don't need a monster computer to do any of this. Even with something like your phone, you can run small models already. And one thing that carries over from cloud AI is that the context windows to the model short term working memory is going to actually affect the amount of space that's being taken out. So if you have long conversations, bigger documents, etc., that's also going to fill up your memory.
Now I'm going to explain this a little bit more in detail because this is super important. But the basic idea is that the one number you need to be aware of is how much memory either Vram or unified memory is on your computer. When you look at a model, whether it's quantized or not, it needs to be able to fit comfortably within that memory range. So if you have 32GB of Ram and a model is 25 gigs, that's fine. You'll be able to run it, right.
However, one thing to keep in mind with memory is also the speed of the memory. So while you will be able to run much larger models on things like modern Macs that have, you know, 128GB of unified memory, the inference speed of those models is going to be much slower than a similar model that's running on something like a dedicated GPU or an Nvidia RTX GPU. The reason for this is the speed of the memory. So while the memory will dictate the size of the model, and the more memory you have, the smarter models you can run.
The memory speed matters for the tokens per second, and the inference speed that you're going to be able to generate. There's a lot of different techniques here and things that I could get into. But generally speaking, if we talk about dedicated graphics cards again, typically in the Nvidia family, these are much faster, sometimes 2 or 3 times faster at inference speed, but they usually have less capacity. So for example, I have 24GB of Ram in my 4090.
And it's very fast and can generate, you know, 200 tokens per second for some of the models that I run. However, I can't run models that are 70 billion parameters like I might be able to on my Mac. However, on my Mac, those models are really slow because the memory speed is significantly slower. So you're going to be looking at memory speed as well as memory capacity. Those are the two things that are going to dictate what you can do with locally.
And there's always going to be a trade off in terms of the size of the model and the inference speeds that you're getting. Typically, smaller models are going to be much faster. And again, the most important thing is that whatever model you run, it needs to fit in this memory. Well, it still can run if it's not in there. It's going to be so slow that it's practically unusable. Generally speaking, if you're looking at models between 14 and 35 billion parameters, those are going to be a really good sweet spot.
They're going to give you pretty decent performance, and you're not going to feel like you're missing out on too much. If you go up to huge models like 120 billion parameters, 250 billion parameters, you're talking about needing extremely high end hardware running at slow inference speeds. And it's very difficult to actually run those at scale on your own machine. Anyways, with that in mind, let's keep going here. I want to tell you about something really interesting.
Now, the whole reason you'd even bother running models locally comes down to one main thing, which I know you all want, which is control. You pick the model, you own the set up, and nobody can change the deal. But here's the problem right? Almost every tool that you'd actually want to use locks you into one provider's model. And that's exactly what Minds Hub Co-Work, who's the sponsor of today's video, is built to fix.
Now it's open source, free to use, and it has a real model route. So you can run GPT Gemini or the same local models that I've been talking about right here, all inside of this workspace. So the workflow is simple. You brief the built in agent harness, which is Anton, walk away and come back to finished work. Now, I asked it to research the latest coding models and build me a comparison dashboard. And this is what it came back with.
An actual dashboard that I can open and share, not just a wall of text. And this is the part that connects to everything in this video. When a better model drops, whether it's local or cloud model. I can swap it in and I don't have to change anything. So I have the same workspace, the same work. And since it's fully open source, you can clone the repo, spin it up locally in just a few commands, or just download the dedicated Mac or Windows app.
The whole thing runs on your own hardware, which is basically the end game of what we're testing today. Now they also have a hosted version, but honestly, I love the desktop app as it's very easy to use. So I'm going to leave a link to it in the description. Try it out again. It really goes nicely with these local models. And now let me show you how we can actually run local models. So there's a lot of ways to run a local model.
And just like anything in software, it really comes down to how much control you want. So I've broken this into four different tiers. Now at the top we've got LM studio. This is a regular desktop app. You can click, you can download, you can touch everything. And you don't really need to go into the terminal. Now below that we have a llama. This is a really popular option especially for developers. And it's just one command inside of your terminal where you can talk with models, spin up a local server.
You get the idea. Then we have Docker model runner. Now this is really good because it treats models like containers, which is perfect if you're actually going to be deploying these alongside live applications. And at the bottom of my list here we have full code. This is where you're running a model in pure Python. And you see every single piece. Now no matter which way you want to run these models here, you're going to be making three decisions.
You're going to pick a model. You're going to pick a size and quantization that fits inside of your memory, and you're going to decide how you want to talk to it. So whether that's a chat window or something like code, if you keep that in mind, every single one of these tools is going to make sense. So let's start at the top. And again I'm going to go through all of them and show you exactly how to run local models. Let's dive in.
So the first tool on my list here is LM studio. I'll give you a quick walk through. But this is one of the best ways to run local models. Now, once you download the tool again, it's completely free. You're going to be brought into a view that looks something like this. From here you're going to go into the model view. It looks like kind of a robot icon on the left hand side. And you build a search through all of the available models you can download directly here.
Now what you'll want to do is search for a model that matches the kind of relative size or amount of memory that you have. Again, if we're talking about larger high run machines, you can typically get away with 27,000,000,035, you know, 30 billion parameters, etc. if you have eight gigs of Ram or 16 gigs of Ram, look for ones that are 8 billion parameters or 4 billion parameters, much smaller sizes. So for example, we have Gwen 3.8.
You'll notice that if I click on this I can view different levels of quantization and see the change in size right here. You also see kind of some icons or indications of which model is the best for you based on your hardware. So you can see when it says full GPU offload possible and a little thumbs up. That's the one that you would want to go with. And notice this number of Q4. That's the level of precision or the quantization level.
So the lower so like Q4, Q2, Q1, the more quantization you have, right? So if you have QA, you can see this is bigger than Q six or Q4. It's a pretty drastic difference. You also going to want to look at the capabilities of vision tool use reasoning if you need it, to be able to analyze images, you need vision if you want it to work in an agent harness or an agent mode, you need tool calling. Hopefully you get the idea.
There's so many models I can't possibly go through all of them, but you get the idea. You can browse through here and look at the ones that are going to match your specific hardware in sizes. That makes sense. Now, once you download the model, you can view your models from the model tab right here. Now, in order to use these models, you do need to load them. So if I select a model like GM, A4 and I just bring open the sidebar here, it will give me some options for actually loading and running my model.
So what I may want to do here is go to the load tab and start changing some of the values. If you're a beginner, don't change anything and just run this directly. The one thing you can have a look at is the context size here. And keep in mind, the larger you make this context size, the more room is going to be taken up in your computer's memory. Because all of this context needs to actually fit in the GPU memory, right?
Or the memory that you have for running local models. There's a bunch of other settings you can use, but in this case, what we're going to do is just load the model. And when we load the model, same thing. It asks us for the settings. We're going to go ahead and load that. It will take a second. And then we will be able to view that here from this terminal view, and also chat with it directly from the chat window. You can load multiple models at once, and you'll be able to see the models that are loaded up here, as well as the size and then to check them.
So you can see that I'm currently using 5.58GB of the 63GB of Ram that I have. However, we're talking about GPU memory here, so that's not really 100% accurate. So if we go here to the terminal, we can now see that this model is loaded. I can view all of the API stuff for this in case I'm a developer. And I want to directly chat with it using something like a curl command. If that doesn't make sense to you, don't worry.
If you just want to chat with the model, you can go over to the chat view. So from here we'll press new chat. We're just going to select the model that's already loaded to Google GEMA for. We can very modify things here. For example like the system prompt if we want to do that. And then we can just start chatting directly with the model just like we would inside of something like ChatGPT. Now you can see this one is extremely fast, right?
We're getting 120 tokens per second because it's very small. And again, I have high memory bandwidth because I'm using a dedicated GPU. You can load multiple models as long as they all fit into the memory here. And then again you can adjust all of the parameters. And if you want to you can start using them from this server, which is useful especially for coding. And you can see the full logs of everything that's gone on.
Tokens per second speed. You get the idea. LM studio is very good. There's a lot you can do with it, and if you want a full tutorial, leave a comment down below and I will go into it. So the next tool on my list here is a llama. Now this is a little bit more popular for developers and it does a very similar thing to LM studio. However, it's a little bit less visual and gives you a bit less control. Now, a llama is a very popular way for downloading and running local models.
In order to use it, you do need to download the tool, so you can just go to olamide.com, and it will be available inside of your terminal as a command. So if you type a llama in your terminal and once it's downloaded sorry, you should see something like this where you can launch it for all kinds of tools. Or you can directly chat with different models now, as well as the terminal or CLI based tool. There is also a visual tool that you can open when you download the desktop application.
From here again, you can launch a llama inside of any of these harnesses and use models that you've downloaded. You can go into the settings, right? Or you can actually just start chatting with different models by selecting one of the ones that you have. Now. In order to download models in a llama, what you're going to do is start by finding the model that you want. So you're going to have to go to the llama hub. So by doing that, you can go a llama and then models.
From here there's a bunch of models that are available for a llama. Same thing. You can search through them. You can ask ChatGPT to help you find one. And if you find a model that you want. So actually let's go. Maybe Nemo Tron 3.5 lightning. Here we can see all the different sizes. We now understand what B stands for and what quantization is, right? So we can have a look at them here and we can pull them directly inside of a lot.
So the way this works is the following. First you can type a llama list. If you type list this will show you all the models you currently have downloaded. And then if you want to pull a model, you'll type o llama pull and then the model ID that you found from the model hub. This is going to download it to your computer and then allow you to start using it. If you want to run a model, you can type a llama run and then put the model ID.
So I'm going to put Nemo Tron three like this and it will start the model. It will load it into my computer's memory. This is why will always take a second at the beginning, because it actually needs to load it. Then you can start chatting with it directly from this view. So here you can see I can type something like Hello World, and I can just directly start using this local model. That was well as that I can do that from this terminal view.
So if I go to let's go Nemo, Tron Nano or something, I can type. Hello. Same thing. We need to wait for it to be loaded and then it will give us a response. Sometimes it takes a second, especially on the first load, but you can see we get the thinking and then we get the response. And if we go back here. Hello. How can I assist you today? Now this is great, but a llama will also expose all of its services on an API. So for example, if I type a llama help, you're going to see an option of all of the different things that you can do.
As long as a llama is running, it will actually serve all of its models available on a default port. I don't remember exactly what the port is, but I believe it's something like 11,434, which means that you were actually able to send curl requests and use a llama from other tools. As long as it's running in the background. I'm not going to go into a full tutorial of it, but if you're a developer and you understand what a Rest API is, a llama provides that already with access to all of the models, it will automatically load any model that you ask it for.
Whenever you try to send a request to it, it actually has what's called an open AI compatible API, which means that you will be able to send requests in the same format that you would to something like ChatGPT or Anthropic. Anyways, that is a llama. Let's go to the next example. So the next tool on my list here is the Docker model runner. Now this is available as an experimental feature inside of Docker desktop. There is a bunch of restrictions with it.
However, if you are going to be doing this on a Linux machine specifically and you have Nvidia hardware, it works very well. It can work on CPU as well I believe. However, it's extremely slow. So with Docker model Runner, if you go into Docker desktop, there is some settings I believe you need to enable this experimental feature. You should see this model start from the models tab. You'll be able to go to Docker Hub and then here you can pull all of the same type of models as you would be able to inside of like a llama or LM studio.
Once you have a model here, you can chat with it directly. So I have geometry for example. Same thing. It will automatically load the model for me and then I can type something like hello. Now similarly to all of the other tools, this will also expose a rest API on a different port. I believe it's 12,434 or 343. That will allow you the ability to chat with these models without being directly inside of this interface.
You can also inspect the model, see all the information about it, etc.. What's interesting about the Docker model runner is that it actually treats models like containers. What that means is that you can write Docker files, you can write compose files, and you can actually have models shipped directly with your applications and be dependencies exposed through Docker kind of services, which is a lot more complicated than I'm going to get into in this video.
But if you do use Docker and you're familiar with this and you use it for your apps and you want local models, this is a really good way to deploy them. Now I also show you that there is a CLI based tool. So similarly to what we had before. If I type something like Docker model, you can see that we can configure, inspect, install the runner, push arm, view the models, you know load unload. You guys get the idea. And you can view models directly inside of here as well.
So this is a really powerful feature. And if you want a full tutorial on it I have actually done that on my channel. You can see the easiest way to run alarms locally. Docker Model Runner tutorial goes through all of the features, and we'll even show you all of the Docker files and how to set it up with the automatic deployment. Okay, that is Model Runner. Now let's move to the last one which is full code okay. So the last example I have for you is actually running models using just code.
So this means that we're actually going to bring in our own inference engine, in this case llama CP, and not rely on something like a llama or Docker model runner or LM studio to do this for us. Now the big surprise is that llama CP is the engine that pretty much all of the tools that we just looked at are already using. But if we want to invoke it directly ourself, we can do that. So for example, you'll see we have this Gwen 2.5 model which I've downloaded locally on my computer.
Again, this is literally just a file that contains a bunch of numbers. Now if I want to run this normally I would need a llama or something like that. But I can actually write code that when voted directly for me. So you'll see that I can just load the model. I can then create a response using this package. And if I just run the code here, you will see that I get the following in my heart. I run AI on my desk running free and it wrote me right a haiku or whatever you call this about running AI on your own computer.
Now I can change this prompt to be hey, who is Tim or something? I don't know if that's going to give us anything meaningful, but let's run this. And Tim is a character from popular video game whatever, right? Like so this is a very small model. Of course, it's not going to give us good responses, but you get the idea. We can run it fully locally. Now, one thing to keep in mind is that as well as doing this, we can actually chat with models that are running on our own computer through services like a llama.
So like I was mentioning, if a llama is installed and running, you can specify the model that's actually available in a llama that you've downloaded. And then similarly to before I can run something like this, in this case it's going to be a bit slower because llama two is much larger. And you'll see that we actually get the response right. And it says, hey, someone might choose to run an AI model locally, blah blah, blah, blah.
And it's using that old llama back end service. And if I wanted to, I could even change this to the lm API, right? Or LM studio API, or the Docker model runner API, and do the exact same thing right from code. So this is kind of the more manual method, but most developers are going to end up managing their models through something like a llama and then invoking them in code using a method like this. Okay, so that wraps up the demos.
Now let's talk about which method you should actually use. So here's my honest take. If you just want to chat with a model and you never want to see a terminal, then you can use something like LM studio. It's genuinely one of the easiest ways to download models and has some of the most amount of features. Now, if you're a developer and you want a model running on your own machine, that your scripts and apps can talk to, you definitely use a llama.
That's what I reach for most days and it works really well locally. If you're already living inside of a Docker container and you want model sitting in that stack right next to you, then use Docker Model Runner. It's really good in production if you're actually building and deploying things. And lastly, if you want to understand everything that you're doing and run models yourself in probably the most efficient way, then you can use your own code to do so.
Of course, you don't need to use Python. This is just a quick example. And with that in mind, just remember that all of these tools at the end of the day are using the same building blocks that we talked about earlier. They have a model which again, is literally just a bunch of numbers in a file, and they have an inference engine and a bunch of other fancy features on top of it. If you understand that, you understand local models, and hopefully this video helped get you off the ground and running your first one on your own device.
Anyways, guys, that's all that I have for you. If you enjoyed, make sure you may like subscribe and I will see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.