Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
4:144.8x the video's typical replay level
their translations. It's because they want users that aren't just English speakers. Serve more users in their language at soydev.link/gt. So let's talk about why I think the local model hype is dumb. There's a lot of layers to this but I want to close out the one I was hinting at at the
Said at 4:08
Most replayed moment #2
0:344.1x the video's typical replay level
are so cool and they are essential to an ecosystem to flourish at all. Open weight and open source are meaningfully similar enough that people who think open weight isn't open source, not interested to me. I don't care. Open weight models are essential for us to keep evolving the AI ecosystem and
Said at 0:28
Most replayed moment #3
20:512.8x the video's typical replay level
all of your options. And they're all priced pretty much identically and the performance is relatively similar. Funny enough, it wasn't similar until I crashed out at Azure and now it's a lot better. Apparently, the hosting of Open AI models on Bedrock entirely broken
Said at 20:46
The graph counts replays. It does not show where viewers stopped watching.
Words
5,544
Runtime
28:11
Speaking pace
197wpm
Reading time
23min
197 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
I'm sorry for this one in advance, but I need to crash out a little bit. I'm not proud of the things I'm about to say, but if I don't, I'm afraid nobody will. And I can already tell that this comment section will be a disaster because the last time I talked about this in other places, I got like 50 death threats in a day. The people who I'm about to talk about are genuinely insane. But before we do that, I want to be clear about what we are talking about here and what the scope
99 words, the words spoken in the first 30 seconds at 197 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 395 |
| Average words per sentence | 14.0 |
| Longest sentence | 70 words |
| Questions asked | 24 |
| Sentences containing a number | 104 |
Most used terms
Filler phrases
71 in total: like 47 · actually 11 · kind of 3 · literally 3 · you know 3 · basically 2 · right? 1 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
I'm sorry for this one in advance, but I need to crash out a little bit. I'm not proud of the things I'm about to say, but if I don't, I'm afraid nobody will. And I can already tell that this comment section will be a disaster because the last time I talked about this in other places, I got like 50 death threats in a day. The people who I'm about to talk about are genuinely insane. But before we do that, I want to be clear about what we are talking about here and what the scope of this video is.
Open weight models are awesome. I love them. They are so cool and they are essential to an ecosystem to flourish at all. Open weight and open source are meaningfully similar enough that people who think open weight isn't open source, not interested to me. I don't care. Open weight models are essential for us to keep evolving the AI ecosystem and landscape, especially now that the government is starting to take away certain models.
We need good open weight models. So, what is this video about if I'm pro open weight? I think GLM-52 is awesome, by the way. Like unbelievably good for what it is. It's actually close to catching up to like 5-4 and Opus 4-6-4-7 in an open weight thing you can download yourself. That's crazy. But there's a catch. I said download yourself. I didn't say run yourself. Models like GLM-52, despite being incredible, are nearly impossible to run the full proper versions of. 5-2 is conservatively 400 GB.
So, if you don't have that much VRAM, you are not using the full model. This is a model that runs great on crazy data center setups and it's capable of unbelievable things, but it is not a local model. And if you're curious about the full precision GLM-52 BF16, 1.5 TB. That's not running on anything anyone watching this video owns in their house. And if you can't run this on something in your house, I'm very curious why you're living inside of Colossus 1 or 2.
Even the quantized and pruned versions of models like this are still in the 200 GB range, which isn't running on almost any consumer hardware. And if it is, it's barely going to run. And this is just the start of why I think the local model thing is really overrated and dumb. Local models are not GLM 52. It's things like a quantized Gemma 4 that barely functions. And we're just at the tip of the iceberg here for all of the problems that exist in the local model world.
If you're hearing people talking about how they can replace Codex and Claude code with things running on their own GPUs, you're being lied to, and I'm going to explain why I'll for a quick break for today's sponsor. If I told you that today's sponsor could 4x your potential customer base, you would probably call me insane and a liar. And that makes sense, because I am kind of insane and I am kind of lying. It's actually 5x, because General Translation it makes it trivial to launch your stuff in every single language.
If you're only shipping in English, you're leaving behind 81% of the world. Only 5% speaks English as their native tongue, and only 19% speaks English overall. And a lot of that 19% prefers other languages, too. So, supporting all of them is huge, but it's always been way too much work. Take it from me, I helped build the system for this at Twitch, and it was miserable. General Translation makes it too easy. It almost feels criminal.
You can run their NPX GT at latest command, and it will get you going fast. Or if you're old school like me, you can go look at the code yourself. You install the package in the project for whatever framework or tool that you're using, you import their config, you choose the languages you want to support, and then you wrap your app in their provider. And from this point on, it is trivial to add translation because you use their helper component, T.
You don't have to wrap every string, you just wrap the sections where there are strings that should be translated. And now their systems can yank out all of the copy that matters, generate translations for it, and then give them to your app when you're ready to ship. It all just works as part of your CI, and it couldn't be easier to do. They provide all the annoying components you might need for something like this, like a local selector, or more importantly, the ability to to variable content where things shouldn't be formatted or a number where the formatting varies a lot depending on where you are or even things like date time and currency which are so annoying to get right by hand.
There's a reason that companies like Cursor, Ramp, ClickHouse, Partiful and more are leaning on general translation to manage their translations. It's because they want users that aren't just English speakers. Serve more users in their language at soydev.link/gt. So let's talk about why I think the local model hype is dumb. There's a lot of layers to this but I want to close out the one I was hinting at at the start which is the gap between runnable and good.
The models that you can run on your laptop are incredibly impressive that they could exist at all but they're not doing real work. There are some exceptions here like DS4 by antirez. Antirez is the creator of Redis and he has been trying to build a from scratch C based runtime in order to be able to use the Deep Seek V4 models on his MacBook. It is a native inference engine optimized first for Deep Seek V4 flash with some support for V4 Pro on very high memory machines.
In order to even use flash which remember is the small version of Deep Seek V4, you need at least 96 gigs of RAM on a Mac or an Nvidia CUDA or Strix Halo machine. But Theo, I have 128 gigs of RAM in my gaming PC with a 5080 in it. Do you know how much RAM the 5080 has? Cuz that's what actually matters. 16 gigs. You're not fitting [ __ ] in this. Oh, well, I can upgrade to the 5090. Do you know how much VRAM's in that? 32 gigs.
You still can't fit [ __ ] on it. It does not matter how much RAM you put into your fancy gaming PC. It is not benefiting inference beyond when the VRAM is saturated. Well, why is your MacBook with 128 gigs of RAM better? Because it's also VRAM. It's unified memory. There's basically three real options you can consider as a consumer for getting lots of unified RAM which is useful for this type of thing. Option one, as I said here, is a MacBook with enough RAM.
Would really be a shame if the prices happened to have just gone up 30 to 50% and more than 2x for the more RAM. The 128 gig option is now 3,000 additional dollars on most Macs. So, that's ruled out. Maybe I'll grab a DGX Spark. If you want to run these things at like 12 TPS at best, sure. I have a DGX Spark. It kind of just sits there. I should probably sell it. It's not useful. It's cool if you want to debug CUDA [ __ ] with a lot of VRAM in quotes, and you don't care about the performance.
Other than that, not useful. So, what we're left with is the Strix Halo, which you can get inside of a handful of laptops, as well as within the Framework desktop, which is probably the best option you have if you don't want to go Mac. Those are your actual only consumer options to get enough RAM to even run these things. And since all of those aren't real GPUs, as in not real high-end Nvidia hardware, they're not going to run great.
There's a very interesting problem that I see a lot on my Mac here, because again, I have these things working on my MacBook with 128 gigs of RAM, because it has enough RAM. If you have a model that fits on this and fits on my 5090, it'll run three times plus faster on the 5090. If I take like Gemma 4, for example, it'll run decently on here. It'll run way better on my 5090, because it fits in the VRAM on both. But, if you switch to a big model like Deep See V4 Flash, not even Pro, it'll run okay on my MacBook, and it will run like absolute [ __ ] on my desktop, because it doesn't have enough VRAM, because it's bottlenecked on the GPU.
So, your options are get enough RAM on one of these unified boxes and have way less compute, or you can get more than enough compute with a high-end chip like a 5090, which is as good as an RTX 6000. It is comparable to an A100, which is a much more expensive enterprise chip, except for the VRAM limitations. And Nvidia does that intentionally. The 5090 is just as powerful as an A100 if not better in everything but the VRAM.
So, your options are high performance [ __ ] models or [ __ ] performance slightly better models. But as soon as you're talking about legitimately good models for our real dev work, we are far outside of what you can do on side of consumer hardware in any format. That's when you go from this 96 gig limit up to like 200 plus. And just for reference, the 5090 currently goes resale for around $4300 for just the GPU. That price is a little absurd.
Okay, it looks like they are actually in the four grand range. That's insane. That's absurd. I never paid near that for mine. But if you do need more VRAM, you have an option, the RTX 6000 Pro. The Pro 6000 Blackwell has 96 gigs of VRAM. So, you can now fit huge models like Deep Seek V4 Flash. And ready for the price on these? $13,000. And it's the same performance. It's just more VRAM. The chip on this is nearly identical to the 5090.
And it's 3x more expensive than the 5090, which is already too expensive. Not good. Don't worry though, you can just buy a tiny box. And for only $75,000, you can have four RTX Pro 6000s. This could almost run GLM 52. It'll have to do some crazy things because the VRAM doesn't get pulled as you might hope when you have a lot of GPUs. You need to find some way to replicate some things and split up others. Yeah, good luck.
Have fun. Don't worry though, the red is way cheaper. It's only $12,000. And this one works by using a bunch of AMD GPUs instead of the Nvidia ones. And it's got a whopping 128 gigs of VRAM. So, again, you can use Flash. And almost nothing else. So, hopefully I have helped establish both that the models you're seeing these crazy open weight drops on, things like GLM 5.2, as incredible as they are, they're not running on consumer hardware.
If you want to spend $75,000 on a tiny box, you can run one instance of it at a decent-ish speed. Well, when I hear local model, I think about things I can run on consumer hardware of some form. Even my $10,000 MacBook can barely run a lot of those, and this is better equipped than most machines. So, the first problem I want to make sure we fully understand this. Just because a model is open weight does not mean that you can run it at home.
And the models that are open weighted unbelievable are very different from the models that you can run at home. The other thing, which I just established, I should have put on the list before, hardware costs are insane. If you're buying these types of GPUs, I think it's a better purchase as a thing you plan to resell than purchasing it to actually use. Oh god, do we have cope in the comments? What the [ __ ] is Ornith?
Everyday I get sent some new family of locally runnable-ish models that I've never heard of that I don't know anyone using. I know a lot more people plugging these than using them, and apparently the new one is Ornith, which apparently can get comparable numbers to Opus 47 and 48 in STV bench verified, which is a saturated, abused, entirely like compromised benchmark. Or Claw Eval, I don't even want to know what that is.
When we look at some of these like NL2 repo, we see that it's like half-ish the score. Or if we look at something like Deep SW which is one of the few actually good benchmarks for these things, and we look at the open models they have here like Kimmy, it scored a 30% where 54 scored a 52. Opus on low scored a 41 and was cheaper to run because these open weight models just burn [ __ ] tokens. So, that's the other problem is even if you have them running fast, they're going to be slow as [ __ ] because of the inefficiencies that exist in all of those models.
Cool. Kimmy K 27's like on par with Sonnet 46. Awesome. That's so cool. But, you know what? I'll entertain it. Let's enter a hypothetical world where the local models do catch up to be roughly the same performance as what we get from models like Opus and GPT 5.5 today. Hypothetical, but let's just assume it does happen. That we get to the point where a 30 bill per am model that you can run on most higher-end MacBooks can be equivalent capabilities-wise to those big, beefy Frontier Labs.
Let's just assume it can. I don't think it can, but let's assume it can. That's when you run into the next problem, parallelism. When I'm working, I'm not going from zero agent to one agent, I am randomly fluctuating between one and 40. Each of those are doing their own inference. Maybe I have two threads running in T3 code or Codex. Maybe I have an agent break up a bunch of sub agents to go explore different things in the project.
Maybe I just want to use Claude and Codex at the same time. There's a lot of reasons I'm running more than one agent at a time, and if you set up your system in such a crazy, impressive, capable way where it can actually run GLM 5.2, can it run it twice? Can it run it 10 times? Can it run it as many times as I can run sub agents in my current agentic workflows? I'll give you a hint, the answer is [ __ ] no. It absolutely can't.
This is separate from the other problem, which is that you're just wasting the money when these sit idle. If you buy enough GPUs to run five agents in parallel, any moment you're not running five agents in parallel, you're just not benefiting. Those boxes sitting idle are just money wasted. But, if you do need to scale up to six, and you only bought enough hardware for five, you're blocked now. You got to wait until one of those workloads is done to have that lane available for the next.
And if you want to run two different things at the same time, like you want to run a bigger model for some tasks and a smaller one for others, maybe you have a big model orchestrating and then smaller ones reading through the code and doing subtasks, you need 10 times more VRAM to even possibly enable that workflow and a lot more compute in order to power it. The number of models I am running at any given time varies a lot.
And I'm not able to do that on local hardware at all. And I see a lot of cope and chat about caching and VLLM and batching and all of these types of things. No, none of that actually solves this problem meaningfully. And if you think it does, you don't understand the problem. And again, we are presuming the models catch up to current state of the art when we talk about all of this because state of the art isn't just it can solve this benchmark well, it's it can queue up a crazy workflow that has lots of different stages and manage it as other agents complete those stages and do the whole end-to-end loop.
Part of that end-to-end loop is often computer use. And GLM-52 doesn't even have vision. It's an unbelievable coding model, but it can't look at things. I can't give it a screenshot. And that's state of the art for open weight. Insane. It's still really good. GLM-52 is awesome. I'm so thankful it exists. But you're not running that on local hardware, you're not running that with the workflows that I do every day, and you're not going to write code with it for UI stuff even though it is much better at UI because it can't take the feedback.
It can't even see what it made. It just knows the code. There's also the reality that the best models are getting bigger. Something like Fable is not a one trillion parameter model. It's between two and 10. We have no idea, but it is massive, massive, massive. You're not running that on anything consumer, ever. And even if you can, we have another problem. Electricity costs. Let's ask a good open weight model about this.
GLM 52. How much would running an RTX 5090 24/7 cost electricity-wise in San Francisco? Give it search. Well, at GLM 52, do some math here. The cost for running my 5090 24/7 in San Francisco is around $5 a day just for electricity. That's two grand a year in electricity costs. Do you understand? It's not cheap. Just cuz you own the hardware doesn't mean you're getting away with this for free. That's just one GPU. If you want to run these bigger models, you need a lot more than one.
The costs are crazy. I do see one other type of cope I need to address in chat. Not everyone needs a frontier model for everything they do. I don't disagree, but I think the things you're talking about here that frontier doesn't benefit are not the growing use case. You need to be realistic here. Do you think the growth in token usage month-over-month and year-over-year is mostly going to these small cheap models for those types of tasks?
Tasks that could be solved by small cheap models got solved by small cheap models a while ago. The cool thing about the development of frontier models is they increase the types of work that can be done with models. When new better models come out, my personal token spend goes up massively because all of a sudden more work is able to be done with these models. More ends of the work can be reached with them where I can start at an earlier place before the LM comes in and then grab it again at a much later space after the LM's done more of the work.
And you do need better models for things like vision, for things like accessing APIs to get the feedback from the code reviews and then addressing them. For owning this long process that is doing millions of tokens instead of dozens, you do need things like the current frontier. GLM-52 is the first open weight model that can actually run for a while, and it's really impressive that it can do that. It changes the economics of using models on a fundamental level in a way that's genuinely really cool and exciting.
That's why I love open weight models. But if you think consumers using ChatGPT are where most of the inference that OpenAI is doing is, and not the small number of devs and companies that are cranking the highest end models doing crazy long jobs, I can send a single prompt that does 10 million tokens. No consumer can. So, is it cool that 1% of token use could theoretically be on a local model that runs on your laptop or phone?
Yeah, until you consider the fact it's going to kill their battery and cause the thing to overheat. Like Mark says here, things like Gemini 3 and fitting on tiny tensor chips on your phone is really cool because it can do message summaries and weather summaries without having to go to the cloud, which is a real valuable use case. Not needing to have your data, which is fully unencrypted once it reaches the servers, by the way, go to the server to serve the inference.
That is awesome. And I do want ways to run models without having to send the data to another company or cloud, which is why I am excited about some amount of this getting better. And like the ways Apple does things on your phone using a little bit of local intelligence to not have to use a server, as well as the secure compute stuff, which is even cooler. There is a problem, though. Your phone only has so much power.
This both limits the capability of the model that you're running on your phone, but it also means that you're going to destroy the battery on your phone in really egregious ways, which would be much nicer if you could just hit an API. But Theo, computers are getting so much more powerful. Phones get better every year. Aren't they going to get good enough that you can run basically anything on them? I'm going to ask you a question.
How much faster do you think the highest end Android phone is today than it was 3 years ago? Better question. How much faster do you think a $200 Android phone is today than it was 3 years ago? Just get like a rough number in your head. Here are the numbers. Sure, iPhones continue growing and it looks like Xiaomi has come in and actually made some meaningful improvement on the Android side, which is cool to see. That was not the case before.
When you look at the cheap devices here, there has been literally no improvement in performance. If anything, there has been slight regressions in the cheaper phone tier. And that's before the price hikes because of RAM and demand and all of the manufacturing stuff getting [ __ ] This is going to be way worse in 2026. If you're just counting on our devices to get powerful enough to do this, here's a chart that tells you you're wrong.
It's not happening. Budget device CPUs have not meaningfully improved since 2022. We're not anti-open weight models, we're just annoyed at local models and the people who think they are the future. This message from Fry nails this. I don't think the major benefit of the open-source models is to run them locally. They enable competition on the provider and hardware side. Hmm. No provider can offer Opus at a cheaper price, even if they have spare margins, because of licensing with Infrobic.
Hmm. There are providers running on places with cheaper electricity that offer cheaper inference of 5.2. Competition is a win. Ding ding [ __ ] ding. You can see the real benefits of this when you go somewhere like OpenRouter. Let's look at GPT 5.5 to start. We have options here. We can use Azure, we can use Azure in Europe, or we can use Open AI's models, or or we can use Open AI's hosting. There are your options. There's all of your options.
And they're all priced pretty much identically and the performance is relatively similar. Funny enough, it wasn't similar until I crashed out at Azure and now it's a lot better. Apparently, the hosting of Open AI models on Bedrock entirely broken right now, so I was right in my prediction that uh it was not going to work on Tranium. I called it, everybody said I was wrong. I was [ __ ] right. I'm going to take that win.
I'm going to be proud of it. But we aren't here to talk about closed models. We're here to talk about open ones. And just cuz they're called open AI doesn't mean their models are open, sadly. Let's look at 5.2. Oh. That's a lot of options. And they have different pricing. Some of them are way faster and cost more money. Like Wafer Fast at 115 TPS and 1025. Or Fireworks Fast, which has a little bit of reliability problems, but still is only 660 out and 130 TPS.
That's awesome. Or Friendly, who apparently is doing 117 TPS for the same 440 that a lot of the others are doing. But then there's companies like Deep Infra that'll go even cheaper. They're a little slower at 30 TPS, which is a compromise that makes sense for some people, but you can choose to be cheaper there. You can choose to spend a bit more to get more reliability or more speed. You have options, and you have a lot of them.
This is where open weight models really shine. They allow for competition in the hosting space, which allows for the most of the problems I was talking about before to be solved. Let's like go back to this list again. Okay. Gap between runnable and good. Doesn't matter when a cloud can host it for you. Hardware costs are insane. Doesn't matter when you're effectively renting. Parallelism. Doesn't matter at all when you're renting.
Electricity costs factored into the token costs you're paying, and they can bring these GPUs and these data centers to places where power is cheaper. Almost all of the problems I have with this local model [ __ ] are solved by just throwing them on a cloud host. You do lose one of the coolest benefits, which is that no one else can see what you're doing. It's just yours. And I'm really excited for Secure Compute to make that a little more viable in the future.
But for now, the best part of these open weight models is that you can get them for good prices with incredible performance on various different hosts. And believe me, I use those a lot. I think it's awesome that you can take these incredible developments in the model space that are all open weight, all available to download and use and retrain and do whatever you want with it. And then just run them in the cloud. It's great.
It's awesome. If your work can be done by these types of models, like if you're doing work that works in Opus or Sonnet, but also works fine in GLM 5.2, and the failure rate isn't a meaningful difference, you should probably move that workload to 5.2. It's insane how much faster and cheaper it can be. There is one last catch though, because those prices aren't the whole story. Deep Infra has GLM 5.2 at $3 per million out versus Opus 4.8, for example, which is 25 per mill out.
So, obviously, that's way cheaper, right? 3 to 25, that's a huge discount. It's almost 10x cheaper. What if I told you it was nowhere near that much cheaper? In reality, Opus on X High is about $8 per run with Deep SWE, and 5.2 on Max about $4. So, it's only a 2x gap. How the hell does that make any sense though? Isn't it literally 10 times cheaper per token? Here's the problem. Open weight models tend to be a good bit heavier on the amount of tokens that they burn during their runs.
Opus 4.8 High more efficient than 5.2. Every single GPT 5.5 run, even X High, used less tokens than GLM 5.2 on High. This is the problem. These open weight models are not as efficient as the frontier models. So, even if the price per token is way cheaper, the massive increase in number of tokens is rough. And on top of that, the speed is worse because if it has to generate three times more tokens to get you an answer, doesn't matter if it's 20% faster if it has to generate 3x more tokens.
This is the problem that kills 3.5 Flash, by the way. It is absurd how many tokens it uses. Those asking, "Why is the chart backwards?" It's cuz top right good. You want to be up and to the right. That's how people read charts. Most of these types of charts tend to go that direction. One of the few sources that does this chart the other direction is artificial analysis. So, the top left is the good area. Do you notice something?
Nothing's in the good area. 31 Pro preview just barely touches the corner there, but there's very little that is efficient cost-wise and also smart intelligence-wise. Once I clean up the chart a bit here, things get much, much clearer. GLM 52 Max is actually very interesting where it is slightly cheaper than Gemini 35 and also slightly smarter. That is good. That is awesome. That is why we like open weight models. This is a real competitive dot in this chart.
But, watch what happens when I turn on GPT 55 on, I don't know, medium or low. Huh. Looks like 55 medium is also neck and neck with GLM 52 and is cheaper despite being way more expensive per token. And I'm not saying this to [ __ ] on GLM 52. It is an incredible model and I'm so thankful it exists. It's even better at UI than GPT 55 is somehow, which is crazy. They killed it with this model. It's really goddamn good.
But, it's not this magical solution people seem to think when they stare at the prices here and think that's all that matters or they hear they can run it locally and they get excited about that and then they try to and realize they literally cannot do that. Think I've said all I have to here. This is a long overdue crash out. I have been annoyed about this stuff for a long time. Open weight is awesome. The competitive marketplace is necessary.
If we don't have good open weight developments, we're not going to have an ecosystem that has any incentive to keep getting better. Drops like Deep Seek R1 shifted the whole industry forward in a really powerful way. And I do not want a world where we are not trying to make the best possible open weight models as an industry. I just think it's insane to believe you're going to run these things on your local network on your own GPUs and get performance even close to reasonable, much less comparable to what we get from these frontier labs.
And if you take this video and frame it as Theo hates open weight or open source, I really hope someone clips this part and shows it to you because I love open weight models. I built T3 chat because of how much I love the deep seek line of open weight models. I am a huge fan of open weight. I just think it's delusional to pretend anyone's going to run good frontier open weight models on their own hardware. And I wish you guys would stop pretending this is a realistic path forward when it obviously is not.
So please, local model people, stop hurting the open source and open weight movement by pretending these things can run on your computers. We need to be able to talk about how good open weight models are without tricking people into thinking they can run them on their own systems. They can't. It's stupid. It's unrealistic. It is really fun. If you're just doing this because you like the idea of having a bunch of GPUs in a cluster and playing with these things, awesome.
That's so cool and fun. And you should talk all about it. But if you're pretending this is a thing everyone needs to do in order to get out of the pockets of Anthropic and OpenAI, you're insane. And I really hope you stop. >> [sighs and gasps] >> I've said all I have to here. This was a very, very overdue crash out and I am sure the comment section isn't a disaster at all. Let me know how you guys feel and until next time, peace, nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.