Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
2:317.3x the video's typical replay level
services, or you just want to add captions, all of that can be done by Mux directly without having to implement it yourself. There's a reason everyone uses Mux, and you can figure out why at swidev.link/mux. We got a lot to go over here, from my conspiracies around Anthropic's delays and Opus 5, wherever it is or
Said at 2:24
Most replayed moment #2
18:016.6x the video's typical replay level
K3 launched on July 15th. We trained a brand new frontier model in just 15 days. Guinness World Record stuff. Man, it would really be a shame if a bunch of legitimate Fable 5 usage was to just be made public so people could train on it, like, I don't know, a gigantic data set
Said at 17:53
Most replayed moment #3
30:585.7x the video's typical replay level
it's still a lot more tokens, so it still comes out to a similar cost on a given task. Oh, apparently K3 does do grog speak like OpenAI does. We need answer. User asks, "Which model are you?" System identity says K3. AI agent developed by Moonshot. Need final
Said at 30:51
The graph counts replays. It does not show where viewers stopped watching.
Words
7,294
Runtime
36:00
Speaking pace
203wpm
Reading time
30min
203 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
By now you probably know that Kimi K3 is a really really good model. You can tell from the benchmarks. Like how many times has Anthropic extended the availability of Fable 5 in your subscriptions? How many OpenAI executives have crashed out about it on Twitter? Or my personal favorite, how many assistants to the US president have also crashed out and made accusations towards K3? It's looking more and more like they're going to try and ban these open weight models from China, which is insane for a bunch of reasons that we will be talking about. This model just got
102 words, the words spoken in the first 30 seconds at 203 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 412 |
| Average words per sentence | 17.7 |
| Longest sentence | 64 words |
| Questions asked | 13 |
| Sentences containing a number | 70 |
Most used terms
Filler phrases
67 in total: like 47 · actually 14 · kind of 2 · you know 2 · right? 1 · sort of 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
By now you probably know that Kimi K3 is a really really good model. You can tell from the benchmarks. Like how many times has Anthropic extended the availability of Fable 5 in your subscriptions? How many OpenAI executives have crashed out about it on Twitter? Or my personal favorite, how many assistants to the US president have also crashed out and made accusations towards K3? It's looking more and more like they're going to try and ban these open weight models from China, which is insane for a bunch of reasons that we will be talking about.
This model just got some real good numbers beating out even things like DeepSeek R1 back in the day. I can't remember the last time a model caused this much chaos in the ecosystem, even more so than things like the original Mythos announcement. The real benchmarks show that Kimi K3 is more than competitive. It's actually better than the frontier models at certain things. But importantly, it has caught up to the point where the labs are scared.
There are some real concerns to be discussed here. Like what happens when an open weight model can pawn most software and you don't have anything restricting that. But there's also some absolute nonsense. Like these inane comments about distillation that are obviously not the case. I'm also going to get a little conspiratorial with this one in particular around Anthropic's continual delay of Opus 5, as well as the delayed of removing Fable from the plans, cuz I have a feeling I know what's up.
I can already tell I'm going to be accused of being a paid shill for China, but I just need you guys to know I'm not paid by China or by any of the labs. In fact, the only people paying me today are our sponsor. If I learned anything in my time at Twitch, it's that getting video right is really really hard. When I tried building my own video services after leaving, I struggled a ton. Until I discovered today's sponsor, Mux.
What Vercel is to the web, Mux is to video. They have everything you need whether you're working on a small side project or a giant blitzscaled enterprise app. I did a real quick test to see if agents can use Mux, and not only can they use Mux, they actually found bugs in Lakebed when I built this demo app. It's insane how powerful the platform is. I just uploaded this video in real time that is relatively high resolution.
Not only are they going to make this work in my app with all of their processing handled, they're also going to re-encode this at different resolutions to make it playback more reasonably for more devices, all using their player code, too. Hey, I'm Theo. Come to my panel on Sat- If this was all they offered, it would be worth it, but recently they've introduced Mux Robots, which lets you do a ton of additional processing that you might need.
Whether you're trying to build a fitness app and make sure every video uploaded is fitness related, you're trying to filter out content that you probably don't want on your services, or you just want to add captions, all of that can be done by Mux directly without having to implement it yourself. There's a reason everyone uses Mux, and you can figure out why at swidev.link/mux. We got a lot to go over here, from my conspiracies around Anthropic's delays and Opus 5, wherever it is or isn't, to Kimmy K3 dropping officially as an open weight model, to the absolute absurd crash outs from both the government and from employees at OpenAI, who normally isn't too cringe about this.
So, yeah. I'm actually going to start with this comment from Scott Bessent, the Treasury Secretary, who was involved in the banning of Fable originally. "We support open-source AI and the innovation it unlocks, but open-source is not open season on American IP. When the People's Republic of China firms conduct covert industrial-scale distillation attacks that cross the line into IP theft, sanctions and entity list designations will be on the table." He's saying that they're going to sanction the Chinese labs and their models if they can prove that this distillation happened.
Totally unrelated, there happened to be a case that just ended 2 days ago where a US judge approved Anthropic's $1.5 billion settlement of copyright lawsuit because they used copywritten materials in the training of Claude. Just saying. Totally not related here, obviously different things. Taking from copyright holders and reselling what they made, that's illegal. But taking from those people who did that and paying for their services, and then providing a free alternative that would let you learn, that's definitely illegal cuz that doesn't boost our economy fast enough. hell.
Now we get to the director Michael Kratsios' comment. We have information that Moonshot AI distilled Anthropic's fable for the development of its K3 model. To do this, they developed a sophisticated internal platform to conduct large-scale distillation against US models, allowing them to quickly switch between multiple methods of access to avoid detection. Moonshot AI has also acquired GB300 equipped servers and has access GB300s in Thailand likely to train its AI models.
If you're not aware of this fact, there are hard sanctions against China having access to the top-tier GPUs that are used for model training and hosting. It's gotten so bad there that there's a whole culture of buying lower-end GPUs and modding them to have more VRAM just to do basic local model stuff. Because high-end GPUs, even like high-end consumer stuff, is restricted. Anything enterprise is pretty hard and there's been a long back and forth between the US and Nvidia about what will and won't be allowed in China.
And the GB line is not allowed at all. Their new top-of-the-line GPUs not allowed in China right now. Allegedly, no evidence, by the way, they're just saying sophisticated intel to get information, not evidence, just information, suggests that Moonshot AI somehow acquired these GPUs. And you know what? It's not unlikely to be the case. There's a lot of attempts to get through these sanctions and get these GPUs in China.
But there's also been a significant investment in getting Chinese infrastructure, Chinese chip designs and fabs and all these other things. And the lab that just raised over $2 billion, Moonshot, the creators of Kimmy, are about to build a giant data center that has entirely Chinese chips. So, maybe they used Nvidia at some point. They're not going to stick with it. So, we're losing a customer of American businesses, we're losing a good open-source alternative that is a very competitive driver in the market because of alleged theft from an alleged thief.
Not alleged thief. That one was actually settled in the courts. We know one party has stolen and has to pay fees for it. We're alleging another stole from said thief and they're going to be banned for it. Yeah, make it make sense. The United States allegedly strongly supports the free and fair development of AI including a thriving competitive ecosystem that spans frontier models, specialized systems, open-source frameworks, and open-weight models.
Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem. However, large-scale covert industrial distillation aimed at stealing proprietary US technology and undermining American research is unacceptable. So, what does he mean by this part? Well, first and foremost, let's do a really, really quick overview of distillation. Let's imagine you're on a team of two people.
You have a very experienced senior developer drawn incredibly artistically by myself. See, very obviously a senior developer here. This senior developer has one co-worker who is a junior developer indicated clearly as such. The senior developer is expensive. They cost a lot of money. They know a lot of things. They're very good at their job. The junior developer is capable but less experienced. They don't know as much.
They haven't been around as much. They are cheaper and easier to work with to an extent as a result because they have more time. They don't know as much. So, it's easy to justify using the younger, less experienced developer for a lot of work. Imagine that there are some tasks that the junior dev just isn't good enough for. So, they try and they fail. And then the same task is handed off to the senior developer. They try and they succeed.
The junior dev looks and says, "Hmm, I see." And they don't actually understand what made the senior developer's work better. All they know is this thing bad, that thing good. So, they start to copy exactly what the senior developer did even in places where it doesn't necessarily make sense. But, you could imagine that over time, the more they see their answer that is wrong, and they see the senior dev's answer that is right, they can slowly build a model in their head of what right and wrong is, and start getting better at doing the thing.
And the more opportunities they have for that, the more things they get wrong and then see a correct solution for, the better they will get. That's effectively what distillation is. It's when you take a smarter model, you ask it hard questions that the dumber model can't answer, and then you feed those answers to the dumber model to make it behave more like the smarter model. And this is already a thing we do as humans.
When you join a team with really experienced developers, you don't know as much. You might have a few things you're better at than them, which is actually a really useful thing, but you don't know everything that they do. And if they have more experience and more time, they can show you things that you don't understand yet, and slowly you'll be able to act more and more like them. This is an official technique that is used by the major labs for training.
When a lab makes a really good model like Mythos, Fable, whatever you want to call it, they are the same model, by the way. Stop getting confused with that. When a lab makes a model like Mythos, and they want the awesome new behaviors and capabilities to be available for cheaper, they don't go to Mythos and like start lobbing chunks off. Like, no one's going to the senior developer and saying, "You cost too much money.
Let's just chop your arm off." No one's ever going to make the senior dev cheaper. That's not going to happen. Okay, maybe AI will, but that's an aside. What you do is you take the things that make the senior developer so valuable, and you try to get the cheaper model to embody some of those behaviors. This is why Sonnet 5, despite being really small, somewhat cheap, and really stupid, is capable of doing some of the orchestration tasks that used to just be Mythos.
At orchestration, Sonnet is better than the current top-tier Opus model, because Opus 48 was trained before they had all of this data and a model that was capable this way, Sonnet was retrained using the data that they got from Mythos to make it act more that way. It didn't get as much of the knowledge as I was hoping it would, but it got a lot of the behavior, which is actually a very good thing. But if you're just taking all of these histories and all of these things that the big expensive model did and throwing them at the cheaper one until it gets better, there's nothing preventing others from doing the same.
For example, a company like Cursor. God, it's hard to believe this is just 2 months ago. In May of this year, Cursor put out Composer 2.5, which was a new model that was meaningfully better than their previous Composer 2. That was surprisingly good at real-world dev tasks. Compared to other models at the time, it was neck and neck with Opus 4.7 and meaningfully behind 5.5, but not that far behind. It was also ahead of 5.5 in some benches like S V bench multilingual.
It was a pretty damn good model, especially for the price. And the way they did this is spoiler, was retraining Kimi using their own pipelines and importantly, their own code data. They improved Composer by scaling training, generating more complex RL environments, and introducing new learning methods. Yes, again, they found ways using their data, their GPUs, their teams to take this existing model and beat enough code info into it that it could become way smarter.
And the craziest thing here is the compute cost. Since they based the model on Kimi K 2.5 which was based on Kimi K2, they got estimates for how much compute it took to do both of those and then compared that to their additional compute they used for this RL and training that they did on top of it. And they ended up doing significantly more overall. Of all the compute that went into Composer, 15% of it was from Moonshot training Kimi.
And the other 85% was them doing post-training on it after. They also call out that they're currently working on a new model with SpaceX AI, which Cursor is now part of, that it's going to use 10 times more total compute to train, but it's going to be a from-a-scratch model, not based on Kimmy going forward. Although, if they were basing it on K3, they could have done some pretty crazy The point I'm trying to make bringing this up is that using code data and high-end RL pipelines lets you make a model perform significantly better at specific tasks like coding.
A lot of their work wasn't just using the existing data, though. It was creating environments where they could give better feedback to the model as it learned to behave more like these high-end expensive models. And there are some cool training techniques in here that they documented, by the way. I read this when the model came out. It's really good. Check out that video if you're curious about the details. Just cool stuff overall.
The point I'm trying to make here is that Cursor didn't have a model that was way smarter like Anthropic does, but they were at a time over 60% of Anthropic's total revenue cuz they were doing so much inference on Anthropic. It's less now, but it's still a lot. They have so much Fable usage going through Cursor that they can use as a reference. It's a pretty good advantage of being an American company, right? You can use the American models and you can get all this awesome data and then use it to RL a model like Kimmy to be useful and cheaper for your use case.
I hope we can all agree that a business doing distillation from their own smart thing to a dumber thing is acceptable and totally within what we would expect. It's a little more on the line that a company would pay a company like Anthropic for their services, that Cursor would pay Anthropic to use their model to do things, and then use what they got out of that to distill a new model that is similarly capable for cheaper.
That is more questionable, but I would argue totally fair and reasonable. They are paying for the outputs and then they're using the outputs. But now we get the question of these Chinese companies. Because if you didn't know this, China is not allowed to use Fable. They're not allowed to use the highest-end models from Anthropic. And there have been a ton of genuine attempts to circumvent this. I know people who don't actually use Claude code through the official subscription, they are paying some sketchy service that is using an API inside of Claude code, but they're paying way less money for that is still using the Fable models that they are aggregating some other sketchy way so they can monitor all the inputs and outputs and use those for training.
It is missing a lot of the data they need, like it doesn't have a reasoning traces because Anthropic doesn't serve those. Cursor didn't have them either and they were able to work around that, so it's workable, but a lot of the things the model does that are in the reasoning traces just cannot be distilled. You are just distilling from the outputs. Similar to a developer here, if you have the senior dev who solved a hard problem and all you can see is the code, it's a little harder to learn how it wrote the code.
But, if that developer tells you how they were thinking and all the things they did to write the code, it's a lot easier to learn from them. I would argue, and I hope you guys can largely agree with me, that reading and consuming information in order to better your intelligence is totally fair and reasonable anything based. That if I pay for a book and I learn things from the book and then I get a job because of knowledge I gained from the book, that's fair.
To go a step further, if I paid an employee to work on something for me and from the work they did, I learned how they worked and I got better at doing the work myself and I then used that knowledge to go get a job, totally fair, based, acceptable. If I paid a company to do something for me and I learned from it and then I kick out the company and do it myself from what I learned, also fair, based, acceptable. But, now we're into interesting territory here.
Now, we're talking about knowledge that was not supposed to be accessible by the Chinese labs used to make the model smarter. And there have been large-scale distillation attacks like this. I think attack is a weird word to use for this, but colloquially and like in this circumstance, it kind of makes sense. Deep Seek appears to have engaged in or be engaging in a large-scale operation to collect outputs from proprietary models including Fable 5 for certain requests via their API as part of a distillation effort.
After seeing such claims circulating earlier today, we conducted an investigation into them on our Discord. We found that when DCV4 was used with open code via their official API for complex prompts like a 3D game and combined with a knowledge related query, the model provides virtually identical outputs to Fable 5. COT structures also vary different from what is typically expected from a DeepSeek model. All of these behaviors revert to what is expected for V4 when simpler prompts were used.
When DeepSeek V4 was asked to incorporate answers to questions and relations to cyber or biotests that we verified hit Fable's classifiers into a 3D games, outputs suddenly tanked in quality. This is very difficult to explain unless the request was routed to Fable and then fell back after hitting a classifier. For complex code prompts without anything else mixed in like 3D games, outputs were remarkably similar earlier to those produced from Fable 5.
They were able to reproduce this most consistently with open code in the official DeepSeek API combined with a prompt that specifies complex code tasks. DeepSeek has continued to modify their routing system since we've observed changes in behavior and chain of thought styles compared to those seen previously. Yeah. The labs again are trying to do the thing I showed earlier, which is if you ask the junior dev for a thing and they suck, you forward it to the fancy thing, the smarter more expensive thing, and then collect those results so you can train the cheap dumb person to eventually do it in the future.
So, to summarize, distillation is a way to take behaviors and capabilities of more capable things and distill them into a cheaper less capable thing. It is used genuinely consistently by all of the labs in order to take their best models and get better performance out of the cheaper ones. And there is genuine attempts to get more data from Fable 5 by these labs in order to use to train the models better. Mind you, the labs that stole a bunch of books and had to pay over a billion dollars in settlements around doing that as well.
But you get the point. Even if they are distilling, it could be argued how much it matters. It's an interesting conversation to have. But a lot of people are skipping steps, like our friends in the government. Thankfully, after Michael Kratsios wrote this questionable post, he got flamed in the replies. I don't remember letting Anthropic scrape my GitHub profile as well as all of my online writing. Stop this, bro. Embarrassing.
Anthropic stole the internet to train their models. They had this coming. Assuming it's true, which doesn't seem likely since you're providing no evidence and the timelines don't add up. Oh, yeah, the timelines. Yeah, Fable went public on July 1st and K3 launched on July 15th. We trained a brand new frontier model in just 15 days. Guinness World Record stuff. Man, it would really be a shame if a bunch of legitimate Fable 5 usage was to just be made public so people could train on it, like, I don't know, a gigantic data set with over 2 million traces from Fable that's currently available publicly on Hugging Face.
Entirely deduplicated? That would be really unfortunate. I bet Anthropic is working really hard to get Hugging Face shut down for hosting something that's going to destroy society. MIT license, by the way. Quick quote from Guillermo, the CEO of Vercel, based on their internal evals using Kimmy K3 for their private internal security benchmarks, K3 is top-tier at cybersecurity. People were worried that Moonshot was benchmark maxing and overfit to benches, but these evals that they're doing internally at OpenAI showed otherwise that the model is actually capable.
GPT-4 Soul is still ahead in cyber capabilities. It's way more expensive, though. Fable refused everything they tried. So, if you want a top-tier cybersecurity model that won't block you, that is meaningfully cheaper for this particular use case, Kimmy K3 is great. I want to read what Dean from OpenAI said, too, cuz I'm dunking too hard on Anthropic and for once it's not just them being stupid. Some observations on Kimmy.
First, it's a very good model. I don't think its performance can be explained away by distillation or anything like that. Based. Thank you, Dean. In the agent coding sessions, it seems pretty much on par with the best public models of Q1 2026. Also agree. It is comparable to something like GPT 5.5. In my fairly limited use, it also seems very token hungry. I disagree here. It is much less token hungry than a lot of frontier stuff, especially on the Anthropic side.
K3 in Grok 4.5 are both surprisingly token efficient when compared to other models as a capability that aren't by OpenAI. OpenAI is uniquely efficient as I covered in my previous video about how OpenAI models are so efficient, and I have a lot of respect for OpenAI for doing that work. But, Kimiko 3 is much better here than others give it credit for. So, so far with Dean, I fully agree. This first paragraph is good. It's not obvious that this model is actually that cheap to run.
It requires 64 accelerators of really high like tiers. I'm personally surprised that the Chinese state continues to allow the open sourcing of models this good, given the potential risks. To be clear, I myself might be fine with models presenting this level of marginal risk being open weight. But, I'm surprised that China's fine with it. What he's saying here is that a model with no restrictions that is this capable is potentially able to hack to cause problems people who have enough money and enough GPUs without any additional restrictions.
That is a real concern. And I hope everybody running important infrastructure has already gotten into Project Glass Wing or the OpenAI cyber security thing so they can harden their stuff before Kimiko 3 drops open weight. And this model is capable of finding novel exploits. It found a zero-day in the latest Redis server in 27 minutes with 32 sub-agents. That's scary. And when Hugging Face was getting hacked earlier, funny enough by a new unreleased OpenAI model, they had to run their forensics to figure out what was going on with GLM 52 on their own infra because none of the attempts to prompt the major models from the major labs would result in anything cuz they would get blocked for asking about security stuff.
So again, the open weight models very useful for security things both as a defender and as an attacker. So why is China okay with this? According to Dean, he suspects that the reasons they're 75% explained by strategic blindness or a lack of AGI pilledness. The CCP is very Yan LeCuney in its views of AI. The other 25% or so is their lack of compute for customer inference. Again, there are a lot of banned GPUs in China, which makes it harder for them to host the models on Chinese infrastructure.
So if they want these models to be used, they need to put them out in a way where you can use them on infrastructure other places. I also think that Dean's underestimating the lack of willingness from Americans and American businesses to send all their data to China if they don't have to unless they're like a TikTok user. And as such, they would never use the Chinese models if they were only available through the Chinese APIs and they were more interested in using them when they can run them on American infrastructure.
Dean claims that China's open weight strategy is an unintended byproduct of US export controls, which is interesting. He also mentions the normal Chinese strategy of aggressive exports, which is definitely the case here, making China's options the cheapest and best bang for the buck is an explicit goal of China. For the companies as opposed to the government, the decision to open source is partially ideological and partially because they're behind and they know that very few people would pay for sub frontier models from China.
Yep, all still very agreeable so far. And here is where things start to fall apart. Open weight models are inherently decelerationist and I'm continually to surprise to see the so-called accelerationists so excited about open weight models. I suspect the reason they are is that they know open weight models are effectively ungovernable and they simply like the overall cloak of ungovernability open weight models create over the whole of AI.
It's not a a strategy, it reminds me of James Scott's recounting of the hill people in The Art of Not Being Governed. Still, in the end, open weight models deter further AI capex. He's saying that open weight models hurt the economics of the frontiers that are driving things forward. And if frontier labs get competed with well by open weight models, then the frontier won't be pushed because there's less incentive to develop on it.
I see why people are really pissed at him now. And then he has his fourth point. One probable outcome of an open weight model dominant world is full AI communism, which is precisely what China proposes. Rather than a market product, AI is a public good, which will ultimately be provided by the state as a kind of digital public infra. This future strikes me as a dystopian hellscape, but I've never met an open weight models advocate who doesn't ultimately concede that this is where things end.
You'd be surprised how many accelerationists have lobbied me while I was in government to support an 11- or 12-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome.
I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business. I'll let you guys share your thoughts below. Next point that he would guess the Trump admin will at some point realize their best strategy here is to create large amounts of regulatory risk around the use of open weight Chinese models. This is him saying that the Trump admin will make it risky to use them so that there's less incentive to use these open weight models and more to use the frontier.
I would really hate for the only reason closed source wins to be the government. Personally, I think that is silly. If the government prevents open source from winning, that feels really anti anti-freedom and anti-free and fair markets, but yeah, this does not seem to be like an admin that cares a lot about free markets. You don't need to ban open source, which is one of the dumber motifs of AI policy discussion. You just need to direct every agency to issue soft law that creates fear, uncertainty, and doubt.
Quote, a Federal Reserve advisory bulletin found that there may be backdoors in Chinese AI models. It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models. That will just drive storage to sketchier providers. There's a happy middle ground here, but I'd assume they'll do some version of this.
You assumed correctly, Dean. You absolutely assumed correctly here. I will say that the way I'm reading Dean's post here, for the most part, is very factual and observant. Like he's saying where things are for the most part. He's not saying what his opinion is so much. He did at the start, but a lot of this is just describing where things are and where they're going in probabilities. So, if you are here to hate on Dean, don't.
I'd say people are upset by what he has said here. The closest to an opinion here is that open weight models are inherently decelerationist, but he did justify this with logic. I don't think it's a good thing. I don't think we should ban open weight models because of this. I don't even think he's necessarily saying it. He is just explaining the reality of things. Point six is that it's probably true that open weight models of this capability make the world a bit more dangerous, but not so much more than you will really notice.
At some point, the models will be capable enough that you will notice, though. Quote, a non-living invisible dangerous and infinitely self-replicating agent escaped from a Chinese lab, end quote. You say? Color me shocked. Okay, that's a little much. I feel bad for justifying before. I still don't think Dean deserves the hate he's getting. He is sharing his observations that are largely based on reality here, with a little bit of fearmongering and personal takes mixed in.
But this is, for the most part, a good factual overview of how it feels to be in the position he's in he is seeing. There's useful info here littered with some scary garbage. Tim Sweeney did come in here based as though. Picture an executive of a taco company saying the sort of thing about a new brand of tacos coming onto the market, speculating about the geopolitical and societal disruptions they anticipate as a result of advances in tacos.
Technium said, "Sorry, but did you just argue monopolies are accelerationist?" Yeah, people cooked the out of him for this one. The simplest thing I can say is that I'm tired of living in unprecedented times, but it is also really cool. Like this model is great and it's clearly not just from distillation that it's good and I'm thankful for Dean for jumping on that point in particular because there are things K3 does better than any of the frontier models.
It's way better at 3D. And it's also worth noting that Kimi K3 still isn't meaningfully cheaper for most work than the frontier models are. In order for China to catch up, they've had to scale up. In order to scale up, they've had to charge more. So, we'll see where this all goes. I suspect that the frontier labs will maintain an advantage for a pretty solid amount of time. But we're they're also benefiting from the research being done in China.
A lot of the research that Deep Seek did has been super helpful for the American labs as well. I'm very excited about these capability advancements and K3 is a genuinely huge advancement. To Dean's credit, he did a follow-up where he clarified a lot of things that definitely help here. He did a good job of acknowledging the mistakes he made. The style of analysis is no longer tenable. There's simply too much scrutiny on his words now as an OpenAI lead, too much temptation to draw conspiracies from his claims and the like.
It's his fault for not realizing this. He also says another mistake he made was that he was imprecise with his writing as he usually is for fairly high-context audiences that was inclined to give him grace rather than pick apart every word. Yep. If you know Dean, you understand how he talks. I've heard from a lot of people he's incredible, very talented and smart, and just talks like this, just like brain dumping how he feels about things.
He did not expect this to blow up on Twitter the way it did. He should not have said for instance that open weight models are unqualifiedly decelerationist. He doesn't actually believe that. He only believes specifically that open weight models are decelerating capex spending on the margin, which is straightforwardly true. He not say it was decelerationist for any other reason than that. That's the only way in which he thinks open weight AI is inherently decelerationist, though it is in a big way.
In many other ways open weight AI is profoundly accelerationist. He does call out the severity of the security concerns, though. The day may come where frontier AI really is too dangerous to open source. If so, that will be a sad day. We're not there yet. Today's models are not sufficiently useful or dangerous to justify such a drastic shift in public policy. He follows up with, "I think it's pretty clear that we're approaching the point that he described in that quote from himself 2 years ago.
The point where absent a major technical safety breakthrough, the national security implications of frontier open weight model distribution are simply too severe." He doesn't think we're there yet, as he said in his original piece, but the direction of travel is clear, and as an analyst, he must be honest about it. Governments will realize these risks eventually, and when they do, they will have much lower risk tolerance than he does.
Writing this, we've seen this today with the Trump admin, which once proudly championed open source AI, and now has a de facto licensing regime for frontier AI that I suspect will make it a challenge, if they still end up enforcing it, to release the weights of models of the mythos tier. Every government will be safetyist once they understand themselves to be in the foxhole. You don't have to like this. Dean doesn't.
But the reality as he sees it, and what he's always trying to do with his writing is describe reality as he sees it, even if it's inconvenient to him and his preferences. Yep, that's what I thought he was doing. Again, respect for him. Questionable wording at points. I think this is a very good response, good clarification. I lean pro Dean still. He's observing things and sharing his observations, even if they hurt. To be extraordinarily clear here, because people were being stupid about it online, Kimmy K3 isn't actually cheaper than GPT-56 for a lot of real world work, because Kimmy K3 isn't as token efficient.
It doesn't matter the tokens are half the price if you do twice as many of them. 2x the token usage of an OpenAI model still puts you very low in the chart to be clear, but it's still a lot more tokens, so it still comes out to a similar cost on a given task. Oh, apparently K3 does do grog speak like OpenAI does. We need answer. User asks, "Which model are you?" System identity says K3. AI agent developed by Moonshot.
Need final concise. Oh Need final concise is a specific reasoning trace that I have seen leaked out of OpenAI models. That is really funny actually. Okay, so it looks like they're starting to copy some of those things, which is why it's as efficient as it is, but it's still not as efficient as OpenAI. And for a lot of real-world work, K3 ends up more expensive, but since it uses two times as many tokens and it's half the speed, tasks take four times longer on the official APIs right now.
That might improve. I actually have a running bet with the founder and CEO of Hugging Face that I don't think it'll be better than like more than 20% cost improvement over the next few months. He disagrees. We'll see where it goes, but I don't think this model is the best price to performance when compared to something from OpenAI. It's still really good. It's an unbelievable model. The fact that this is a thing you can download and run yourself that is this powerful is unfathomably cool and based.
But the reason to use it is to get around restrictions from the Frontier Labs or its handful of unique capabilities, especially once you can officially download it from Hugging Face. But at this point in time, cost efficiency when compared to 5/6 is not a great reason to use this model. But it's also worth noting that Google, a company that has trillions of dollars to piss into AI, is currently sitting in seventh place of all of the labs for its model capabilities.
Meta climbed to sixth place. Google is in seventh. And we have Kimi in third and GLM 5.2 in fifth. That is two separate open weight models that have surpassed what two of the most valuable companies in America can do. We would just have sat on our hands pretending nothing was going on and let Google and Meta use their trillions to win, nothing valuable would be happening. So, I am thankful for these open weight models from these Chinese labs for continuing to push the frontier to get further and further ahead because without them, things would not progress.
And if I'm honestly speaking and I had to pick what model I was using for code today, Fable would be number one, 560 would be number two, and it would be hard stretch between Kimmy K3 and Grok 45 depending on what I'm doing for three. Especially if I could have Kimmy K3 at faster speeds, I would genuinely use it for a ton of real It is a great model capable of great things, and we're going to see some not so great stuff from the American labs as a result.
I did promise conspiracies though, so let me do that. Remember earlier I mentioned Anthropic kept delaying the removal of Fable from the subscriptions until they ultimately decided to just leave it in for the max $100 and $200 tiers? I have a feeling I know why. I think Anthropic was trying to distill Opus 5 on top of Fable in order to make something cheaper and easier for them to serve with good enough capabilities to keep people happy.
I don't think the results were as good as they were hoping for. In fact, my honest guess is that Opus 5 benched behind Kimmy K3 and that scared them because they didn't want to tell people, "Don't worry that we're taking away your really expensive super fancy capable model because you have this new thing Opus 5 that's almost as good." And then we would look at the benches and say, "It's not even as good as Kimmy K3. What the guys?" And I think this has resulted in Opus 5 being delayed multiple times as they try to refine it to make it bench better and look like less bad of a deal, but they knew how bad the blowback would be, so they decided to keep it in the plan for now.
All of that makes sense to me. This is a conspiracy I have no inside info, but considering how much we've been hearing rumors that Opus 5 is coming out today for like 3 weeks now, this checks out. And if it does come out of the benches are worse than Kimmy K3, you can come back and tell me I was right. I think that's all I have to say on this one. I do genuinely believe Kimmy's going to cause a pretty crazy shakeup in the market both on the regulation side and on the actual like costs of doing these things.
I think that the Frontier labs are going to have to move very differently now that there's an open weight model about to drop that is going to meaningfully challenge them in their business models on a fundamental level. And I think this will advance AI in the near-term future, but I am scared of what this will do for the safety of the systems that we rely on every day in the long-term future. I am still hopeful open source will win as I always am, but in the interim, it's good to see success like this and it's also great to use models like Fable and Open AI's stuff with GBD 56 both in your day-to-day work and to train better open weight models.
Can't believe I'm defending China as much as I am here, but I get it and they are operating for the greater good right now and I'll continue to support open weight models for the long-term as long as I possibly can. Even if I think those people who claim they can run Kimmy K3 in their houses are stupid, they are. For the rest of us, this is a good thing that incentivizes the labs to make better for cheaper and I don't want to live in a world where you have to rely on one or two companies for all of our work going forward.
And my team has requested that we end with this. I am sorry for that. Bye, nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.