Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
28:423.9x the video's typical replay level
noticed that the color isn't great. This is dark, and this isn't, and if anything, that should probably be flipped. And I also am hating the gray more and more since I had to remove the noise cuz of performance things that'll be in an upcoming video. Keep an eye out for that. Fable screwed up the performance of the
Said at 28:35
Most replayed moment #2
2:552.9x the video's typical replay level
already 30% faster for existing Depot CI or Sandbox runs. But more importantly, they managed to move from their roughly 10-second VM spin-up time to sub-second spin-ups, which is crazy for CI jobs. There's a reason companies like Posthog and PlanetScale do all of their builds
Said at 2:47
Most replayed moment #3
20:232.8x the video's typical replay level
doesn't have to waste a ton of RAM in order to preserve the actual weights. At least that's my rough understanding. Smarter people, feel free to correct me in the comments. Remember that supercomputer thing I said? Since inference efficiency likewise benefits from larger high-bandwidth communication
Said at 20:17
The graph counts replays. It does not show where viewers stopped watching.
Words
8,807
Runtime
41:35
Speaking pace
212wpm
Reading time
37min
212 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
KimmyK3 is here and it is a huge leap for open weight models. I'm going to be honest and tell you guys that I just haven't been that hyped about open weight stuff recently because it hasn't been close to where we're at with models like Fable and GPT-5.6. The open weight frontier caught up to where we were before with models like Opus 4.8, kind of, but nothing has come close to surpassing it, especially for day-to-day use on complex coding work, especially once you start orchestrating really long runs where agents spin up tons of sub-agents for complex tasks. I think this might have
106 words, the words spoken in the first 30 seconds at 212 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 461 |
| Average words per sentence | 19.1 |
| Longest sentence | 98 words |
| Questions asked | 3 |
| Sentences containing a number | 115 |
Most used terms
Filler phrases
99 in total: like 69 · actually 21 · uh 5 · kind of 3 · um 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
KimmyK3 is here and it is a huge leap for open weight models. I'm going to be honest and tell you guys that I just haven't been that hyped about open weight stuff recently because it hasn't been close to where we're at with models like Fable and GPT-5.6. The open weight frontier caught up to where we were before with models like Opus 4.8, kind of, but nothing has come close to surpassing it, especially for day-to-day use on complex coding work, especially once you start orchestrating really long runs where agents spin up tons of sub-agents for complex tasks.
I think this might have changed because Kimmy K3 is genuinely on the line for frontier, if not surpassing where we're already at for various different things. The benchmarks are showing some pretty absurd numbers with KimmyK3 beating out GPT-5.6 sole in various tasks, as well as Fable and others. And at the very least, it's neck and neck throughout pretty much every bench I've seen. The benchmarks only tell one part of the story though, so I spent the whole day building as much as I could with KimmyK3 just to push it to its limits and see what it's capable of.
And I'm going to be real, I'm blown away. There are definitely some rough edges and I'll do my best to show you guys how to work around them. My terminal froze because I was doing so much though, so I'm going to have to fix that first. This video is going to have a lot of fun in it from how to maximize your usage of the model to addressing the confusion around the different ways to use the model because there are quite a bit.
Talking about how the world is seeing this and what the impact might be both on how we do dev work, as well as the economy, but also possibly most importantly, the security implications of a release like this because this model is going to be open weight. And when you have a model this capable with no restrictions, there are some real concerns we're going to have to address. I'm going to go fix my terminal and while I'm doing that, I hope you don't mind a quick break for today's sponsor.
If you use GitHub Actions or Docker, trust me, you're going to want to watch this one because Depot is today's sponsor and they made both way, way better. Depot is fully compatible with GitHub Actions, but they also built their own alternative CI engine that is way faster and it can also be called from your coding agents using a CLI, which allows your agents to get feedback much faster than they would if they had to run all that stuff locally or wait for your PR to build it for you.
If you do use them for your normal actions, you'll still see crazy speed-ups, so up to 10 times faster. Docker's where they shine even more though, making your Docker builds 40 times faster for real-world use cases, and not just in the cloud, on your machine, too. The Depot CLI is a drop-in replacement for the Docker CLI that caches all of the layers on their CDN, which means all of your employees that are building the same images can all have way faster builds fetching from that cache instead of having to do the whole thing on their machine.
And somehow this all just got even faster with Depot Metal. As a friend of the Depot team, I am blown away at how far they went with Depot Metal. They went as deep as they possibly could on AWS, specking out machines directly that they own and control, managing the storage themselves as well. And the results show why they made these changes. It's already 30% faster for existing Depot CI or Sandbox runs. But more importantly, they managed to move from their roughly 10-second VM spin-up time to sub-second spin-ups, which is crazy for CI jobs.
There's a reason companies like Posthog and PlanetScale do all of their builds on Depot, and you can figure it out yourself at soydev.link/depot. In case you thought I was joking, I'm not. I actually have to force quit tmux right now in order to do what I was working on. Let's start with what the official Moonshot team had to say about this release. Today, we're introducing Kimmi K3, our most capable model. Kimmi K3 is a 2.8 trillion parameter model built on our Kimmi Delta attention and attention residuals, with native vision capabilities and a 1 million token context window.
These are two really nice, big changes. Things like GLM-52 don't have vision at all, which is one of the most annoying parts of using them. And the 1 million token context window is also super useful, even if it's not available in the subscriptions we'll talk about in a bit. It's very nice to have when you do need it, and this model is huge, so it'll be beneficial for a lot of different things. It's crazy how just a year ago the Kimmi models had no vision, had short context, and didn't even have reasoning.
And they've somehow caught up to the frontier in that time. Saying that they just raised around $2 billion and have raised almost $4 billion total makes sense that they're aiming for the stars. Like this is a moonshot in the most literal sense, especially at the size of 2.8 trillion parameters. If you're curious how big this is, the rough math for FP8 is that a trillion parameters is roughly a terabyte. So 2.8 trillion is 2.8 terabytes of data.
At FP8 it gets cut roughly in half, so it's only 1.4 terabytes. Oh man, that's so much more reasonable. To be very, very clear, anyone who's telling you that this is the future of local models has no idea what they're talking about and they should be ignored forever because a 1.4 terabyte model is not fitting in memory on any computer owned by anyone watching this. And if I'm wrong about that, please contact me. I would love to borrow your machines for some fun work.
It seriously though, this is not running on anything any of us have in our homes unless you happen to live in like the Colossus data center. This model requires supercomputers to be used. We don't know how big the closed weight models from Frontier Labs are, but we've seen estimates between 1 and 3 trillion per ams for a model like Opus, usually in the 1 to 2 trillion range. So having a roughly three trill model that's open weight is insane.
This is a massive leap in the size of models that are available for us to use, and I genuinely feel bad for our friends over at Hugging Face having to host this and deal people downloading 2 plus terabytes of data to use it. They won't have to for a bit though because their planned release date for the weights is July 27th. So the weights aren't available yet, which sadly means we have to use their APIs, which uh they're a Chinese company, so take that as you will.
They're going to have access to your code. Some people freak out about this and I can understand why. Let's talk about the performance a bit. They say that it still trails the most powerful proprietary models like Fable 5 and 5.6 Soul, but it also demonstrates frontier level performance across their evaluation suite consistently outperforming other tested models. I mentioned these in the intro, but we'll go through the benches real quick here.
Deep SWE, which is my current favorite software dev bench, shows 5.6 Soul slightly ahead of Fable 5, 73 to 70, and then Kimi K3 behind them at roughly the same rate at a 67.5, which puts them ahead of GPT-5.5, Opus-4.8, and GLM-5.2. Where things start to get more interesting is when we go down a bit to frontier SWE. I'm coming around to this bench because it seems more focused on how likely the code is to merge, not just how well it solves the problem, and Fable 5 definitely writes the most mergeable code in my opinion in my real-world projects.
And now that I've read more of the bench, I understand why it scores like this. I still think it's a little higher than it should be, and the gap's bigger than I would expect in real-world usage, but Kimi K3 sliding in between Fable and Soul with a 10-point lead on Soul is kind of nuts. It suggests this model is more tasteful is the best I can put it. Like it writes code that better It just looks and feels better, rather than just finding a solution to the problem.
They have their own internal bench where it scored just ahead of Opus-4.8 and behind K3, but it's worth noting their internal bench puts Soul behind 5.5, so I don't know how trustworthy it is. And then over here, we get the terminal bench where it beats out Opus and Fable 5 and is just barely behind Soul. Program bench, which I haven't really seen much of, but it is world-class there, just barely beating out Soul, which beats out Fable 5 by a little more.
And SWE marathon, a new bench that I know a lot are fond of, somehow Opus-4.8 was the lead before, even against Soul and Fable 5, but Kimi K3 has now come out in front. General agent evals are also quite interesting, things like GDP val, job bench, spreadsheet bench, which is now led by Kimi K3. To whoever is really into spreadsheets and open-weight models, it must be a phenomenal day for you. Congratulations. Browse comp scoring so high is of the most interesting to me though, because I'm a huge fan of browser use and computer use now after not caring about it for like 2 plus years, because the models got way better at it.
It's much more interesting. Having an open weight model that can do this is genuinely fascinating, because that means hypothetically speaking, if I could afford the hardware to run it on, I could have a fully offline runner that can control my computer and do real work without having to send my data to Anthropic or Open AI. Generally, if you want to do actual computer use work right now, you're just expecting to send all of those screenshots of your machine to one of those labs.
Not great. You hear solution there. It's expensive, but in the future if the costs come down and the opportunity use bottles of this capability comes more available, that is huge. I mentioned before that it has vision capabilities and it seems to be pretty solid there, too. But, what's really cool about the browser stuff as I was hinting at before is not just that it is industry leading by its score on these benches, it's also comically cheaper than the frontier is.
It's funny cuz I was just glazing this chart in the GPT-56 video about how much better soul was compared to everything else on it. But, I don't know if that's fair anymore, cuz Kimi K3 max gets even cheaper than soul max, roughly the same price as soul high, but a meaningfully better score. That said, an open weight model that is priced as cheap as possible competing neck and neck with 56 soul on cost shows just how far Open AI has gone in reducing costs to the best of their ability. 56 soul is still the fewest tokens to solve these types of problems and benchmarks, but since the tokens are so much more expensive, Kimi K3 eeks out a win there.
We'll talk about cost more in a little bit, but I'll just give the numbers now so you have them. 30 cents for a cash hit, $3 per million tokens in and $15 per million tokens out. For reference, the standard API price for solid models is $3 per million and 15 per mill out, but Anthropic is temporarily offering a discount of 2 per mil in and 10 per mil out right now. So, this is roughly a Sonnet-level priced model, but it doesn't have a lot of Sonnet's problems, which we'll definitely talk about in a bit.
Kimmy K3 is available today on kimmi.com, Kimmy Work, Kimmy Code, and the Kimmy API. I would ignore the two in the middle here. I will talk about the kimmi.com and Kimmy API stuff in a bit when I talk about how to use the model. At launch, it will use max thinking effort by default with low and high effort modes to be introduced in subsequent updates. This is one of the most interesting things I saw about this release is it doesn't offer reasoning controls at all yet.
You just use it and it gets used on max. They're currently working closely with inference partners and open-source maintainers to align technical details and ensure reliable rollout across the ecosystem. This is also exciting to see. I know the Kimmy guys have been pretty good about making it clear which providers are and aren't hosting their models correctly with benchmarks to verify the likelihood of any given provider actually hosting the model properly, which is great because this model is not trivial to host cuz it's an interesting implementation.
We'll talk briefly about that in a bit. They also haven't put out their technical report yet, which will have a lot more of those details. Moonshot has historically had the biggest open weight models with the Kimmy line. They put out Kimmy K2 as a trillion per am model all the way back in July 11th of 2025, and they stayed at that size for all of their releases from that point forward until now, where Kimmy K3 is a huge jump of 2.8 trillion per ams, putting them ahead of everyone else, even huge models like DeepSeek V4, which was 1.6 trillion on the pro version.
I feel bad for Thinking Machines. They were so hyped to put out Inkling, and it's only a trail per ams and didn't bench great, and now with this coming out the day after, oof, I feel bad for those investors even more so. K3 is built on their Kimmy Delta attention and attention residuals models, two architectural updates designed to improve how information flows across sequence length and model depth. They've also scaled up the mixture of expert sparsity, effectively activating 16 of 896 experts while paired with a stable latent MOE framework.
Again, not going to be trivial to host this cuz they invented a lot of their own solutions to make a model this big and capable without the compute cost being absurd. The memory cost is still very high cuz you're going to need a lot of this in memory for it to make any sense at all, but the sparse nature of how these experts are traversed should hopefully keep costs from being too absurd once the model's in memory. When you combine all of those updates with the refined training and data recipes that they produced, they end up with a 2.5x improvement in overall scaling efficiency, which is a pretty big jump.
These guys have always been pretty good about sharing their advancements and the cool things they do, so I'm extra excited for that technical report to come out. Now, let's talk about what using it looks like on the coding side especially. Kimi K3 has strong long horizon coding performance, meaning it can run long tasks with minimal human oversight. One of my favorite test tasks for these much bigger, more capable models is to take my old code base for ping.gg, which is a Zoom app for content creators doing live collaborations, and see if the model is capable of porting it.
It's been running for like 3 plus hours, and it was doing great until like 10 minutes ago where it hit a limit on the context window size. This is a mistake that's partially my fault because I didn't realize that the subscriptions that you can use for Kimi code don't actually give you the full 1 million token context window, and with my setup using it with Claude code, it was hitting a much smaller limit. So, I'm going to really quickly change that max size and get that compacted so it can keep going because I want to see if that run can complete.
Sadly, I won't be able to compact using Kimi K3 because again, it won't respond to the API request. I'll be able to get it compacted with Fable and keep the run going in just a moment though. The fact that it was able to get through 122 tasks with a single like paragraph and a half prompt and no additional insight or effort from me is unbelievable though. I've never seen an open weight model come close to staying coherent for even half of this length, much less like actual real-world massive migration work.
It's impressive. And since it also has visual reasoning, it's way more capable of UI type work, too. It leverages screenshots and visuals to optimize game dev, front end, and CAD. And those visual capabilities are nuts when it comes to front end code. We'll talk more about this later, I'm sure, but just know in advance that Gemini K3 is really, really good at front end, at least according to Arena AI. From my experience using it, it has its quirks, but it is very impressive, and I cannot wait to show you guys just how cool some of the UI it creates is.
On the very opposite end, we have the kernel optimization capabilities, where apparently they use it to optimize GPU kernels, and it did a pretty damn good job. After being active for around 15 hours, it saw slightly better improvements than even Fable did, and quite a bit better than 5.5 and 5.6 Soul, somehow, which is nuts. They have four types of kernel optimizations that they benched here, and somehow K3 came out near Frontier or above Frontier in all of them.
It's also really interesting to see how big the gap is between Soul and Fable in a lot of these as well. This might be a good bench for them to actually like put out in the future, not just cuz they're leading it, but because it's fascinating to see how big the gaps are. This is also scary for companies like Anthropic who have went out of their way to hide the model's capability of helping with ML work like kernel optimization for hosting models.
They went out of their way to keep the model from sharing those things, and it's one of the restrictions they have on both Fable and Mythos. So, it's fascinating to see Gemini bragging about how good their model is at this when they plan to release the weights, because that means this capability is now in the hands of everyone to an extent. Man, Anthropic has to be terrified of this release more than anybody. In the late stages of Gemini K3 development, an early version of K3 handled the majority of the team's kernel optimization work.
That's pretty nuts. They also tested if it could a complete GPU programming stack and compiler from scratch. They developed many Triton, a compact Triton-like compiler with its own tile level IR layer over MLIR. I'm sure much smarter people will know what this all means and be really pumped about it. Yeah, it seems like they did a really good job training the model to do these type of stuff. Here's where I'll be much more useful.
Game dev and digital creation. K3 combines strong 3D reasoning, coding, and vision capabilities to turn concepts, images, and videos into fully playable interactive experiences. It achieves a true vision in the loop by seamlessly iterating between code and live screenshots, instantly seeing and refining outputs. I actually watched it do this live when I was having it do some changes to T3 code, where it would spin it up, get a screenshot, look at it, think about it, and then change what it was doing.
Apparently, it created everything here, including the rider and horse models. So, it like actually understands 3D? I'm going to have to play more. I want to see this in action. I just spun up Pi to go do a 3D port of Fish Slap. We'll see how it goes. Hopefully, not well, cuz I don't want to have to record more after I'm leaving. I'm actually supposed to be at an event right now, but had to film this because it's such a cool model.
Yeah, the fact that it's able to do this type of 3D work and modeling is crazy. Every model I've tried so far is just so rough at 3D stuff that I've been impressed by like circles being placed properly sometimes. If this model can actually do 3D, I'm going to have to that's going to change things. Oh, Oh, Theo from the future here. It got pretty far in Fish Slap 3D, so I wanted to check this out quick. Holy The 3D game is like actually working.
Uh Sorry about the audio. Can I mute that easily? No, I can't. I have no idea what it sounds like. I'm not putting on my headphones to see I I I'm I'm putting on I'm so curious. >> Holy the fish textures. They're not perfect at all, but this is the best fish model I've seen any lab create. And the submarine is phenomenal, too. For a model generated for a 3D web game like this. It's got its issues, for sure, but like it's still working on it.
It just said like the game works, let me do more. But the fact that it's already this far, unbelievable. Okay. Holy Yeah. It even has sound effects and things. This is so much better than I would have expected. It's still working like I'm checking this before it's done, and it keeps pulling in visualizations like it ran it, and I saw the little picture in the pie history where it actually like pulled it up and was like, "It's working.
Let's continue. Let's fix all these things." It's a good model. Apparently, it's also good at chip design. This is interesting. It might have real like beyond what the frontier currently allows capabilities that none of the other frontier labs are focusing on. This is actually arguably one of the problems with this fixation on code that both OpenAI and Anthropic have as they try to win an enterprise. They're not finding new capabilities the same way that they used to.
Here, it seems like this new frontier open weight model is actually able to explore things that the labs here have just not explored. Hopefully, as long as the 3D stuff is as cool as it seems, that's going to be huge. It's also incredible at knowledge work, according to them. Benchmarks like online EXP bench, deck bench, and finance bench, it is industry leading in. A lot of that's probably because of how good it is at visualization stuff.
They had it do a bunch of research work, and it did a pretty good job even at generating the actual reports, which is pretty damn cool. Good at infographic style presentations. It's still got that like AI-generated vibe that a lot of those have. They show some dashboards. It does really like this um Bento box style UI layout, but it makes them in a very pretty way, and the animation taste is actually quite good, too.
I've been impressed with it. Apparently, it's also good at video editing, which is crazy. I will not have any time to try this anytime soon, but uh if my team ends up liking it, I'll be sure to share that in the future. I'm not that into AI-based video editing cuz video editing is actually quite fun and not too tedious if you are any good at it at all. And they had edit a lot of their videos for the launch of the model.
Having it go through lots of clips, handling clip selection, motion matched cuts, frame accurate beat synchronization, audio processing, and multiple rounds of revision. That's pretty cool. They have some cool info here about how they were able to get a model of this size to be trainable in a reasonable time frame compute-wise in in a stable fashion cuz the bigger it gets, the harder it gets. They found some fun tricks like using FP4 weights for the actual stored weights and params, but then using FP8 for the activations once the model's actually running so that it has better short-term memory, and it doesn't have to waste a ton of RAM in order to preserve the actual weights.
At least that's my rough understanding. Smarter people, feel free to correct me in the comments. Remember that supercomputer thing I said? Since inference efficiency likewise benefits from larger high-bandwidth communication domains, we recommend deploying K3 on supernode configurations with 64 or more accelerators. 64 H100s is a bit rough. That's 2.6 million to host the model. Yeah. Very local-friendly. So, that's what we have from them.
Let's take a quick look at what others have said. I mentioned Arena AI gave it an absurd score for front-end stuff. We'll show some front ends in a bit. We're going to start with Artificial Analysis because the model, according to them, is the third smartest ever. So, pretty much out of nowhere, we had two drops in a row that took the frontier away from OpenAI and Fable just going back-to-back forever. And suddenly we have Grok 4.5 and Kimmy K3 coming up for those third-place spots right behind Soul and Fable.
Normally Opus would be up there, too, pushing these back, and it's not anymore because both xAI and Kimmy have gotten their so together that they are leapfrogging. And they are so far ahead of what Google is cooking, it's hilarious. I honestly think Google if they had any reasonable way to buy a Chinese company like Moonshot, they probably have to at this point cuz they are just so behind in comparison. Okay, slight correction.
Grok 4.5 is behind both Tara and Opus, as well as 5.5, so it's not really up at this frontier level, but Kimmy K3 is. It is just behind Soul and Fable according to Artificial Analysis's benches. But things get much more interesting as we dig in more. I'll read what they said on Twitter first. Kimmy K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and 5.5, but remains slightly behind Fable 5 and 5.6 Soul.
Moonshot has expressed plans to release the 2.8 trillion parameter model's weights, which would make it the leading open weight model. It's got really strong agentic performance, as I mentioned before, it's meaningfully better than models like GLM 5.2, going from 1668 on GDP val from 1514, which was the open weight frontier before, and even ahead of models like Opus 4.8. It's the second highest score I've ever seen on Artificial Analysis's knowledge bench called briefcase.
It will absolutely be the leading open weight model, not a surprise there. The cost per task is similar to 5.6 Soul, but it's half the price of Opus and higher than all of the open weight peers. Not just cuz the model pricing is more expensive, but also remember the amount of tokens it uses is important, too. For reference, in output tokens per task, Kimmy K3 is a decent bit lower than Fable 5, which is the most token efficient model in Fable's put out in a while, but compared to things like 5 6 Terra or even Soul all the way back here at 15k tokens per task, 23 is more, but also not much more when you consider the price difference and it's so much less than other models, especially other noisy open weight ones, things like DeepSeek V4 Pro or even worse V4 Flash, which was 45k tokens for the same tasks.
It's pretty token efficient and that's awesome to see because historically only OpenAI has really focused on token efficiency. Now we have both Grok and Kimi surrounding them in their little cheap token section here, finally driving the industry towards more efficient reasoning. 21% more efficient than K2 6 was even though the model is bigger and the tokens are more expensive. It has native multimodal capabilities as I mentioned before, very, very good there.
The AI Omniscient score is also very interesting. If you're not familiar with this bench, it's Artificial Analysis' attempt to measure hallucinations, specifically if the model is going to tell you when it doesn't know something versus will it make up something instead. And Kimi K3 is one of the best open weight models at not hallucinating. In fact, I think it is the best by quite a bit. Even 5 2 is a much worse score here.
And this test is fun because you can get into the positive when you say I don't know the answer and you go negative when you lie. So even some very good models like 5 6 Soul end up scoring a lot lower than they should here because they're a little too quick to lie when they don't know and the extra points they get for the things they do know get canceled out a bit by that. And 5 6 Luna gets hit real hard with this, getting into the negative as a result.
One of the most honest open weight models we've ever seen and that's genuinely exciting cuz open weight models have not been good at this historically. But if your goal is to use this model to save a bunch of money, you probably shouldn't be too excited yet because that $15 out cost is not cheap and since it uses twice as many tokens as something like Soul, cancels out the 50% discount that you're getting and you also lose a lot of the niceties that you get from Modern Frontier Labs.
Moonshot themselves even said that the model still has a noticeable gap in user experience compared with Fable 5 and 5.6 Soul. So, if Soul is as cheap if not cheaper for your use case, it might end up being better just cuz it's nicer to work with. And I have noticed that with Kimi myself. I noticed things like when a workflow finishes, it responds to the workflow finishing instead of giving me the context I need. Those are things you can work around, but you're going to have to do that with this model because it's not as R L on the expectations we have as users and they just don't have as much data cuz they're not getting feedback from people the same way cuz the vast majority of users of the Kimi models are using them on other providers.
So, they don't get any of the telemetry they would need to improve these things. I'm sure it will get better over time, but not surprisingly that there's a real usability gap here. They also call out that it's too proactive. Reminds you of a certain Rottweiler model, as well as the sensitivity to thinking history that it needs its thinking. Thankfully, it's an open weight model, so we get the thinking data, too. It's not like Infropic or Open AI where the thinking is hidden on some server.
All reasonable, really cool call outs. I love this level of transparency from a lab that's releasing something this important. Okay, enough of the research side. Let's talk about the actual front-end capabilities cuz I know a lot of you guys are excited about this. I did one of my recent favorite demos, which is to have things redesign the T3 code marketing site. This is the design I came to after working with the Claude design product, my own brain, and a lot of back and forth.
I got it here. There are little things I would change. Obviously, I want to make the cursor icon a different color so it fits better, but you get the idea. It's not bad. I love my little carousel here with all the nice things people said about T3 code. So, let's look at what it did. I told it to do five different designs on different URLs to keep them varied and here's what it gave me so far. This is the first one. It's doing pills as a certain Open AI models really like to do, so that was interesting to see.
I did this with Open Code as the harness if you were curious. So, that's number one. Not bad, but not something I would ship. Here's two. It looks It's nice except for the fact that this type of design has been copied by so many models for so many things that it's a not as cool anymore. I also can't help but notice that this top bar doesn't have enough content, so it's cycling, but it's half empty now as a result. Has a lot of nice little animations when you hover over things, which is cool.
Still not perfect, though. But like considering this is an open weight model, unbelievably cool. Switch over here, and now we get the cringe terminal one. I was so shocked that this snuck in that I actually asked if it had pulled in my UI skill or something, so I thought I had removed it. And it hadn't. It's almost like the Claude design skill got baked in, though, which is For the fourth design, I channeled more of what I already had, but it with a surprisingly not too cringe glow in the corners, also fixing this icon to be the right color.
Still don't necessarily love how it structured things here, too bubbly, but you could steer this somewhere good, for sure. I do think the gradients are pretty cool. And then we have the Bento box version, cuz I knew it would do one of these. I don't think it fits for the type of product that T3 Code is, but it's not a bad design at all. So, from just this one pass, I would say that for doing a marketing page, it's slightly better than what I get out of open AI models, but slightly behind what I would expect from Claude.
But marketing pages are far from the best way to measure the UI capabilities of a model. So, I gave a slightly harder task. It was actually something I was already working on. I've been overhauling the sidebar in T3 Code. I want to find a better way to do it, something that's a little more uh flexible based on how it's being used. So, I already had this build where I redid the sidebar, and this build has a much better experience there, but I've noticed that the color isn't great.
This is dark, and this isn't, and if anything, that should probably be flipped. And I also am hating the gray more and more since I had to remove the noise cuz of performance things that'll be in an upcoming video. Keep an eye out for that. Fable screwed up the performance of the app and I had to remove a bunch of to fix it. So, I asked it to do a darker redesign with true black instead of grays as often, and this is what it made.
I will be frank. This is better than what we had. It made a better design than what we were already doing. With an open weight model that is a third the price of the frontier models from the other company that makes ones good at UI. I think we finally have a model that's good at solving real-world UI tasks without having to pay Anthropic massive amounts of money. Very exciting. Another cool UI example that Mac shared here is a recreation of macOS 27's liquid glass styles in a real web app that you can load that Kimmy built all of from scratch.
And they're also hosting it on kimmy.page, their little like web hosting thing. And this is nuts. To have something that looks and works this well that like a open weight model generated for not too expensive is kind of crazy. There's a lot of little things that are broken in it, but like this is not bad at all for something that a model generated. Seriously, this is dope. What was much more exciting to me is how well it did this work.
I did this work through OpenCode's bindings in T3 Code because, as I've mentioned many times, T3 Code is not a harness. You have to bring something like Claude Code, Codex, or OpenCode in order to use T3 Code. So, since Kimmy isn't available in Claude Code or Codex, wink, I'll show you something soon, I used OpenCode. And this was a fun bit of work. I told it to rethink the core UI for T3 Code. It's currently too gray.
In dark mode, I want a cooler black and white layout with hard blacks inspired by ChatGPT. I gave it a screenshot of the ChatGPT UI so it could have that as a reference. And then an hour and 5 minutes later with one steer where I told it that it was opening things incorrectly cuz it kept trying to run the dev command to see, but when it did that it opened my browser and burned the code one-time token that we use in T3 code to verify who is connecting.
So, it couldn't access it the way it wanted to. But, once I told it that and it stopped letting it open the browser, then it was able to access it, check its work, and keep refining until it got somewhere that it liked. It also apologized for opening new tabs constantly cuz it kept forgetting the no browser command. Sorry about that. Everything I started is killed in the temp homes have been removed. Also, a parallel sessions dev server, coding and PP ports got caught in a broad cleanup kill.
I restarted those with their original flags. Nice. So, it made mistakes, it told me about them immediately, and then cleaned them up. Other models will make similar mistakes, but they don't necessarily tell you when they do. So, that was very nice to see. I asked it to do a more machine task, not like building a project, but fixing a config because I wanted to run this in isolated environment so it didn't affect my existing T3 code install, just so I could open it in the browser and play with it a bit as I did here.
But, in order to do that, I needed content. So, I told it to figure out how to get things from my official T3 history, which was breaking this build cuz I have all my sidebar changes. And it managed to pull over the history, modify it, and get it in a state where I could actually work again. Thought that was really cool that it could do that type of thing. I then asked it in the same thread to file a PR. Follow Reborn Adventures for PR naming, minimal description, include before and after photos, use the file upload skill.
If you don't have the skill, look at the global Codex and Cloud skills, you can find it there. I added this because I didn't have the skill in open code and wanted to go find it, and it did, and it got it working. It again struggled a bit with getting the browser going, but once it realized it could do the no browser command and also open in Puppeteer, it pulled it together, and it made this PR where it shows the before and the after.
You can very clearly see how much better the after is here. I'm I'm impressed. It did a great job with this work. I didn't hold back in any way. I didn't give it an easier thing to see if it could do it. I worked with this the way I work with other models, and it did all of the work very well in one thread without having to do any custom anything. Since it did so well on those tasks, I decided to push its limits with these bigger overhauls.
I did have the issue with the token window and compaction with the previous set as I mentioned. I'm working on fixing that now, but that did not stop me from having success with some other really big jobs. This was one of those jobs. I asked it to do a deep audit for security issues on my cloud product that I'm building called Lakebed. This is one of those tasks that I get a lot of refusals for when I use Fable and Soul for it.
So, I was curious to see if it would refuse as well as what it would find cuz I have a concern with a really powerful open weight model like this coming out. And it did it. It spun up a bunch of agents to go do discovery on specific things. It then followed up with a verification pass for all of those things that it found doing 25 verification agents. And then it synthesized it all at the end, ran out of token space, I compacted, continued, and it succeeded.
It gave me useful feedback on things I can do to secure further. This is good on my end cuz I can use this to secure my systems, but it's going to be really bad for us in the future when attackers start using it for similar things. All of that effort that Anthropic and OpenAI have been putting in to use their models for defense work and blocking them from doing offensive work is very helpful because that is no longer going to be enough to protect us now that we have an open weight model this capable.
And if you're curious how Moonshot is thinking about security and safety, the word security appears inside of some of the demos on their page, but not in the actual contents of it. And the word safety does not appear on the page at all. As I mentioned before, they have not put out a system card yet. They plan to put out the technical reporting in the future, but we have no info about safety and security and what the impacts on those will be from this model yet.
We don't even know how they thought about safety and security during training. So, that is a thing that is worth being concerned about. And if I talk about just how concerned I am, this video will be much longer. So, go check out my other videos about security and my concerns in that general space. Since the reasoning is visible in this model, I did take the opportunity to read said reasoning a little bit and uh I don't know if the other labs do this too cuz most of them don't share the reasoning, but it was interesting to see the dumb things the model would get stuck on.
For example, in my global agents MD, I'm very strict with the models that I don't want them to write any all over the place. I want them to actually use TypeScript for its types. So, when it was working on a plan in a markdown, it read the detail that we need to ensure no any per user preference. The plan is markdown, not code, so okay. Like it took a second to think about that. Time was spent reasoning about the fact that it shouldn't use any, the TypeScript syntax, inside of a markdown file.
When you combine that type of bad reasoning with the model's 20 TPS out, which is less than half of what we see from other labs, the result is that it feels slow. It's spending time doing things it shouldn't. Others might as well, but we don't see that data. At the very least, I would guess OpenAI doesn't cuz their models are so efficient with reasoning. But that is time I spent and money I spent on a thing that doesn't actually affect the quality of the outputs.
Here's another weird experience I had. I mentioned this one earlier, where I had it doing a big workflow to write a plan up for me. And at the end, it didn't tell me what to do. It ended with no new report here, just an idle status ping. The plan in plan/overhaul.md remains current with all three audits incorporated. It was done. It had completed the work, but it was responding to the sub-agents and the ping, not to me, the user who was setting this off asking it to do the work.
Just a weird thing that I had to notice and be like, oh, it's not going anymore. I need to do a follow-up. The reason I tend to use these models in Claude Code is because I really like workflows. I've talked a lot about this in previous videos. Go watch any of my videos about Claude Code versus Codex or why I use 5/6 in Claude code probably has more detail there as well as how I did this. That method worked perfectly for Kimmy by the way.
It was very easy to just do slash model Kimmy K3 and it just worked. But what really blew me away was that Kimmy K3 can use sub-agents and workflows and all of these things correctly, but it did it differently. Normally when I tell a model to use a workflow, it will use one workflow to break down a big task. Kimmy made a bunch of workflows for various phases of the work it was going to do with the port of Paint to a modern stack.
It broke up all these different phases, created workflows for them, and started running them. But the most interesting thing is that it allowed the sub-agents to complete tasks on the to-do list a level higher. So it had a to-do list, it turned the to-do list into workflows, started running some of them, and I noticed the to-dos just like randomly getting crossed off throughout. And that was because it gave those sub-agents the ability to check them off.
Never seen that before. Even Fable doesn't do that inside of Claude code for my experience. So that was really cool and fascinating to me. Probably time to talk about how you can use the model today. Since the weights aren't available yet, the only way to use it is through the Kimmy platform. To be fair, this has been the case with Anthropic and OpenAI historically. They do usually set up their models on Bedrock with AWS as well as GCP and Azure, depending on which lab and which models and which provider.
So there are some other options, but at the very least you always have a lot of options that are based out of the US. That is not the case for Moonshot. Moonshot is a Chinese company, which means either of these solutions are going to go through Chinese servers some amount. I would not use either of these methods for real important sensitive data at all yet. Wait till the weights are out, you'll be able to do that then.
But for now, these are your two options. You have the subscription on kimmy.com or the API, which is platform.kimmy.ai. These are both hard to find, so I ended up accidentally spending $100 on the API because I forgot I had a $40 subscription on Kimmy. They have an interesting breakdown of subscriptions. They have a $20, a $40, a $100, and a $200. And I ended up hitting the $40 pretty fast with K3, at least the 5-hour window.
I hit it in about 30 minutes of my early testing. So, I bumped up the $200 tier. And despite my borderline abusive usage, I have not even gotten to 50% of my 5-hour or 10% of my weekly since doing that upgrade. It seems to be quite generous, which makes sense cuz the model only costs $15 per mill token out. So, I did a bunch of calculations on how much my usage has cost. For the $40 sub that I burned it upgraded, the $10 or so of API usage I did before realizing this, and all of my usage in that sub so far, my estimates from using CC usage and having my whole fleet analyzed through Soul, is around $63 of usage.
That is not bad at all considering how much work I've been getting done with this model. I'm impressed. For context, I'm doing like a $1,000 a day with OpenAI models right now just on weird side project I was burning and looping on, and I do at least 3 to $400 a day with Fable for my day-to-day work. I've done a similar amount of work with this model and it ended up being way cheaper. So, that's good to see. Hopefully, not similar to the amount of work I do with Soul, to be clear.
I do way too much with that. But, the cost to actual work done compared to Fable, it's a really good ratio. And it does also instinctively so far seem slightly cheaper than 5/6 on lower reasoning levels for my day-to-day work. What this means for you is that the API is actually a pretty reasonable option because the costs aren't too bad. But, I would start with a subscription, at least like the $40 tier to play with it if you're curious enough to do said play.
And if you do sign up for either of these, you can bring the API key from them into Open Code, therefore making it work in something like, I don't know, T3 Code. It's also worth noting that the subscriptions work inside of the CLI proxy that I use for routing my Claude code to both Claude code and Codex, and now also to Kimmy. That's some pretty nice, too. If you'd prefer to wait until the weights are released so you can use this on American servers, I totally understand.
But if your reason for waiting is cost, don't. I ran a bunch of numbers in previous Kimi models, when you compare the price they released them at to the price that was available on other hosts and providers, they were cheaper, but it was like 11 to 15% at best. These models are just massive, so they're not cheap to host, and I would not expect this model to suddenly become $5 on other providers unless they quantize the hell out of it and make it way dumber.
That doesn't mean the open weight nature of the model isn't great for the market as a whole. It's going to make pricing way more competitive, and I would be genuinely surprised if this doesn't force Anthropic to rethink some of the Opus release that they have coming in the not-too-distant future. If they still price it the same they did before and the numbers end up worse than Kimi's, Anthropic's going to be in a rough place.
Think I've said everything I have to about this model. It's unbelievable at 3D. It's really damn good at visual stuff, especially coding and UI, and it's just overall a really good model. It no longer feels like it's good for open weight. It now feels frontier-class in many different ways. And while talking to it and working with it isn't quite as pleasant as the best models available, it's close enough that I am blown away, and I'm so excited for the future of models like this and where they're going to go.
Let me know how y'all feel about the Kimi K3 launch, and until next time, peace, nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.