Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Cole Medin · @ColeMedin
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Cole Medin's most watched videos.
Most replayed moment at 10:58
3.0x that video's typical replay level
cache is performing in production with real user data. And the best part is better DB is open source and free to get started. So I'll have a link in the description. I'd highly recommend them as a tool to help you scale manage your costs for agents you're deploying to production. And so now Google is saying with
Said at 10:50
Most replayed moment at 3:20
3.2x that video's typical replay level
doesn't end up becoming the standard down the line for personal agents. There's going to be something like this. And so it's good to understand this now. Okay. Now, let's really get into OKF. So there are two things that they're standardizing here. The first is how we are organizing information like our
Said at 3:13
Most replayed moment at 9:16
3.5x that video's typical replay level
not extremely difficult to get all this set up like it used to be. And the best part is the agency CLI is free and open source. You can take these skills, bring it into any coding agent, and see how easy it is right now to build any AI agent. I'll have a link in the description. I'd highly recommend
Said at 9:09
The graph counts replays. It does not show where viewers stopped watching.
Words
3,022
Runtime
14:08
Speaking pace
214wpm
Reading time
13min
214 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
This week, I've hit my Claude Code rate limit way faster than I ever have before, and I'm not even doing that much more work. This is my max $200 a month plan, and the worst part is, I also have a $200 month Pro subscription for Codex, and I've almost exhausted that limit as well. And both of them aren't going to reset for another 3 or 4 days. And so, if you thought it was just you hitting your rate limits faster, don't worry, it's not. We're all suffering from this, and unfortunately, we knew this was coming. It's been a trend this entire year,
107 words, the words spoken in the first 30 seconds at 214 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 159 |
| Average words per sentence | 19.0 |
| Longest sentence | 59 words |
| Questions asked | 7 |
| Sentences containing a number | 23 |
Most used terms
Filler phrases
43 in total: like 25 · actually 7 · I mean 4 · right? 4 · basically 1 · literally 1 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
This week, I've hit my Claude Code rate limit way faster than I ever have before, and I'm not even doing that much more work. This is my max $200 a month plan, and the worst part is, I also have a $200 month Pro subscription for Codex, and I've almost exhausted that limit as well. And both of them aren't going to reset for another 3 or 4 days. And so, if you thought it was just you hitting your rate limits faster, don't worry, it's not.
We're all suffering from this, and unfortunately, we knew this was coming. It's been a trend this entire year, our rate limits getting worse. But it's just finally come to a point where I can't just be more efficient with my token usage. I mean, it's ridiculous that I have to wait 3 to 4 more days before I even have a reset on my max plans. So, over the past half year, I've been doing a lot of experimentation with using different models, especially open ones, in my AI coding workflows.
But now, that's gone from just experimentation to, "Okay, I absolutely need this as a crucial part of everything that I'm doing." I mean, look, the proof is right here. And that's not just me. As an industry, when we get more into harness engineering and loop engineering and software factories, whatever your AI coding workflow looks like, we're trying to scale the output with our coding agents. But now, even more than before, the biggest limit is just the number of tokens that we can spend.
That is why we need to strategize and figure out for our larger AI coding workflows, not single agent, where do we absolutely need the largest and most capable LLM versus other places where we can use something that is smaller, cheaper, faster, and still get basically the same output And yes, that is possible. Trust me, I've been deep in the lab this week, expanding on all the research I've been doing this year, figuring out for a traditional AI coding workflow where you have your planning, implementing, validating, and reviewing, where do we need the best LLM versus where do we not?
That's the big question to answer right now, so you don't hit your rate limit 3 days into your weekly reset. So, for all the testing that I've done this week, I've done it with my AI software factory. I've covered this system a lot on my channel recently. It takes all my best practices for AI coding with planning and implementing and validating. It packages it up into a complete system where I can send in a scope of work and it just builds and validates everything without me in the loop at all.
It's the ultimate test of how autonomous can we make AI coding while still keeping it reliable. But more importantly for this video, it allows me to very easily test different combinations of large language models with the full development cycle. So, what I've been doing is building a bunch of different applications with the AI software factory including this video game that I'll show you today. It just makes for a nice and visual core example.
But all the other things that I built, the results I got is pretty much the same as what I got with the different combinations for this game. So, building the same game but with only open models using Claude and using Codex. A lot of other testing that I did as well, but these are the three core columns. And so, essentially I just took the best model from each group, right? With open models, DeepSeek V 4.1 Flash is pretty much the best of the best.
So, if you look at like the live LLM benchmarks here, as far as open models go with that open tag, DeepSeek V 4.1 Flash is the best right now. And then for Claude, Claude Fable 5.1 Max Effort is the best. And then for Codex, GPT-6 Astra Max Effort is the best. And so, going back here for my core testing with open models, I'm using DeepSeek Flash for planning and reviewing and then GLM 5.3 Flash for the implementation.
And I actually did this as a baseline where I used it for implementation for every single test. And then for Claude, using Fable for planning and reviewing. And then for Codex, using Astra. The most important thing here is you'll notice I'm always using the more capable LLM for planning and reviewing and then the smaller model for building. This is something that I've determined in my experimentation this week and a lot of what I was doing earlier this year as well.
The most important step of your AI coding workflow is always the plan step. If you have a well-written plan, the implementer following that, it doesn't even have to be that capable of an LLM to pretty much get the best results possible. Right? Like using Fable for planning and implementation gets you essentially the same results as using Fable for planning and Sonnet for implementation. So that you'll see that very intentionally laid out in all the testing I show you today.
And by the way, for implementation, generally that's the most token-heavy part of your workflow. So for using the smaller model for that, that is what prevents us from hitting our rate limits super quickly. Now, obviously for some work, if it's very, very difficult and unexplored, you might want to use a large model for everything. And I don't do this in my software factory, but if you do have human in the loop, really small and bounded tasks, you can use an open model for literally everything.
So, there's different combinations that make sense in different situations, but this is generally what I recommend. And I'll show you a lot of my testing that has led to this, but for me, my favorite flow right now, my favorite combination, is to use GPT-6 Astra for the planning and the review, and then GLM 5.3 Flash for the super fast and cheap implementation. This is my core workflow for everything that I'm doing right now with my software factory and even outside of that.
The sponsor of today's video is Scrimba, and they just released something incredible called Explain. You ask a question and it gives back a fully narrated video lesson in 2 to 3 seconds, built live while you watch, and it plugs right into Codex and Claude Code. So now over in Claude Code, I do {slash} MCP, we can see I have Scrimba Explain connected. It was just a single command to do this. And so now I can ask Claude Code to make an explainer for anything.
And so I thought this would be a neat demo here to show it making a Scrimba explainer for a more complicated pull request from my open-source project Arkon. So here Claude reads the diff, it builds up the lesson, and then it streams to Scrimba slide by slide. So I get the link right away, I can watch it as it builds the explainer. And I got to say, the explainer that I got back here at the end, it genuinely impressed me.
Let me show you. So, this is the original pull request that I made the explainer for. And then here is the video. So, I'll just shut up for a second and play a couple of the clips here. PR 3416 gives three adapters one shared way >> I'll go to a more high-level diagram here. >> is the contract that the CLI adapter and the web adapter implement. >> And it goes into the code as well. Let me click a little bit further here. >> lives in core.
The helper accepts any object, which lets the headless platform pass its structurally identical work >> So, I'm not going to play the entire thing, obviously, but yeah, it sounds great. And it's breaking things down really nice and simple for me. I love it. Explain also works from ChatGPT and as a Chrome extension for any article. Give Explain a try for free with my link in the description. Now, the question you might have at this point is how exactly do we build a larger AI coding workflow where we're using different models and providers?
Cuz maybe you want to implement the same way I am, you want to experiment like I am, whatever it is. There are a lot of different ways you can do this. The most manual, but simple way is just to have each coding agent output a handoff document as markdown, so you feed that into the next model or the next coding agent. And so, you could just go through one at a time, opening up different coding agent sessions. There are also harnesses out there like Omnigen that make it very easy to work with models and different providers in this way.
But what I do for my AI software factory and my coding workflows in general, is I simply use my open source harness builder Archon. This makes it really easy for us to build these workflows that I've been running hundreds of times the last week to combine different models and providers. So, for every single test that I ran, every single combination of models, the shape always looks like this as an Archon workflow. And I'll show you what the workflows look like in a little bit.
So, the input is always a GitHub issue that describes what we want to build. Then I have the most powerful model do the planning like GPT-6 Astra as my favorite. We take that plan, and then within the Archon workflow, we automatically pass that into the next node, which is GLM 5.3 Flash doing the implementation. Then, we have Astra review, and then any findings it sends back to GLM to correct, and we have that loop with obviously max retries.
And then, we have the check at the end, which is just running our tests, making sure the build is green before we do the merge at the end of the software factory. And so, this is a pretty traditional AI coding workflow. There's nothing crazy here. This is the exact shape for every single test I ran. There are also a lot of different ways to access open models for our AI coding workflows, but my favorite recently, especially cuz Archon supports this, is to use Py as the AI coding agent harness.
And then, for accessing the LLMs, using Neon's AI gateway. I've always loved Neon. I love their Postgres database solution. They've been adding a lot of other really cool AI features as well like AI gateway and all their other new back-end features like object storage, authentication, and functions. But yeah, for right now, I'm using AI gateway, which gives us really reliable and at-cost access to pretty much every open model that you could hope for.
So, for example, we have Kimiko 3. Uh what else do we have here? We have GLM 5.3 Flash, which I've been talking about a lot. And so, this is just my easy way to access all the models that I need within my Archon workflows and with Py. But regardless of how I am accessing these open models, the important thing is that I am. I'm not just using Astra or Fable for everything like a lot of us are tempted to do. And you can do the setup with anything.
You don't even need Archon. It's my tool. It's free and open source, and I use it to run the AI software factory. But also, I'm not telling you have to go use my tool at all, right? Like this is just what I've been doing for my experimentation to give these results to you and help you think about what models to use, what combination. And so, for my Archon workflows, this is a bit of a simplified view, but pretty much every single one of them looks like this, right?
We have the planning step at first, where we're using the more powerful model, and then we take that plan markdown document as a handoff to go into the build step with the cheaper model like GLM 5.3 flash through neon, and then review with more powerful again, and then fix anything that came up with the cheaper model. This is the shape, even beyond my experimentation here, for just generally how I work with coding agents now.
Okay, so I wanted to spend good time explaining my workflow, my experimentation, and the models that I recommend, but now I want to get into the actual game with you as the core example here, just to give you a visualization of the results that I've been getting across all the different apps I've been building. This, again, like I said at the start of the video, is very much the results I got no matter what I was building.
And so, the first version of the game looks terrible. You can probably guess this is what I built with only open models. And so, by the way, this game is an idea from a friend. It's really cool, so I thought I'd put this through the factory. It's a weather chasing, like storm chasing game on Neptune. It's actually multiplayer. You can have different people join and man the stations here. I won't set that up right now, but it's pretty comprehensive game that it gave us that PRD into the software factory.
And so, this version of the game, I used open models, so DeepSeek v4.1 flash for the planning and reviewing, and then for the building I used GLM 5.3 flash. Now, the visuals are so bad you can't even really tell that I'm moving. There's a couple of clouds you can see outside the windows, but yeah, we can see the heading changing in the top left as I'm turning around as the pilot. But overall, this isn't even really a good starting point for a game.
And then for the next version of the game, this one was built entirely with Claude, because I wanted to show you an example here where we aren't using open models at all. I used Claude fable 5.1 to build this entire thing. And yes, for the sake of tokens, I couldn't do something super comprehensive, so I unfortunately wasn't able to make something that looks like super beautiful, but I mean, this still looks pretty good.
Definitely a lot better than just the open models. And let me go over to the pilot here. I wish I could run, but I'll go over to the pilot and show you that things look a lot better. Like if I take the controls here and I change the heading, you'll see that like I mean it's a little bit choppy, but I'm moving towards the clouds and then when I go forward, the clouds are actually getting closer. Like the navigation on the planet is actually working pretty well here.
It'd be cool to show you all the other stations as well, but it would take a lot of time. The point is that Claude Fable 5.1, it did a pretty good job for the very initial proof of concept for this game. I didn't allow that many tokens, but yeah, it's pretty good. The next version of the game, I am genuinely impressed. Not that it's like actually amazing, but the fact is I used GLM 5.3 Flash to write every single line of code in this version of the game.
And then for the planning and reviewing, I used GPT-6 Astra. And so this version of the game was about four times less tokens or like less cost overall because I'm using GLM for the implementation. And it plays pretty much the same. Actually, the moving is even smoother here going towards the clouds. I would say this version of the game actually is better even though I was using an open model to write every single line of code.
So, of course, it helps to have Astra review, but I'm still saving a ton on my cost/rate limit by using this combination of models. And so, yeah, a little bit of a cheesy game, very much just a proof of concept, but the point here is this is the results that I got for pretty much every application that I built. And remember, going back to the live bench here, Claude Fable 5.1 and GPT-6 Astra, they're pretty close in capability.
And so, it's pretty awesome that we were able to build even a little bit of a better game and other apps with this combination compared to just using Claude. That is the proof you need. You don't always have to have the best large language model for every step of your process. It is well worth your time, and I hope I've really proved it here, to analyze your AI coding workflow, figure out where you need the best model and where you don't because that is what's going to prevent you from jacking up your rate limits for your subscriptions like Claude and a Codex.
And so, my AI software factory is how I've been doing all my testing, how I run my workflows now, but it whatever you're currently doing, this this lesson applies. Like you should really explore open models at this point because it's becoming clear that we can't always run on our frontier models. So, hopefully you found this testing interesting. As I'm going through all these different combinations, building out different apps.
If you did, I would really appreciate a like and a subscribe. Follow along as I continue to build out my software factory and test these combinations. And with that, I will see you in the next video.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.