Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
22:016.2x the video's typical replay level
would feel like a superpower to them, and now they're up to 90. Waste less time and hire better engineers at soidev.link/g2i. It is worth noting that a lot of analysis is showing that Terra isn't really a great option, which is disappointing. I was quite excited about it. Artificial Analysis said the
Said at 21:54
Most replayed moment #2
19:185.6x the video's typical replay level
with with results there. And crazy enough, Tibo even blessed it saying that if anyone is banned for using my weird hacks to get my Codex sub into Claude code, that he will give you a reset. So, yeah. This is a blessed solution. Very fun video on that coming in the near future. Just wanted you to know that
Said at 19:10
Most replayed moment #3
24:434.2x the video's typical replay level
didn't talk about this one too much here, but it is really fun. You can let other agents steer. For example, I set up my Claude code to allow for it to call Codex. I've went much further with this since and that's the Codex in Claude code video coming soon. Where I'm actually using Soul as the model in
Said at 24:37
The graph counts replays. It does not show where viewers stopped watching.
Words
6,281
Runtime
30:09
Speaking pace
208wpm
Reading time
26min
208 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
GPT56 is an incredible model, but it's also incredible at burning through people's rate limits. It used to be really hard to hit 5-hour limits on GPT55, even X high. Now a 56 solid draining super fast even on medium reasoning. I've already burned through three 5-hour limits. Seriously, this has to stop. I've now set 56 from high to medium, not fast mode, and I'm still burning through my rates at an insane rate. My 5 hours are almost gone again. I've hit the limits every single time since Soul was released, and I don't think I ever hit it with 55. Bro,
104 words, the words spoken in the first 30 seconds at 208 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 410 |
| Average words per sentence | 15.3 |
| Longest sentence | 56 words |
| Questions asked | 3 |
| Sentences containing a number | 111 |
Most used terms
Filler phrases
39 in total: like 28 · actually 6 · kind of 2 · basically 1 · literally 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
GPT56 is an incredible model, but it's also incredible at burning through people's rate limits. It used to be really hard to hit 5-hour limits on GPT55, even X high. Now a 56 solid draining super fast even on medium reasoning. I've already burned through three 5-hour limits. Seriously, this has to stop. I've now set 56 from high to medium, not fast mode, and I'm still burning through my rates at an insane rate. My 5 hours are almost gone again.
I've hit the limits every single time since Soul was released, and I don't think I ever hit it with 55. Bro, I've waxed a weekly in a half on one message in X high. With GPT55 on the $100 plan, I could barely use up my quota. Now, one small PR with 56 Soul consumes about half of my 5-hour limit. Not going to lie, GPT56 Soul is great and all, but the usage drain is whack. Used to be able to work consistently all day with 55 on X high.
Now I can't even get through a single 3-hour session without hitting limits. Whack. Not good for long-term work. As you can tell, people aren't happy, and I'll be frank, a lot of this is OpenAI's fault. There are some things in Codex right now that just don't make sense. I want to do my best to help you guys out to make sure you can get the most out of 56 without hitting these limits. I made a bunch of these changes myself, and I've actually noticed the quality of outputs going up while also using as little as a fourth or a fifth as many tokens as I was before.
I took the time to write an article sharing my advice here, but since then quite a bit has changed even though it's been 24 hours. I've also seen a bunch of really bad advice going around that has been confirmed by many, including the actual team at OpenAI, to not be good. So, if you want to get the most out of your subscription plans with Codex, I hope this is helpful. And even if you don't, maybe you're using something else like Cursor or Claude Code, this video will still have a lot of tips that could help you be better at using those in an efficient effective way.
I have never burned as many tokens as I have with 56 Soul, and I've really come to know how this model works, what its strengths and weaknesses are, and what gets it to burn. I want to take the time to share this all with you, but this experimentation was not cheap, so So, hope you can pardon a quick break for today's sponsor. As AI writes more and more of our code, the chances of it getting things wrong goes up, too.
And as smart as the models are, there are certain places you just shouldn't take the risk. And if you're trying to land enterprise customers, the auth layer is one of the worst places to take that risk. That's why WorkOS has always been a great option. They really understand enterprise, and they don't compromise on the developer experience to get there. They have really good SDKs, really good integrations, really good docs, and everything else you need to set up your apps in a great way.
They also give you up to a million users for free, which is just insane, and I can't imagine many of us having that problem. I I wish I had enough users to hit the limits and need to pay money, but I don't. So, it's a really generous offer. But nowadays, users aren't your only users. Your users have agents. How are those agents going to auth? And if a company like Microsoft wants their agents to use your service, how are they going to get in?
WorkOS is one of the few companies thinking about this, which is why they partnered with many others to build Auth MD. They're working with companies like Cloudflare and Firecrawl to get all the pieces necessary, so your agents can auth for you, can create accounts for you, and then let you attach them to your service account, whatever, in all of the logical ways that are needed. To be very, very clear, Auth MD is not a WorkOS feature.
It's an open standard that they helped build. They're the authors of the protocol, and obviously that means you can one-click turn it on in your WorkOS apps, but this is a standard that anyone could implement, and I'm excited for everyone else to, as well. WorkOS is great because they let you sell to enterprises now, but they also set you up for a future where the agents are signing up instead. Get enterprise proof and agent ready at soidev.link/workos.
Time to talk about how to use Soul without hitting limits. The model is great, and the $200 plan is still great. Even the $100 plan is pretty generous. I've not sure what the state of the $20 plan is right now, but the $100 one is fine, and the $200 one is really good, as long as you don't make certain mistakes. And again, to your credit, OpenAI is not doing this right. There are a lot of little things that exist in Codex right now that are kind of forcing you to overuse your usage, and this is their fault.
I feel bad. If people in chat are saying the $20 plan is actually pretty good for coding, that's good to hear. I One person said it's horrible, but others are saying it's fine. So, yeah. Take it with a grain of salt. Apparently, it's okay. I am on the $200 plan, and I was hitting limits aggressively when they made this move. Also, full disclosure, when I was testing 5-6, it didn't count against my usage, which is why I was able to push the model so hard and do a genuinely absurd amount of inference on it.
But, as soon as it came back just a day or two before it went public, I started hitting limits immediately. It was 20 minutes after they said, "Okay, you have it again." that I hit a limit for the first time. And the reason I hit that limit was a certain new feature called Ultra. I have a dedicated video about Ultra coming very soon. It might even be the next video after this one. So, as always, make sure you subscribe and hit that bell if you want to better understand all of these pieces.
If I was to talk about Ultra in depth right now, this video would be 2-plus hours long. I want this to be to the point and easy to apply to your work and send to co-workers and whatnot. So, I won't do that here. But, know that Ultra is not worth using at the very least until you wait for that follow-up video. So, for now, Ultra, no. Ignore this. Don't touch it for now. This could might change in the future. I am hopeful that it will, but for now, Ultra is best avoided, not used.
With Ultra out of the way, I do want to talk about a thing that just changed cuz it's quite important. OpenAI is aware of the issues. I've been working with them a bunch on getting these things fixed. Hopefully, more of them are fixed before this video is live. Probably not too many, though. We'll still have useful tips long-term, but I want to give you guys context on where things are at the moment of filming. Teemu posted this morning the following update, which is very relevant to what we're talking about here.
The last 48 hours of Codex and ChatGPT work have been intense. There's three important updates. First is that they temporarily removed the 5-hour usage limits for all plus, business, and pro plans. If you're not familiar, your sub is broken up into two limits. You have a 5-hour limit that resets every 5 hours starting from when you send a message. So, let's say you do a bunch of work at the start of the day, you start at 9:00 a.m.
You have a limit to how much you can use between 9:00 a.m. and 2:00 p.m. And then let's say you hit that limit at 1:00 p.m. You have an hour until you can use it again. But if you don't start until 3:00, the next 5-hour limit doesn't start until 3:00. One of the tricks I have is I have a cron job running that does something every 5 hours so that I'm always burning one of those 5-hour limits. It uses a very small model with a really simple like hello, hi prompt.
But now you don't have to, at least in the interim, because the 5-hour limit is gone for now. The weekly limit is the one to be scared about. The weekly limit is equal to roughly four or five of the 5-hour limits. So, if you go hard enough to max out the 5-hour limit four plus times, you will hit the weekly limit. Weekly is the only one applying right now, which is good because the 5-hour limits were a little too easy to hit, but it's bad because if you accidentally ran an ultra fast run that is way heavier than expected, the 5-hour limit would have stopped it before it used more than 25% of your weekly.
Now it won't. So, it's possible to run one prompt and blow through your whole weekly limit. I'm tempted to do it as an example and then burn a reset, but I'm I'm hoarding other resets. I'll be real, guys. So, with all of this known, let's go through the rest here. The second point he made is rolling out changes that will make 5.6 Soul more efficient across the board and it will be reflected in less usage being used so it can take you further.
Exact impact to be quantified and shared. Still working on that there. And his most exciting point here is that they have 6 million active users and they're about to land a usage reset. That one already landed. So, as was established prior, they got rid of the 5-hour for now and we just have the weekly limit. I fired one prompt off before streaming, so I don't have much used here yet, but I'll probably burn through this in the next day or two and I'll be sure to share how, why, and what I'm doing to minimize and maximize usage.
With all of this established, it's time to give practical advice. Part of how I want to think about this is with 5.5 versus 5.6. When 5.5 came out, I was far from the biggest fan. I saw how capable it could be, but I didn't love using it because it felt like it would just lose track of what I was doing. It was really easy to screw up its context. If it read the wrong file and had something in its history, it would fixate on that instead of doing what I asked it to do.
And this happened a lot. It also was really, really bad at compaction, which was just obnoxious when you had long-running threads. I've never had more threads than I did with 5.5 because I felt like I had to in order to keep it on task. None of that is why 5.5 was cheaper, though. The main reason 5.5 was cheaper is that it stopped and asked for permission or feedback all of the time. It would say, "Here's step one, here's step two, here's step three through 10." And then you would say, "Okay, go do all of the steps." It would get through step one, it would get halfway through step two, and then it would stop and say, "Okay, I finished step one and I'm halfway through step two.
Let me know when I can keep going." It's like I never told you to stop, bro. There's all of this behavior is that 5.5 didn't really use your limits heavily on a per message basis. From my experience, a single message on 5.5, even on the highest reasoning levels, would use between like 0.1% and 2% of your 5-hour limit per message. It wasn't that bad. 5.6 fixed these problems, but in doing such, massively increased how much usage you're getting. 5.5 was a price bump from 5.4.
They went from $15 per mil out to 30 per mil out, which would have hurt a lot more if it wasn't for this sudden stopping behavior from 5.5. 5.6 fixed that behavior, but as a result, it can go much longer per message, which can result in much bigger usage per message. From my experience on higher reasoning levels, especially X high and max, 5.6 can use up to 15% of my 5-hour limit in one message. This is a big jump, obviously, but honestly, it's not too too bad because I would end up using similar amounts of 5.5, but it would be prompt after prompt after prompt.
Because each message only would use so much of my limits, I didn't find it too brutal, and I would often use it with fast. Fast mode is pretty cool because you get 1.5x faster speeds, but as a catch, you go through your limit 2.5x faster. And this wasn't too bad when I was using 5.5 because when I sent one message with 5.5, the 2.5x more burn would make that 2% into 5%. That wasn't too bad. 2.5x burn on a message that was 15% though, that's a little more brutal.
That is closer to half of your 5-hour limit from one message. And that's not even with ultra. That is scary, and that is a problem I think a lot of people are having right now is that they are continuing to use the model the way they did with 5.5. And with 5.5, since each message was less usage, fast mode didn't feel too bad. So, if you're hitting limits and you're using fast mode, please stop. It's not that much faster.
The reason I think fast mode felt so good with 5.5 is because it would stop itself constantly, and you had to talk back. I never watched my agents run quite as much as I did with 5.5 both because its context was constantly getting screwed up, and I had to prune it, clean it, make a new thread, but also because it stopped all the time, so I would have to respond and tell it what to do next. As such, the speed mattered a a because I was sitting there waiting for it to respond.
With 5.6, I spin it up and then I go do something else. Maybe I spin up another thread. Maybe I go do code reviews. Maybe I check email. Maybe I play Power World. I don't care. With 5.6, I do other things. With 5.5, I have to sit there and watch. So, fast mode felt important. With 5.6, it's running for so long anyways that the inference is barely the slow part. It's usually the tool calls, the test runs it's doing, all the other things it has to do.
I've barely noticed a difference in speed since turning off fast mode, but I have noticed a massive shift in the amount of usage that I'm hitting with it. For context, there's lots of people in chat including those on my team like Maria that says almost every prompt and thread they've done with 5.6 has lasted over 8 hours. Yeah. You can get these models going for a while. Even without using {slash} goal to be clear. I almost felt like {slash} goal was necessary with 5.5 because otherwise it would stop and the goal would basically tell it to keep going over and over instead of having a human do it.
With 5.6, not necessary. I never use {slash} goal anymore unless I'm just trying to burn tokens. Generally with 5.6, it will complete the thing without needing extra encouragement. So, now establish two things. Don't touch ultra. You should probably turn off fast mode at this point cuz it doesn't matter as much as it used to. But, there's more that we can learn from. Reasoning level should be an easy enough one to cover.
I'll dive through this super quick. You have five options. They've been rebranded, but they used to be low, medium, high, ex-high, and max. You might think ultra is a reasoning level. It's not. Again, video crashing out about ultra coming very, very soon. Hit that button if you haven't at the bottom. Subscribing makes it more likely you see that when it drops. So, with ultra removed, we have these five options. I did a really good analogy about what these are and how they work on my most recent podcast episode that should hopefully be out in the next day or two.
If you're not already listening to the podcast, check that out as well. Ben and I just nerd out about the details here and it lets me go a little longer than I try to on the videos. Trying to keep these shorter lately. Hope you appreciate the effort. So, instead of breaking down all of the details here, I'll give it to you simple. Ignore everything past this line for now. X high can be really cool. Max is less likely to be.
Low, medium, and high are all very good options. I mostly just default to high right now because the model is so efficient even on high reasoning. If the task is simple, it will stop really quickly and not reason too much. When I was benchmarking the model, I thought there were bugs in how I was passing the reasoning level because the difference between low and high was like 50 to 100 tokens at most. It just didn't seem right at all.
Turns out it is actually just good at that. It will use less tokens if the task is simple. So, leaving it on high has been fine for me. I do want to look at the numbers quick though. So, let's do that with Deep SWE. This is my favorite code benchmark right now. Very much subject to change eventually. I'll throw in Fable for a comparison here. So, with GPT-4-6 Soul, on low, you score 45% and in this bench it costs a dollar per task.
Medium bumps all the way from a 45% to a 61% for a dollar and 86 per task. And then high bumps you up to a 69. Nice. Which puts you neck and neck with Fable on this benchmark, which gets a 70. But, the cost is $3.47 per task versus Fable at the same score being $13.00 per task. So, that is super efficient. All three of these are meaningful jumps in score. You get from a 45% on low to a 61 on medium to a 69 on high. But, then we see the next two levels like X high, which goes from a 69, nice, to a 71, less nice, but also bumps the price meaningfully.
Going from 347 to 470 per task. And while we're not paying API rates when we use it through the subscription, the API rates are a good measure of what you're burning. So, yeah, that's not great. And then when we bump up to max, the price doubles to $8.39 for an additional two percentage points. I don't think that's good. Medium is $1.86 and gets a 61. Max is 839 and gets a 73. And again, high is 347 for 69%. That's a more than double cost for a 4% bump.
Not worth it. Just stick with high. If you disagree, that's fine. I just hope you have the budget to handle your disagreement there. And for full transparency, there are other benches that have shown different numbers. For example, in cursor bench, high only scored a 63.5 and max scored a 67.2. And the price gap there was 279 to 569. So, it wasn't quite a doubling of price and it was a much more meaningful percentage bump.
So, this will depend on the work you're doing. It's also worth noting that in cursor bench 3.2, Fable ended up scoring quite a bit higher than Soul did, which we'll talk more about Soul versus Fable in the near future. That will be a big video, too. I don't want to be sidetracked by that. The point I'm trying to make here is high is really good and past high is when you lose this like vertical line where the cost jump isn't that big and the success jump is.
I just decided to quickly check the artificial analysis intelligence versus cost index. And when I turn off everything other than 5.6 Soul and Fable 5, 5.6 Soul high is the only thing in the good quadrant of cost to intelligence. It does meaningfully go up for each bump, but again, 5.6 Soul seems to be this really solid end of the vertical curve. So, reasoning levels are now covered. Now, we need to talk about what is probably the biggest behavioral difference with the 5.6 model specifically both Soul and Tyra.
We need to talk about sub-agents. There's a lot of different ways to implement subagents, but you generally need a tool that allows for your main top-level agent to spin up other work for other agents to complete. I could say that Codex's implementation of subagents isn't good, but that wouldn't be fair because they have two implementations of subagents called V1 and V2, and they're both not good. I have a lot to say about that.
Again, the ultra video will go more in-depth there. I will resist the urge to crash out hard about it, but what I will say with relatively high confidence is use with caution. There's a problem, though. You might want to use subagents with caution, but 5-6 was trained to use them with less caution. As such, it's not uncommon for 5-6 to spin up subagents for things where it doesn't really need it. If you notice this happening, I don't think you should rush to go change anything about this just yet because the subagents can be pretty good.
And I'm also hoping that the Codex team fixes the issues I have with it as they currently stand. But if you are noticing limit burn after applying all of these other recommendations and you're seeing subagents spinning up a bunch, I would recommend doing a slight adjustment to your global agents.md file. The easiest fix for now is to add the following to your agents.md. Only use subagents if the user explicitly requests them.
This change will pretty much stop them entirely unless you request, and I usually request when I want them. You'll eventually figure out when to or not to use subagents for work like what work justifies splitting things up that way. So, don't be hesitant to experiment with them. I think they're really cool, especially when applied carefully, but right now 5-6 is a little too eager. If you need to tone it down, here's the strategy to do such.
It is worth noting that the subagent implementation in other tools like Claude Code and Cursor is meaningfully better. So much so that I've actually done some hacks in order to get Claude Code working with my Codex sub, and I've been really, really happy with with results there. And crazy enough, Tibo even blessed it saying that if anyone is banned for using my weird hacks to get my Codex sub into Claude code, that he will give you a reset.
So, yeah. This is a blessed solution. Very fun video on that coming in the near future. Just wanted you to know that it's worth exploring other options. I know pie is a decent one as well if you're curious enough to explore, but I do expect Codex to fix this in the not too distant future. So, if you're doing it out of fear, just wait. If you're doing it out of curiosity, go nuts. One other thing I forgot with the reasoning level bit, and I'll just be super quick about this, is model selection.
TLDR, don't use Luna. It's just not for us to use for code. It's really useful as a thing that you access programmatically, like you're using it over API to filter data and whatnot. Not really what you should be using for coding. Tyra seems to be a good middle ground, but honestly, I would rather just use Soul on lower medium. They're really efficient. I just recommend using Soul high as your default, and if you find things that it goes longer than it should for, try out solo.
Maybe play with Tyra, but I have been totally fine just using Soul for everything. And once I made these other changes, I was not coming close to my limits. But there is one last change we have to talk about. The most important tip in this whole video. This is definitely the change that affected how Soul worked for me the most, and I'm really excited to share more about it right after a quick break for our sponsor. Today's ad's going to be a little different.
You've already heard me talk about G2i. They help you hire world-class engineers at whatever scale and time frame you need. Normally, I tell you all about how crazy it is that they can get you engineers in under a week that are actually experienced and ready to go, but instead I'm going to tell you about one of their customers, Batteround, because they've hired nine engineers through G2i, and they learned about G2i through me.
So, I think that's pretty cool. These guys needed to hire, and if you've ever hired, you know how grueling it is. You just lose all of your spare time, and it no work gets done. If you're spending time hiring, you're not spending time working, and it's this weird unintuitive thing where when you realize you need more help, you end up needing more help sooner because you're stuck doing other stuff. And they wanted to hire, but they couldn't afford to lose bandwidth.
Batter Around's leadership team had decades of experience hiring engineers in other sectors, but their network in the media space was small. It was getting hard to find qualified candidates, which is why they reached out to G2i. They wanted fewer resumes from better candidates so they can more easily evaluate. G2i already had a large network of engineers ready to go, and they filtered to the best ones fit for the specific needs of Batter Around in this mixed media space.
With G2i, Batter Around has contracted nine engineers, had a 90% success rate, reduced leadership time spent screening candidates, integrated engineers directly through their existing Slack workflows, and they've minimized disruption through the clean transitions and fast backfills when they need to shift around their staff. Batter Around themselves said that a 50% success rate would feel like a superpower to them, and now they're up to 90.
Waste less time and hire better engineers at soidev.link/g2i. It is worth noting that a lot of analysis is showing that Terra isn't really a great option, which is disappointing. I was quite excited about it. Artificial Analysis said the following: "5-6 Sole Luna are ahead of Terra at every point on the intelligence versus cost per task chart." This is the chart that I had just shown before. And here you can see that Luna on max scores slightly better than Sole on low while being just barely more expensive.
It almost looks like a pretty consistent line from Luna low, mid, high, x-high, and then Sole low, mid, high, x-high. But Terra underneath here, not looking great. The model needs to know when to stop. As I was hinting at earlier, 5-6 is very eager to get work done. It will keep going as far as it possibly can unless you give it a reason to stop going. I gave some good examples in my article, so I'm going to reuse those here.
This is the style of prompts I've learned to write when using Sole because without it, it can go further than I want it to pretty often. Example one is the following: "I want you to build this new feature. Start by writing a plan. We When you finish the plan, stop and ask for feedback before proceeding. I have told the model explicitly, do this thing and be done at this point. That doesn't mean you have to give the model prompts that are less work, though.
You can have the stop point be way, way, way further on. For example, the plan looks great. Let's build it out. Use computer used to test your implementation. Keep going until the code works and you're happy with the implementation. Put up a PR, baby sit it for the first set of review comments and address them. Stop after the first set of review comments. I'll handle it from there. Previously with 5.5, the model stopped way too often.
With 5.6, it's so unlikely to stop that I feel like I have to put up the stop signs myself. But now that I've been doing that, it feels much better. It's nice that I can start building my own intuition for how far the model can go without my intervention. And also that I can kind of choose when to stop it. Rather than trying to do this through reasoning levels or changes to your harness, so much of this can just be done via prompting.
And I will say I am proud of myself for getting this far in the video without my instructions being prompt better. Because your prompt can absolutely cost you more money, but there's other things that I thought were more important. This piece though, this can change how you use the model fundamentally. When you realize that the end is a thing you can define in the prompt rather than a thing you have to set up with tools, you can let the model go way further.
And the challenge I would present to you is see how far off you can make that stop sign without the results in the quality going down. I didn't talk about this one too much here, but it is really fun. You can let other agents steer. For example, I set up my Claude code to allow for it to call Codex. I've went much further with this since and that's the Codex in Claude code video coming soon. Where I'm actually using Soul as the model in Claude code.
Seriously, ClaudeX is way better than I expected it to be. And as silly as it is seeing this 5-6 soul there, it's been working great for me. Most importantly here, the thing I want to encourage is more experimentation. You shouldn't be trying to copy my skills, my agent.md, and my prompts directly because this should be different for everyone. You should experiment more with these different things. Try prompting different ways.
Try different reasoning levels. Try adjusting your codex, agents.md, and your claude.md files a bit to see how it changes behavior. You should be spending a good bit of time in the dot codex and dot claude directories on your computer if you're using these agents and models. It's so fun to customize cuz it's not like you have to be deep in the weeds of the details of how it works in all these config files. It's a markdown file and changes to that can fundamentally change the experience that you're having.
And the more you move things around, the more you channel our inner Dan Abramov to experiment and shift until things feel the way you want them to, and have your own feelings change in the process, that's what's going to give you the best experience using these things. If you blindly follow my recommendations or you go set up somebody else's recommendations like you install any of the oh my whatevers, you're never going to learn what does and doesn't work and you're going to be limited by the creativity of someone else.
Go play with it yourself. I don't have a bunch of other people's skills and things installed. I just dick around and change things and read the traces for my agents when they don't do what I expect and make adjustments until they do. I am currently deep in the process of rewriting my agents.md and this is all by hand. Not a line of this was LM generated and I think that is one of the best places to take the time because you'll be amazed how much these things can change how the model feels to use.
What I'm trying to say is that you need to stop looking for someone else's solutions on a shelf for you to buy, copy, paste, and hope for the best. You should use resources like this video and like these other examples people have as inspiration to get it the way you want yourself. I mentioned at the start that I've been seeing some bad advice. This is probably the worst one. While I don't trust opening eye to configure everything perfectly in Codex for us, they absolutely covered the context windows themselves just fine.
I've seen a lot of people making recommendations like this, specifying that you should limit the context window and when it auto compacts manually inside of your config. Don't do this. The model was trained on specific levels of compaction in tools like Codex. This is going to make the model dumber and it's not going to save you any money. If anything, it's going to cause compaction to happen more than it should and compaction is expensive.
Tibo jumped in and said as much. This is not correct. Do not do this if you do not understand exactly what you're doing. We do not charge extra above 270k context in the context threshold has been tuned for 56k sold to be perfect with a default limit. It is, trust him. Don't just blindly follow random advice on Twitter. At the very least make sure the people giving the advice aren't getting their advice countered by people at OpenAI.
I hate filming live sometimes because literally right after I finished that part and was ready to go offline, Tibo tweeted the following. We've laid in first optimizations to make it cheaper to run. It's going to be 10% savings, cool, awesome. They also noticed that changing the context size limit in the product to 372k up from 272k resulted in more usage being charged than intended. This means that their usage tracking was broken even though they planned to let this stay.
They've temporarily reverted to 272k and will roll back in the days to come. This should be a big change in how fast your usage drains, cool, awesome. Stupid that we're here, but here we are. They confirmed that the leaks around juice values being changed were true, but they've been reverted since. And they also call out that there's been more use of multi-agent than intended in high and X high reasoning efforts and they're fixing it going forward.
Also fixing small other things they noticed with auto review where they can be more efficient. Ah, reporting on the fly is annoying, but you now have the additional context. That all said, every piece of advice I gave in this is still true, so don't think this means my advice is invalid. Just know that this is a moving target. Best practices are still best practices. One more point in favor of lower reasoning levels, the open code folks are obsessed with 5.6.
They barely even use Fable, they like 5.6 so much. They like it so much that most of OpenAI has been promoting their posts throughout. I assume they were probably using it on higher reasoning levels because who wouldn't initially? They hit an issue though, one that I've hit when I was testing these things over API. They changed the key in the JSON for reasoning levels at some point, and a lot of tools don't pass through reasoning level properly.
I ran into this a lot during benchmarking. Turns out Open Code did too, and during all of their usage, a month of usage, they realized that they had it misconfigured to always use medium. Even though they thought they were using X high or other levels. And despite that, the whole team agreed it was their favorite model. Definitely an endorsement for medium as the default. So if you do like Open Code and or you trust the Open Code team, they have confirmed medium is incredible.
Medium and high are both really, really good options. So go play, go configure, go customize, and see what you can do. I have a feeling you'll be surprised. Keep playing and keep prompting, and until next time, peace nuggets.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.