
Never Hit GPT 6 Astra Usage Limits Again transcript
Dubibubi · @Dubibubiii
Words
3,073
Runtime
16:20
Speaking pace
188wpm
Reading time
13min
188 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
I reduced my Codex usage by up to 91.75% and I did this by making GPT-6 Astra run hundreds of tiny experiments to find ways to optimize the way I was using Codex. And what I found might actually surprise you. I used to hit Codex usage limits every single day. However, after implementing a few practical rules for how I use Codex, I haven't hit my limits in over a week. >> Oh my god. >> So, without further ado, here are 11 things I changed that stopped me from hitting Codex's usage
94 words, the words spoken in the first 30 seconds at 188 words per minute.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 217 |
| Average words per sentence | 14.2 |
| Longest sentence | 50 words |
| Questions asked | 2 |
| Sentences containing a number | 33 |
Most used terms
- astra35
- codex32
- usage19
- actually17
- rule14
- gt12
- tokens11
- file10
- job10
- luna10
- work10
- agent9
Filler phrases
38 in total: actually 17 · like 14 · literally 2 · you know 2 · basically 1 · kind of 1 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Transcript
I reduced my Codex usage by up to 91.75% and I did this by making GPT-6 Astra run hundreds of tiny experiments to find ways to optimize the way I was using Codex. And what I found might actually surprise you. I used to hit Codex usage limits every single day. However, after implementing a few practical rules for how I use Codex, I haven't hit my limits in over a week. >> Oh my god. >> So, without further ado, here are 11 things I changed that stopped me from hitting Codex's usage limits.
And at the end, I leave it to a live demo so you can see just how crazy these savings actually are. Oh, and if we haven't met yet, I'm Duby. I build apps every day using Codex and I've generated over $50,000 in 80 days using these AI tools. So, the very first thing I need you to do in order to make your Codex usage last longer is to throw everything you know about AI out the window. No, I'm actually serious because with the rollout of Astra, OpenAI has made tons of tiny tweaks that make a lot of your previous token saving tips completely useless.
So, let's play a game. Here I have a bunch of popular token optimization methods and I want you to guess how many of them still apply for GPT-6 [music] Astra. Now, if you guess three, you'd be completely wrong. The answer is none of them. None of these apply to Astra. >> Embarrassing. >> Keeping context under 270,000 tokens is not even necessary anymore. Catch resets every 5 [music] minutes? Well, now they've extended it to 30 minutes.
Oh, and don't worry about changing effort levels in the middle of your conversations because Codex [music] won't reread your entire chat history anymore. And the point I'm trying to make here is that as AI evolves, we also need to adapt and create new methods that work for the next generation of models. So, let me start with rule one. Give Astra a usage budget. With the rollout of Astra, OpenAI released a very important update that now lets your models see their own usage limits, meaning Astra can actually check how much usage you've got at any given moment.
So now, unlike before, you can actually budget. Hey, complete this task in 3% of my weekly usage limit, and Astra will try its best to stay within that budget. And I found even just mentioning token budgets will make Astra work much more conservatively, but here's where it actually gets really interesting. Because usage limits are now visible to the model, prompts such as run experiments to find out how we can reduce this workflow down to 1% are actually viable now.
And OpenAI's own documentation supports this, saying Astra exhibits behavior patterns that can be optimized through prompting the model for your use case. Now, you may need to give up a little bit of your usage up front, but for repeatable workflows like reading emails and drafting scripts, you could easily reduce your token usage by 80 to 90%. Just to give you an example of how effective this is, I'm filming another video right now where I use Codex to run a business selling digital products.
The first product we created used 4% of my weekly allowance. So, I gave Astra some room to experiment. I told it to optimize the workflow with a goal of getting a digital product made for less than 1%, and Astra was able to absolutely smash that goal and get it down to three products made for only 1% total of my usage limits. With these optimizations, I was able to get 300 digital products made from the same allowance that initially would have gotten me only 25.
That is the power of asking Astra to optimize the way it works. If you're finding the video helpful so far, be sure to give it a like. I make videos like this every week, so subscribe if you want to see more. Rule two, let Astra manage cheaper agents. Now, rule two is extremely important because unlike previous models, GPT-6 Astra specifically has been trained to be able to divide and delegate work to sub-agents that work in parallel, meaning OpenAI expects you to utilize sub-agents when working with Astra.
Because the truth is, you don't need Astra to personally execute every part of your plan. I had Astra orchestrating Luna Max agents, and Luna's token rates are around 98% lower than Astra's. OpenAI actually cut Luna's prices by 80% recently, which makes it seriously worth trying for routine work. Benchmarks even show that Luna Max is equivalent to Sonnet 5 at Max, but for 1/7 the price. The trick with Luna models is to give it a very narrow job.
Say you want AI to help manage your inbox. Instead of one model handling the whole process, like Astra, you break it up into steps. For example, one Luna agent could sort the emails and flag what needs attention, another could draft the replies, then you have Astra review the responses and handle anything complicated. And I'm not telling you to manage all three agents manually. No. Just tell Astra to manage Luna Max sub-agents for narrow execution and review at the end.
That's literally it. If Astra ends up redoing everything, then the task was clearly too big and you'll need to actually break it down further. Rule three, quick cheat code, if you're on the $20 plan, set a cron job to literally just ping Codex with one tiny message at 6:00 a.m. While you're still in bed, your 5-hour window starts then, not when you sit down to work. So, let's say you work at 9:00 a.m. and you hit your limit at 10:00 a.m.
Well, your reset lands at 11:00. Rule four, turn repeated mechanics into reusable skills. Once you found a process that works, save it. Don't make Codex figure out how to do the same job from scratch every single time. With your digital products, we kept repeating the same steps: package the artwork, create the previews, build the download, and check the files. So, we turned these steps into a reusable workflow. You can do this with a weekly report, a client proposal, your content research, whatever you regularly create.
Just tell Codex, "Turn this successful workflow into a reusable skill." The next time you need that job done, just point it at the skill. And unlike the old days, skills are actually cheap to keep around now. At the start of a task, Codex only reads each skill's name and a one-line description. It opens the full instructions when it decides to actually use that skill. So, skill you don't need today costs you almost nothing.
You only really ever need to iterate and go through a workflow manually once with your AI agent. Then every time after that, the skill can explain it. >> Rule five, save your preferences once and stop introducing yourself to Codex every single time you open a new task. For example, act as a professional whatever. Keep it short. We're on next.js. That's all token burned on repeat and usage spent teaching it the same thing again.
There are two places to save this. To save instructions universally, go to settings, personalization, and custom instructions. Here you can save whatever you want. I like to add notes around my writing style and how I want answers formatted. For my terminal users, this updates your personal agent.md file, and Codex reads that at the start of every task. If you're looking to add instructions that are more project-specific, just ask Codex to do it.
Save my writing preferences in this project's agent.md, so future tasks use them. It will then create and edit that file by itself. And if you keep correcting the same thing, save that correction. Otherwise, you're paying to teach it again in every new chat. Now, it's also worth noting saved instructions still take up context on every task, so keep them short. The savings come from skipping the repetitive prompting. But if your instructions are bulky and unused, you'll just lose unnecessary tokens again.
So, to avoid things like that happening, every month or so, just ask Codex, "Review my agent.md and skills. Find anything duplicated, outdated, or making simple tasks unnecessarily complicated. Show me what you'd change." Rule six, correct it while it's still working. If Astra is building the wrong thing, don't just sit there and let it finish. Go to settings, general, follow up behavior. You got two options here. A message you send mid-run can steer the current run, or you wait for the next one, which is queuing the message, the default.
Pick steer. When you type no mid-prompt, Codex gets the correction while it's still working, instead of finishing an entire job you're about to throw away. It doesn't refund the tokens already spent, but it does stop you from wasting more tokens on a bad run. Rule seven, turn off the features you're not using. OpenAI's pricing clearly states model choice, context, reasoning, tool use, retrieval, and catching all affect usage.
So, every tool Codex has switched on costs you something, whether you use it or not. Start with plugins. I find myself installing plugins for one-off use cases, and I always forget to turn them off. It's good to go here and just turn things off. You might be surprised at how many random things you've got here. And also watch image as well. It's the same usage pool as your chat, but burns through your limits three to five times faster than a normal prompt.
So, don't ask Codex for random generations. Now, if you're not sure what's switched on, type {slash} status in a fresh chat. It shows your context usage and your rate limits. If a chunk of your context is gone before you've typed a word, that's instructions, tools, and plugins. It won't tell you directly which plugins are eating up your context, so switch things off one at a time and see what moves the needle. The rule of thumb is if you didn't turn it on intentionally for this job, turn it off.
Okay, if you've somehow made it this far into the video, then you're in luck because I'm looking to grow a small exclusive community of people where I teach you everything I'm learning about AI from the basics up to growing and launching your own apps. Now, I'm not really announcing this anywhere. I'm just going to keep hiding it at the end of my videos because I only want people who one, actually have an attention span, and two, are genuinely motivated to build and launch something.
So, if that sounds like you, I'll leave a link below. Rule eight, stop making it read every successful check. Every time Astra does part of a task, it makes what's called a tool call. Opening a file is a tool call. Running your website to see if it works is a tool call. Searching the web is a tool call. And every tool call sends a report back to Astra. It then has to read the whole report before it can take the next step.
But, here's the problem. Those reports are long by default. Say you ask Astra to look through 200 files for a spelling mistake. It doesn't get back a note saying found two. It gets back a full write-up on every single file, including the 198 that were fine. And it reads all of it, and you're paying for it. Imagine you hired someone to do that job, and they came back and read you a full report on every file that had nothing wrong with it.
So, tell Astra the same thing. Keep the reports from the tool calls short. When something works, just say it worked. Only give me the details when something goes wrong. We tested this across six matched comparisons. The same job with shorter reports used about 6% fewer tokens. And this can compound pretty quickly. Everything still got done, and the problem still showed up. Astra just stopped reading pages that marked if something was fine or not.
And again, if you don't want to say it every time, tell Codex to make it permanent. Now, it's worth noting, be careful because if you cut it too short, it can miss the one line that explains what broke. So, try it on a real job before you use it for production. Rule nine, stop paying for essays you didn't ask for. Every word Astra generates adds output tokens. In English, a token is roughly 3/4 of a word. So, a 100-word response is around 130 tokens.
Now, that's fine when you actually need the explanation, but if you asked it to change a button, you probably don't need six paragraphs about its incredible button-changing journey. Now, there are plenty of ways to reduce Astra's word vomit. Skills like Caveman, which teach it to talk like a caveman, which is also kind of entertaining. Or my personal favorite is I have ADHD, which forces Astra to speak more concisely with line breaks.
If you don't want to install any skills, OpenAI also recommends adding this prompt to your agent's startup D file, which should also do the trick. The rule here is to make the response as short as the job allows, because this will cut out the output portion of your usage. It also saves you from reading a novel every time Codex finishes something. And when you want the full explanation, you can just ask for it. Rule 10, stop making Astra guess where to start.
If something's broken on your website, don't just say debug my website, tell it which page is broken, what happens when you click the button, and what should actually happen instead. If you know which file controls that page, give it the file. People don't realize how much of their usage gets wasted from the model just searching for stuff. We actually tested this ourselves across six comparisons, giving Astra the correct file path used around 23% fewer total tokens.
A few extra words in your prompt can save an entire investigation, and if you don't know where in your code base it is, it can be useful to have an agent map out your code base and create a skill so future bug hunts can find the problem faster. Rule 11, make it leave a record of what's finished. Nothing is more frustrating than watching AI spend your usage doing something it already did. So, for longer jobs, have Astra keep a short progress file what's finished, where the files are, what failed, and what needs to happen next.
For example, if it's drafting emails, record which conversations already have drafts. Then if something crashes or you come back tomorrow, it has a clear starting point. In my current workflow, saved assets and draft ideas let me resume after interruptions. You've already paid for that work, so make sure the next attempt can pick up from it. All right, guys, as promised, just to prove to you that I'm not lapping, here I have live coded a dashboard for Vanilla Codex as well as the enhanced codex, which is going to include all of our rules and optimizations that we talked about in this video.
And we're going to be able to see how much tokens they used to fix this bug. So, what is the bug? So, basically, we have this website here, and you can click through these buttons to change the background. However, we also have a dummy website here, and these buttons don't work. So, the bug is really simple. Fix the buttons so they work. So, we're starting off with vanilla codex over here. We've gone ahead and sent it the prompt.
Let's see how it performs. We're not telling it where to find the problem. We're not telling it to use our sub agents. We're not turning off our plugins. This is completely vanilla how someone would normally prompt their codex agent. As you can see over here in our dashboard, the tokens are going up for vanilla codex. Okay, boom. Looks like it's fixed. Let's go ahead and have a look. And, yep. Seems to be working. Nice.
Looks good. Now, for the enhanced codex session, we're going to be sending it a much more specific prompt, telling it exactly where the components are, giving it a budget to follow, telling it to delegate to Luna Max sub agents. We're also going to be turning off a lot of our plugins that are unnecessary. So, all of these are going off. Also, rather than just giving it the caveman skill, I've just told it to talk like a caveman.
That should also do the trick. So, again, we have a faulty website. And the enhanced codex agent is off. Let's see how it does. Now, it is worth noting when you are using Luna Max agents, uh it does go a little bit slower. So, I am expecting it to be slower than vanilla codex, to be honest. This is about saving your tokens, not spending them as fast as possible. As you can see here, we have one Luna max salvage and running right now.
All right, enhanced codex is now complete, so let's go ahead and refresh the page. Let's see. And yep, boom, working just as expected. Now, let's have a look at the results. Boom, as you can see, we got a total of a 33.3% saving. Now, it's worth noting because the debugging task itself was quite small, it's not going to be like a massive difference. Like, you're not going to see a 90% difference here. But, I guarantee that if you actually give this a long-running task, if you actually give this a much harder debugging tool or you're building an app, this saving is going to be much, much higher. >> And there you have it, 11 practical ways to never hit your codex limits again.
Like the video if it helped, and check this video out. YouTube thinks you'll love it.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Use this transcript
Three free tools that work on the material around a video like this one. No signup, no login.
Hook Analyzer
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Policy Pre-Flight
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Channel Skill Generator
Read this channel's public videos and transcripts, and download a writing brief for it.