YouTube transcripts

24 Hacks to Solve GPT-6 Astra Usage Limits (in 15 mins): video thumbnail

24 Hacks to Solve GPT-6 Astra Usage Limits (in 15 mins) transcript

Jay E | RoboNuggets · @RoboNuggets

Published September 19, 202614:5275K views

Watch this video on YouTube

Transcript analysisComputed from the caption text

Words

2,902

Runtime

14:52

Speaking pace

195wpm

Reading time

12min

195 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.

Opening (first 30 seconds)

GPT-6 Astra is incredible, but it has one big problem. It drains your usage limit faster than any other model, and you get locked out of your work until your next reset. So, here are 24 tricks to help you with that problem, plus two bonus ones that I haven't heard anyone else talk about. We're going from quick setting changes and free plugins you can switch on today, all the way to some advanced power user stuff. So, let's start off with the easy ones. The first is to understand effort levels. So, effort is basically how

98 words, the words spoken in the first 30 seconds at 195 words per minute.

Sentence shape

MeasureThis transcript
Sentences147
Average words per sentence19.7
Longest sentence56 words
Questions asked0
Sentences containing a number31

Most used terms

  • agent38
  • usage20
  • file16
  • number16
  • model14
  • task12
  • agents11
  • files11
  • astra10
  • codex10
  • effort10
  • start10

Filler phrases

25 in total: basically 9 · like 8 · actually 6 · kind of 1 · sort of 1.

A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.

What this transcript is

Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.

Transcript

GPT-6 Astra is incredible, but it has one big problem. It drains your usage limit faster than any other model, and you get locked out of your work until your next reset. So, here are 24 tricks to help you with that problem, plus two bonus ones that I haven't heard anyone else talk about. We're going from quick setting changes and free plugins you can switch on today, all the way to some advanced power user stuff. So, let's start off with the easy ones.

The first is to understand effort levels. So, effort is basically how hard your agent thinks. And from different independent tests, we now know that more effort doesn't necessarily buy you that much more in terms of output. For example, Artificial Analysis AI, a company that tests every model, found that extra high on Astra costs around 30% more than high for about 1 and 1/2 extra points in intelligence. So, if you're trying to save on usage, what you can do is to pick light for your effort first, and only move it up when your result isn't good enough, so that you only pay for the thinking that you need.

Number two is to skip max effort. So, a quick one while we're on the topic of effort, on GPT-6 Astra, going from extra high to max costs 40% more per task, but the intelligence score only goes up by less than half a point. So, unless you have an unlimited budget, the highest that I would really set effort to is extra high instead of max. Number three is knowing what ultra mode really does. So, in Codex, ultra is more than just a thinking kind of confusing since they put it in the same slider.

But what ultra does is max out reasoning and hand parts of the job to sub agents, which are basically these helper agents that work with you in parallel, which sound great, but you do pay for every one of those sub agents. So, what you can do is to just pick extra high effort when you just want more thinking, and keep ultra for the really big jobs that you think would need a team. The 6:00 a.m. ping. So, on a lot of plans, your usage limit runs on a 5-hour window that only starts when you send your first message.

So, what you can do is to tell your agent to set up a scheduled task that sends hi to itself at, let's say, 6:00 every morning so that if, for example, you start work at 9:00 and you hit your limit shortly after that, the window will reset at 11:00 instead of around 2:00 p.m. in the afternoon, which gets you back to work much sooner. And by the way, I've put every one of these tips into this detailed PDF guide, including all of the resources that I'll talk about in this video.

And you can just send this to your AI agent and ask it which ones apply to your setup to make it easier for you. The link for that is just in the description below if you need it. Five, pick your model before you start. So, in every session, your agent keeps a cache of your conversation, which is basically a saved short-term memory that keeps your usage low. But, if you switch models halfway, the new models can't use that cache.

So, it pays full price to read your whole chat conversation again. What you should be conscious of then is to set your model or decide on your model before your first message. So, decide if you really need Astra's power for that task or if Soul or Terra is enough. Let's stay cheap. Number six is about fast mode. So, it's tempting to switch on the 1 and 1/2 times speed setting because you want answers quicker, but it actually burns through your limit a lot faster for the same work.

So, by default, what you can do is to just switch it off and only turn it back on when you're really in a rush. And anyway, the lower effort settings that we talked about previously, they're already pretty fast or at least faster on their own. So, most of the time you won't even need this setting. Number seven, turn on mid-run steering. So, when your agent heads off in the wrong direction, a lot of people wait for it to finish and then correct it.

But, what you can actually do is to open settings in Codex, search for follow-up behavior, and switch it to steer mode so that every message that you send mid-run would change the current run instead of waiting in line until the task finishes. So, that if you want changes to be done in the conversation, you can do that instead of paying for work that you would throw away anyway. Number eight is Caveman. So, Caveman is a free open-source skill that does only one thing.

It makes your agent talk like a caveman. So, short, blunt sentences with zero fluff. From its before-and-after test, it uses about 55% fewer output tokens than a normal answer, which is great for usage savings and also just for keeping agent replies concise. So, what you can do is to paste the GitHub link into your agent and to ask it to install it to start using it. Another plugin that's somewhat related is this free tool called I have ADHD.

So, this is an open-source plugin that makes your agent lead with the answer instead of burying the lead in a wall of text. It has the same benefits of outputting less tokens. So, what you can do is just grab the repo link and give it to your agent. Your replies will now come back shorter and also a lot easier to read, so that you spend less time scrolling and fewer tokens spent on fluff. Number 10 is to ask your agent what's eating usage.

So, one of the easiest things you can do is to simply just ask your agent to look at your setup and tell you what is using up your usage. It'll typically flag plugins that you're not using, effort that is maybe set higher than the job needs, and helper agents that copy your whole conversation every time they start. This is quite an easy habit to get into, and it will pay dividends for you down the line. Okay, so those are the quick ones, and honestly, if you only did those, you'd already stretch your limit quite a lot.

But now, let's go to some intermediate tips. Number 11 is to tell it exactly where to start. So, a lot of your usage goes on your agent searching through your project for the right file. So, whenever you already know which file it should look at, Just copy the file path of that and hand it over to your agent. On Windows, the shortcut to do this is control shift C and it's much faster and uses far fewer tokens than having your agent find the right file.

Number 12, make a map of your workspace. And this is important so that your agent can find the right files much faster and with a lot fewer tokens. I do this myself with what I like to call router files, which are basically short markdown files that just point things live in my workspace. And I've got one for each of the big departments of my life and my work. So, content.md is a map for content, product.md is a router file for the apps we're building and so on and so forth.

What you can do is to ask your agent to write one short router file per area of work that lists where everything lives and also to update its agents.md to list these router files. It then reads these routers and opens only the file that it needs instead of loading everything and searching for things blindly. A related one is to make agents.md a router by itself. So, your agents.md file, it gets read every time you start a new chat, which means that every line in it would drain your usage on every single task or session.

So, what you can do if it gets too long is to move each section of your agents.md into its own file and just leave one line behind that just routes to that file. So, your agent would read now a short list of pointers instead of a long dense handbook and only opens that it actually needs for the job. Number 14, our second brain systems. And when I talk about second brain systems, I don't mean just pointing a tool like Obsidian at your workspace to get a fancy visual.

A proper second brain system, in my view, is basically an index of everything that you've decided, saved, and built where your agent can look things up really quickly in a single step most of the time. So, instead of opening 20 files to find one answer, it looks it up systematically and goes straight to the right file. Now, this is a much larger topic, so I'll link the lesson where I teach how to build this somewhere in this video. 15, switch off unused connectors.

So, every plugin and MCP server you have switched on gets read and actually uses tokens in every chat you start, even when you don't use them. So, if you want to clean that up, you can just send this prompt to find the connectors that you can already pause. And of course, you can always just switch them back on whenever you need them. 16, monthly cleanups. So, instruction files, skills, and even the plugins and MCP servers you install pile up over time.

So, if you ask your agent to set up a scheduled monthly task to review your instructions file, your skills, your plugins, and flag anything that is duplicated, out of date, or too long, then you can either have it delete or archive whatever it flags. As an extra tip, run this at the end of your usage week so that you're spending usage that was about to reset anyway. Number 17 is about image generation in Codex. So, making images inside Codex draws from the same usage pool, and it burns through it around three to five times faster than a normal prompt.

So, if you don't want images draining your usage, what you can actually do is to connect a separate AI image provider to Codex. The one we use the most is Key AI, which gives you ChatGPT's image model for about 3 cents an image, which is relatively cheap and it doesn't drain your Codex usage. To use it, just head to key.ai and under settings you'll find your API key there, which you can just send to your agent to start using the different AI models that they have connected, images and videos and even music.

Number 18, compact with a focus. So, when your chat gets too long, compacting basically squashes that long conversation into a shorter summary. And normally your agent would guess what to keep in that summary, but what you can do is to just ask your agent to do compaction better. And you do this by having it interview you on what you would like to keep. That way, the summary holds on to what you care about, and every message after that would carry the context of your conversation, which would save you a lot from having to repeat yourself, which is not only frustrating, but it also costs you tokens.

Number 19, ponytail. So, this one was made for coding, but it can apply to other use cases as well. At its core, developer, which basically means that it writes the least amount of code that does the job. Less code written means fewer tokens and faster replies. And you usually end up with a smaller code base that's easier to look after. It already has thousands of stars on GitHub, and to use it, as usual, you just send the link to your agent.

Now, it's time for some advanced power user tips. And if you get lost in this part, remember that you can just send that PDF guide to your agent and ask it to explain any of these tools simply, so that you can apply them to your setup. Number 20, RTK. So, whenever your agent runs a command on your computer, it usually has to read everything that comes back from that command. And that output that it reads behind the scenes, it can get long.

So, RTK, which stands for Rust Token Killer, is a free tool that trims down that output before your agent reads it. It already has 80,000 stars on GitHub, and from their tests, it uses 60 to 90% fewer tokens on common developer commands. So, to try it, what you can do is to just paste the GitHub link into your agent. 21, Headroom. Headroom is an open-source compression layer that sits between your agent and the model, and it shrinks everything that your agent reads.

So, logs, files, and even tool results, those get compressed before it gets sent. If you're curious how this is different from RTK, RTK basically only trims what comes back from terminal commands, but headroom is a lot more broad because it shrinks almost everything that your agent reads. Headroom says coding agents use about 20% fewer tokens with it. And the great thing about it is of course you can adjust how much the agent would use it or not.

Number 22, QMD. So normally when your agent searches through your files, it looks for the exact words, sort of like pressing control F to find specific matches for files. But if you wrote the file name a bit differently, what ends up happening is that the agent would read file after file, which can be time and token consuming. Now QMD is a free tool from the founder of Shopify that searches by meaning instead of exact words.

So that if you have a larger workspace with thousands of files already in that second brain, you'll be able to find files faster and easier. So to use it, you can just paste the GitHub link in and say install this and your agent can set up QMD for you. Number 23, give it a token budget. Astra and other models in Codex actually have a view of the remaining usage in your plan. So a good trick is to put a token budget right in your prompt.

For example, you can add to keep this task under 3% of my weekly usage and to stop and check with me if it's going over. That way a big task gets planned around that number from the start instead of you finding out when the limit hits. 24, optimize your repeated workflows. For anything that you want to run over and over like drafting regular weekly reports. If you're finding that routine is draining your usage, what you can do is to point Astra at that workflow and to say to look at this routine and get it down to for example 1% of my weekly usage.

It does cost a bit of usage up front to do this, but if it's a regular scheduled task that's important to you, optimizing it is very well worth it. 25, name your helper agents. If there are task types that you commonly ask Codex to do, like summarizing a long document, what you can do is to set up a helper agent with its own model. In Codex, that's just a small file in the agents folder where you give it a name and set the model, and then you can call on this helper agent from any chat.

You can set this up with just one prompt to Codex, and that way if you set up a summarizer helper agent, it would run on the cheap model without you having to use a powerful model like Astra for menial tasks. 26. The brain and hands technique. So, the principle here is simple. If you have a big task, you can assign your strongest model, which in this case is Astra, as the brain. So, it only plans the work and checks it, and hands the actual doing to cheaper models, which are the hands.

To do this, you can basically say to Astra to plan this task, to split it into subtasks, and to send each one of those to a helper on the cheaper model, then just have Astra review what comes back. That way your top model's usage goes on the thinking, which is the part that really needs it. So, that's all of them. As usual, thanks for watching till the end. And also, let me know if that was useful, and what you think about this type of video.

You probably have guessed that this entire video was edited by AI, so I'm keen to know what you think of it, and I'll see you all next time. >> [music] >> Cheers.

The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.

Use this transcript

Three free tools that work on the material around a video like this one. No signup, no login.