YouTube transcripts

Paste This Into GPT-6 Astra, Never Run Out Of Tokens Again: video thumbnail

Paste This Into GPT-6 Astra, Never Run Out Of Tokens Again transcript

Sharbel A. · @sharbelxyz

Published September 19, 202611:5970.8K views

Watch this video on YouTube

Transcript analysisComputed from the caption text

Words

2,040

Runtime

11:59

Speaking pace

170wpm

Reading time

9min

170 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.

Opening (first 30 seconds)

GPT-6 Astra might be the smartest model ever made, but it's also eating people alive. Most people are burning through their entire ChatGPT subscription in minutes on this thing. So, naturally, people are scared to even use it, and I get it. But, once you understand how the model works, you can cut most of your token consumption out. So, in this video, I'll show you exactly why Astra burns tokens so fast and the fixes that actually move the needle in order of what

85 words, the words spoken in the first 30 seconds at 170 words per minute.

Sentence shape

MeasureThis transcript
Sentences135
Average words per sentence15.1
Longest sentence38 words
Questions asked2
Sentences containing a number13

Most used terms

  • astra16
  • model13
  • message11
  • start11
  • tokens9
  • codex8
  • context8
  • single8
  • actually7
  • cost7
  • effort7
  • people7

Filler phrases

19 in total: actually 7 · like 7 · basically 3 · literally 1 · you know 1.

A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.

What this transcript is

Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.

Transcript

GPT-6 Astra might be the smartest model ever made, but it's also eating people alive. Most people are burning through their entire ChatGPT subscription in minutes on this thing. So, naturally, people are scared to even use it, and I get it. But, once you understand how the model works, you can cut most of your token consumption out. So, in this video, I'll show you exactly why Astra burns tokens so fast and the fixes that actually move the needle in order of what saves you the most money.

I am saving the nastiest one for last because there is a setting that's literally doubling your bill without you even knowing. But, let's get started. Real quick, because you need to know this before we get started. These models have no memory. None. So, every time you hit enter, your entire conversation gets packaged up and sent again from the top. So, your first message is cheap, but your 20th, you're repaying for everything above it every single time.

You're paying for every message you've sent before and every message it has sent you back as well. So, your consumption compounds. Astra makes this worse for one reason. It thinks more. All that extra reasoning is extra tokens you're paying for, even for the things that you never see on screen. That's essentially the tax for it being this smart. The good news is you control most of it. So, let's start cutting away. Before you change anything, you can start by finding out what's actually draining your tokens, because it might be different for different people.

In ChatGPT and Codex, you can either ask it something like, "What's my usage like right now, and what's eating it the fastest?" It will tell you how much of your limit you've burned through, when it resets, and which model and settings is draining you. Alternatively, and I really recommend you do this, you can take a screenshot of the prompt on the screen right now, and send it to your GPT session. Or, for the sake of convenience, that exact prompt and the other two prompts I use in this video are in the first link in the description, so you can just go there and copy them straight away.

Run that first prompt first, and you'll know exactly which of these next five fixes matter the most for you. And once you know your leak, here's the biggest lever, and almost everyone gets it wrong. Astra has a dial for how hard it thinks. There's light, medium, high, extra high, and ultra. And everyone slams it to max thinking that they'll get a better answer, but you usually don't. You just end up paying more. And these numbers aren't mine.

Artificial analysis benchmarks these models independently. They give the models the same set of tasks and measure the results versus their token consumption. On low effort, Astra costs about 80 cents a run, but if you crank it to the max, that same exact run or work costs $3.26. That is four times the money for the same exact job. And on their coding benchmark, max effort runs about $7 a task, just to tie a model that costs a fraction of that.

You are paying a premium for a score you can basically match for much cheaper. And this isn't just me. OpenAI's own guidance says lower effort doesn't mean lower capability, and that cranking the effort higher won't always give you a better result. So, I highly recommend you start on medium and not the default, which is high. Only if the answer you're looking for is not landing, then bump it up. Don't start at the top and then work your way down.

Instead, work your way up. And while you're in there, turn off fast mode unless you actually need the speed for some reason, because that one alone doubles the rate you pay. Now, the next one is where you save the most, and it's about not using Astra for everything. Here's the shift I want you to move towards next. Don't try and every single task on Astra. I know what you're thinking. I came for this video to know how to use Astra more and this guy is telling me use Astra less, but just hold on and bear with me.

Astra is as smart of a model as we get. I'm not saying don't use it. I'm saying be selective on where you absolutely must use it. Use it like the smart senior developer it is. Let it plan and review for you, and then give the grunt work to cheaper models. That way you get the best of both worlds. You get the thing you want Astra for on, and then you get the rest, the grunt work, the things that any model can do in the world done for a much cheaper cost.

This is how I'd set it up. The boring scoping and cleanup, push that down to the cheapest model. You can use a lighter Open AI model like GPT 5.5 for this. The planning and the checklist, you can hand to a mid-tier model like GPT 5.6 Soul or Luna. Then Astra only touches the hard part at the end, the actual judgment call on medium effort. Now, I wouldn't recommend you switch these by hand every single time. Instead, you can set it up once and let it route itself automatically.

Codex lets you build sub-agents, and there are two ways to set them up. The first is super technical, but you basically create a .toml file inside your project folder that looks like this. The second way is much easier. You just paste these three prompts into your Codex chat. Two of these prompts create sub-agents for you, and the third one tells your Codex agent when to call out each of those two sub agents. That way, you're paying for the genius once on the 10% that needs it instead of every single step.

This one change is the difference between a video that cost me 10 bucks and one that cost me one. And speaking of what you feed Astra, there are two things that quietly bloat every single message, your tools and your output. When you connect a tool, its whole instruction manual loads into your context before you even type anything. A big one can carry tens of thousands of tokens on every single message. So, you can click on settings, then plugins, then look through your list and turn off anything you haven't used in over a month.

People have taken a 55,000 token load to 3,000 with that one small simple change. Then, there's output. You tell it to install something, then you have 800 lines come back, but you don't see or read any of it, and now you're repaying for those 800 lines on every message after. There is a free tool called RTK. You point Codex at it, and instead of dumping raw command output into the chat, it hands back a stripped-down version of that.

I ran it on my own setup, a giant file listing went from 7 million characters down to about 1,000. Two things worth knowing if you install it. On Codex, you have to tell it to actually use RTK. It won't grab the output on its own, and it only helps on that big messy output, not the fixed cost every message already carries. There's also a desktop one called Headroom that does the same job across Claude Code and Codex, and its maker says it cuts cost by about half.

Now, let's move on to a few small habits that stack up fast. When you want to start a new conversation with Codex, don't just type in your new prompt in an existing conversation window. Start from a new context window altogether. Remember, you don't just pay for the message you send, you pay for every message you've sent before along with every response you've gotten before every single time you send a new message. So, when you're done with a workflow, always start fresh.

Start from a fresh conversation. Another very important one, pick your model and effort at the start of the session and leave them. Because the moment you switch models or effort levels mid-task or mid-run, none of your cached conversation matches anymore and it reprocesses the whole thing at full price. So, the thing people do to save money, switching to a cheaper model halfway through, is actually the most expensive move on the board.

And also, don't paste in whole repos or entire old chats. Keep your scope tight because that's exactly what can bloat a context window. All right. Now, here's the one tip I promised you at the start and this one will shock you. This is the one almost nobody talks about. Astra's got this huge million token context window they love bragging about, but here's what they don't put on that same slide. The second your input goes over about 272,000 tokens, the price doubles.

Input goes to two times its price and output to one and a half times its price. And it's not just the tokens over the line, it's the entire request. All of it gets repriced. So, that giant context window is a bit of a trap because if you fill it up past 272,000 tokens, you walk straight into doubling your bill with zero warning. That's why I'd recommend keeping an eye on your working context and trying to keep it well under that line.

If you want to keep an eye on your context window, you can type {slash} status and always be able to look at how much context you have left and how much you've spent inside one session. And before you go, let me kill the fixes that don't actually work because a couple of them are the opposite of what you've been told. Number one, writing shorter prompts. That doesn't change much, to be honest. What you type is essentially a rounding error.

It's basically nothing. So, don't bother typing shorter prompts. Number two, compacting to save money. That one's a little backwards because to summarize your chat, it has to send the whole thing one more time. So, the thing you do to save tokens is the most expensive message of that entire session. Only use {slash} compact for continuity if you absolutely need to continue a chat, not for savings. Number three, screenshotting text instead of pasting it because a screenshot can cost you a couple thousand tokens.

The words in it cost way less. So, paste the text instead. And lastly, PDFs. You pay for every page of a PDF twice. Once for the model to read the words and once for the picture of the page that the model takes. Convert it to a plain text file first and then send the doc instead because that same doc costs you only a quarter of what it would have as a PDF. So, look, Astra isn't the problem. It's the smartest thing out there and it's genuinely worth using.

It just doesn't come with a manual, but if you do these six things, at least it'll stop rinsing you and you actually get to use the good stuff without watching the meter every single second or minute. The setting that will save most people the most money is making Astra the boss, not the worker. So, start there if you really want to feel a difference. That being said, tell me in the comments which one of these were you doing wrong.

I'm genuinely curious. And as I mentioned, every prompt I use throughout this video, the all the one, the sub agent setup, all of it is free in the first link in the description. And if this saved you some money, then subscribe because I've got a ton more content like this coming your way. Oh, and would you look at that? The algorithm gods have decided you're going to really enjoy this video next, so click it and I'll see you there.

The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.

Use this transcript

Three free tools that work on the material around a video like this one. No signup, no login.