Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Simon Scrapes · @simonscrapes
Words
4,417
Runtime
20:12
Speaking pace
219wpm
Reading time
18min
219 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
You've probably got a prompt that starts something like this. You're a senior copywriter with 20 years experience who writes all of my marketing copy. >> [music] >> Take a deep breath. This task is very important. I've written openings like that hundreds of times on my prompts, and someone actually went and tested this. If you give Claude a role or a persona, does it actually perform better? They tested 162 different personas across 2 and 1/2 thousand prompts, and the prompts with a persona did no better than just asking the question directly. So, all of those additional lines that you've been writing at the top of every
110 words, the words spoken in the first 30 seconds at 219 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 224 |
| Average words per sentence | 19.7 |
| Longest sentence | 78 words |
| Questions asked | 13 |
| Sentences containing a number | 42 |
Most used terms
Filler phrases
118 in total: actually 62 · basically 27 · like 18 · right? 5 · kind of 2 · you know 2 · literally 1 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
You've probably got a prompt that starts something like this. You're a senior copywriter with 20 years experience who writes all of my marketing copy. >> [music] >> Take a deep breath. This task is very important. I've written openings like that hundreds of times on my prompts, and someone actually went and tested this. If you give Claude a role or a persona, does it actually perform better? They tested 162 different personas across 2 and 1/2 thousand prompts, and the prompts with a persona did no better than just asking the question directly.
So, all of those additional lines that you've been writing at the top of every prompt for the last year, the model isn't even reading. And the Claude code team basically confirmed it from the inside because recently they've cut their own system prompt by around 80% and now say even that stuffing it with examples is no longer best practice. So, that's mistake number one of 19 that I'm going to show you today. Because pretty much every rule you've learned about using Claude has been rewritten in the last 6 months, and most people are still running on the old rules.
So, to finish number one off, you just delete those lines and spend the same words on the three most important things have to say in your prompt instead. And these come straight out of Anthropic's own prompting guidance. So, number one is telling it where to look. So, actually we're just going to give it advice on where the additional context it can go and get is. Number two is giving it that definition of done or the done criteria.
What does it look like when the output is finished? What should the output look like? And then third is giving it a self-check. And their docs literally suggest the phrasing, "Before you finish, verify your answer against." And then we add whatever our check should be. So, say you're getting a client proposal written. Instead of you are a senior copywriter with 20 years experience who writes all my marketing copy, we're going to get rid of that.
So, you write all my marketing copy. We're going to get rid of take a deep breath. This is very important. And what we're going to say is, "Actually, you need to look in the onboarding folder for the client notes. This looks like a one-page proposal covering their three pain points with price listed at the bottom when it's done. And before you finish, this is the most important thing to check against. Every number against our notes and flag anything that you can back up." So, this is roughly the same number of words as we had before, but now every single one of them is going to be used by Claude, which brings us to number two, which is one more from the Claude Go team around your prompts, which is stop writing do not do X, like do not return this as markdown text, because it basically confuses the model when it conflicts with what you're actually asking it to do.
And instead of writing something like do not return it as markdown text, we actually want to just say write it as smooth flowing text paragraphs. So, we're telling it what to do, not what not to do. Now, number three is about routines running with full tool access by default, which is not only unsafe because Claude can basically access all of your tools through every routine, but it's also more costly because we're loading in the contacts and the names, etc. of those tools.
So, you can basically come into the desktop app or your Claude account, go into the code, and go into routines, and if you look when you set up a new routine, all of the connectors are added by default. So, when you set up a new routine, you can go through and basically remove any of those connectors before you start. You can also retrospectively go into your routines and remove access to any of those connectors, basically redacting tool access from those.
Now, staying with connectors for number four, if you've got a decent number of connectors switched on, and your tool access is set to auto, which it is by default usually, the tool definitions for all of those connectors are actually being loaded into the context up front on every single message. Now, some people have measured this, and it's costing them thousands in tokens for every single session they have. So, there's a quick fix, just go into the chat mode, go to the plus icon, go to connectors, and make sure that inside tool access it loads tools when needed, not tools already loaded.
And now nothing is going to load until Claude actually needs to use that tool. So, go and do that one, and that's going to save you a bit of money in the long run, too. But, in contrast to that, with number five, there used to be advice about removing any MCP servers that you use infrequently because all of MCP schemas are loaded in on every single session. And I've definitely given a version of that advice myself in previous videos, but that's actually out of date now.
So, tool search has been on by default for a while, which means only the tool names actually load in from those MCPs, which is around 120 tokens, I'd say. And then the actual schemas get pulled in now on demand. So, adding more MCP servers now has minimal impact on your context window. You can actually check this by running {slash} context and see exactly what these MCP servers are actually costing you in terms of number of tokens.
And you can see next to that now, {slash} MCP, which are all loaded on demand, so it's costing us nothing from the outset. Now, number six, you'll have felt the pain of if you work between the terminal or something like VS Code on Claude Code directly, and then working inside Claude Code Work, for example, in the desktop app. So, if you've ever built a skill in the terminal, they get nested under {dot} Claude {slash} skills and then the skill itself.
So, let's say this trending research skill, which researches data from the last 30 days across Reddit, X, the web, et cetera. Those are accessible directly by Claude Code from this folder tree. But as soon as we move to Claude Code Work, the skill access no longer has access to those skills, even if you give it access to that folder directly. What you have to do instead is go up to this customize window and add it in as a zip file or create the skill again inside this customize window.
And you can see we've had to duplicate here, basically our trending research skill, so that Code Work can actually use that skill, too. Now, hopefully it's not long before the functionality between code and code work merge much more closely, so that we can actually just have skills in one place and use them across different platforms there. Oh, and by the way, there's one more thing that actually caught me out here, which is the Code Work loads all of those skills at session start.
So, enabling a skill or adding a skill halfway through needs you to actually restart before you can use that skill. So, then number seven is some advice that I used to hear quite a lot. So, using sub-agents was deemed important to keep your main context clean. And that is still true, but it's probably the most repeated Claude tip of the last year. But what never really comes up is what it's actually going to cost you by having a lack of context injected into that sub-agent.
So, Anthropic's own docs say that agent teams use approximately seven times more tokens than standard sessions when teammates are running in that plan mode because basically each teammate maintains its own individual context window which has a standard set of context that's injected there and runs as a separate Claude instance. So, the way that I tend to think about it now is that sub-agents are really brilliant for looking at things, broad searches, parallel investigation where not a lot of context is required, but they're actually expensive for actually doing things. you were actually going to delegate specific tasks where a high amount of context is required, then actually I'll just keep that to the main context session because you're having to pass that context that's then un-cached in that separate window on the same model as the main window and all of that adds up over time and that's why it tends to cost you seven times as much when you're using sub-agents.
Plus, we've got built-in modes that delegate properly and efficiently like Ultra Code now anyway. Now, number eight is a quick tip not to get tricked by. So, there's actually a mode called fast and as it sounds it basically is 2.5 times quicker, it runs on Opus, and you get this little symbol here to dictate that you're on fast mode. Let's say send me a response and you'll notice there that actually switched off fast mode for me because I didn't have usage credits enabled.
So, that brings us to the first limitation which is actually this fast mode is based not on your subscription plan, but on your API credits. So, if you have usage credits enabled, then using something like fast mode is going to cost you a lot of money and it will automatically stay on by default until you write fast off and that's even when you switch sessions. So, just be careful of that one. And the second limitation then, and this could really cost you, is the first time you enable fast mode in a conversation, you pay the full fast mode for un-cached data.
So, if you've had a load of context in that conversation, you've cached it or Claude has cached it in the background, you are paying less every message that you send and you're paying that on your subscription plan. You switch it to fast mode and suddenly all of that cached data becomes context that's now un-cached. So, you're going to pay the full fast mode uncached input price for that entire existing context. So, flipping fast mode on in the middle of a conversation could be a really dangerous and costly move.
Now, let's move on to hidden context. But before we do that, I am scarily close to 100,000 subscribers now. So, if you're getting any value from this, then please subscribe below if you haven't already. So, this next one's really simple, which is don't edit your Claude.md mid-session. So, I'm mid-session chatting and I decide actually I just want to add a rule inside a new section here. Always return your outputs in bullet points.
Say I'm really pedantic about how I want my outputs to be structured. And you'll notice that if I start a brand new session, it will abide by the rules that we've added. But if I continue in my existing session, then it will not acknowledge those new rules, right? So, until you clear the existing session or start a new session entirely, it will ignore any edits to the Claude.md. And that's because straight from Anthropic's docs, the project root and user-level Claude.md files are read just once at the session start and held in memory.
So, restarting, clearing, or compacting all basically get that to re-index in the information. So, if you ever want to add a rule mid-session, you're going to need to compact or restart that. Now, number 10 might seem obvious, but not many people actually stick to this. And you've probably heard keep your Claude.md short loads of times now, but what you usually don't get is the actual diagnostic of the impact it has.
So, here it is directly from Anthropic. So, it's been tested and proven. If Claude keeps doing something you don't want despite having a rule against it in your Claude.md, the file is probably too long and the rule is getting lost. And in the past we've always been told, "Okay, you can stress the importance of a specific rule by, you know, capitalizing it or writing important before it." But the real fix is actually just make the Claude.md shorter in the first place.
And just for some scale, Claude code's own system prompt is around 50 instructions, and models are supposedly following somewhere in between 150 to 200 instructions reliably. And the instruction files that perform best sit around 300 to 350 words. And that's words, not instructions. So, if you look through yours, is yours less than 100 lines? Is it less than 200 lines? The shorter, the better. The more you can outsource different contexts to different files and reference that, so Claude can pull it in at a specific time when it needs it, the better the model will be able to perform.
And there is, of course, a quick hack that you can run, and you can run the {slash} doctor command. And it's basically a health check for the user's Claude code setup and fix issues. So, let's say, "Doctor, Doctor, propose trims to my Claude.md that won't impact performance." Run something like that, you're going to get a suggested list of what you could trim from your Claude.md to immediately cut it down to only the most important information, and therefore you'll get better performance over time.
And that comes back and gives us a really nice set of suggestions uh from seven sections inside our agent.md and our Claude.md that would save us, you know, 4,000, 5,000 tokens every single session. So, it's definitely worth doing. Now, related to that with number 11, what actually gets passed over to a sub-agent? Like, what does it actually know when you pass a task over? So, it basically gets its own system prompt, whatever you wrote in the task that you were giving it, and the full Claude.md hierarchy, right?
Well, kind of, most of the time anyway. So, it doesn't even get conversation history. It doesn't get output style. It doesn't get any of the auto memory that's injected from the main thread, and none of the files that have already been read into context already. And it turns out some agents actually get even less than that. So, the built-in explore and plan agents, so these are automatically delegated if you've not created your own agent, like a lot of the exploration or the planning agents that are delegated by Claude themselves, actually skip the receiving of Claude.md.
So, the important thing to know here is when you're passing instructions to a sub-agent, if you are using sub-agents, if there are certain rules in that Claude.md that those sub-agents need to abide by, then you need to actually pass those in the instructions of the sub-agent directly if it goes and spins out any exploring plan agents. So, unfortunately, there's no great fix for this. It's just to restate those critical constraints inside the prompts that you're giving the sub-agent directly.
And if you are doing a side task that genuinely needs the full context, then actually you can just fork the conversation instead, and therefore it will inherit the full context of that conversation. Now, onto number 12, in a previous video I actually gave you incorrect information or it's been since updated. So, in my hidden settings video, I said Claude code waits until your context window is around 95% full before it compacts, and that you can actually go and change that percentage and override that.
And that's completely changed now. So, let me just correct what actually has changed. So, the current documented behavior is that if you don't set an auto compact window limit, then Claude code is going to compact when the conversation reaches the model's context limit. So, it's no longer a percentage, and there are per model exceptions on top of that. So, cloud sessions compact as they approach and some models compact at the full boundary, let's say 200,000 tokens in the window.
So, now what you can do is run the auto compact setting to set how full the context gets before it auto summarizes. And you can leave this automatically, or you can say 50K tokens or 100K tokens, and actually you set the auto compact window now to 100K tokens. Now, as you go through the conversation, it's going to hit 100K tokens and start auto compacting. So, instead of percentage now, you can now set a fixed amount.
And this is important because quality degrades long before that context window is full. And while we're on context size, for number 13, historically everyone's been chasing these huge context windows, right? We were all so happy when we saw that we had models that were now able to actually take in a million tokens in that window. But Anthropic's own published benchmark from testing Opus 4.6 on a long context retrieval test, where it put eight hidden things inside the context and tried to confirm if the model could find them, it found that at 256,000 tokens, it could score 93%.
So, actually pretty remarkable you can get 93% accuracy at that high number of tokens. But then at 1 million tokens with the same model with the same test, it actually dropped to 76% accuracy. So, suddenly you've got a three out of four chance of actually retrieving that information, which in plain English is roughly one in four of those actually failing. So, one in four of your rules not actually being able to be retrieved.
So, a bigger context window isn't really more memory. It's basically then going to become harder to actually retrieve the right information. So, I definitely recommend using things like auto compact as well. And by the way, for number 14, compaction isn't all or nothing. I used to think it was basically all or nothing. But if you hit escape twice in an existing conversation, it puts you in rewind mode, which a few of you are probably familiar with.
So, you can actually jump back to a previous part of the conversation and start again that conversation from there. But as soon as you go into one of those, it actually gives you more options. It allows you to restore the code and the conversation, restore the conversation, restore the code. But it also allows you to summarize from that specific point in the conversation or summarize up to here. So, when we use something like summarize up to here, we can add additional context.
But what we're actually saying is auto compact everything before this point. And now we can continue the conversation from that point onwards. So, we can choose which points of the conversation we actually want to auto compact. Now, moving on from context, this is one thing that hardly anyone seems to know and will actually save you quite a lot of time with such a simple change. So, when you're in plan mode and you're saying something like plan some pricing page updates, 10% off for more users, Claude goes off and creates a plan for you.
You kind of have these different options, right, to respond. So, you have yes and switch to bypass permissions or yes and manually improve edits or you can basically add in additional information. But I've always missed this, control and G to edit directly in VS Code. So, if we hit control and G, then what it basically does is open up the plan inside the browser so that we can go and edit the plan directly. Before, what I'd always have to do is actually ask it to go and edit the plan to change one or two little things.
But, if you hit control and G, you can actually edit the plan directly, tell it you've updated the plan, and then it can go and actually act on that new plan. Saves you a lot of tokens back and forth trying to get it to rework that plan for some simple changes. Now, number 16 is actually crazy because it's all about cheaper models actually costing you more. So, it catches basically everyone out because it looks like the sensible thing to do.
When you're doing a simpler job, switch to a cheaper model, right? Cheaper models are usually faster as well, so it makes a lot of sense. And this is all about the way that the data is cached and uncached. So, if you're, for example, 100,000 tokens into a conversation with Opus, and you want to answer a question that's fairly easy to answer, it would actually be more expensive to switch to Haiku at that point than to have Opus answer because basically you'd bring it to build the prompt cache for Haiku with those 100,000 tokens.
And then if you switch back to Opus, you're also going to be rebuilding that cache with 100,000 tokens. So, every model has its own cache, basically. But, the moment you switch, it gets uncached and you have to pay for it all again. Now, moving on to number 17, this is a real power users tip. So, Anthropic actually have a power user guide, and they say something in it that I think most people skim straight past. And their exact words are, "The single most impactful tip in this guide is verification.
And if you only adopt one practice, make it verification because if there's no check, the Claude can run itself, and you're basically acting as the verification loop for Claude, which means every single error just sits there waiting for a human to notice the errors. So, there are four levels here, and they escalate in terms of how you can get Claude to verify things. And the simplest is just asking for the check in the exact same prompt.
Then you've got {forward slash} goal, where there's a separate evaluator that rechecks a condition after every single turn. Above that, you've got stop hooks, which physically block the turn from ending until your script passes that stop hook. And then at the very top level, the best verification is actually an adversarial review agent. Now, when you employ an adversarial review agent, it is as it sounds, going to find gaps.
So, that is actually a a problem if you just told it to go and find problems because it will basically keep going until it finds problems. So, you basically need to tell the adversarial reviewer agent to flag only what affects correctness or your stated requirements or your done condition and treat the rest as optional. Otherwise, it will continue to try and find problems. So, if there's one thing you need to understand is use a verification layer instead of using yourself as that verification layer and there's four different layers you can leverage that.
Now, the next one I've actually lost a ton of data to in the past, so make sure that you correct this today. You can basically run resume right to resume previous conversations. When you hit that, you'll see 13 sessions. I haven't only completed 13 sessions since I started using VS Code and Claude Code right. And that's because by default it only keeps 30 days of your previous conversations. And those 13 for me have been mostly today because now I've mostly switched using the desktop app and Claude Code there instead.
You can see if I scroll down, we've got a few from 1 month ago, but before that all of the conversation history has been completely wiped. So, inside your settings.json if you want to stop this, then you can actually put clean up period days and then a number. So, if you put 365 in there, then you've got a year of historic data. And by the way, zero doesn't mean unlimited, it'll basically wipe everything as it comes in.
It's worth knowing separately that your auto memory files are actually exempt from that sweep. So, anything stored in the auto memory files do get kept. It's just you can't resume those conversations. And then the final hack, the last one is basically five for the price of one. There are commands that hardly anyone ever types. So, we used doctor earlier, which is basically your full Claude setup check in or check up.
So, it'll find skills, MCB servers, and plugins that you're not using, show how much it's costing you in context, dedupe your Claude.md, all of those different things. It can propose trims, etc. And those are all really helpful. Then we've got insights, which will basically generate a report analyzing your Claude code sessions. So, it's a full HTML report analyzing up to 200 of your recent sessions. Now, this is only going to be valuable if you're not clearing sessions after the last 30 days, right?
You've then got by the way or BTW, where you can ask a quick side question without interrupting the main conversation. It will come back with an answer. Hey, what's the date today? And it will answer that. You see in the background of Y and we can close that with escape whilst it's still running that insights command. And then if you're mid-conversation and you've got a bunch of context that you don't want to lose, but you want to do something else with it, then actually you can run branch and it will create a branch of the current conversation at that point.
You can take that to a new conversation. So, it copies your conversation to take it in a different direction whilst the original conversation stays exactly the same. So, if there is a point that you want to branch out multiple conversations with the same context, you can do that with branch. So, that is the 19. Next video I'm going to show you 14 ways to supercharge your Claude code setup. Give this video a like, subscribe, and I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.