Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
16:124.1x the video's typical replay level
and I totally agree. I can't wait till we get to patch MD though. That's a plan up and cooking for a while that I've not had time to do. And while prompts decay silently, even janky code will tend to be relatively stable if it's left untouched. But, every model upgrade can turn a functional prompt into something
Said at 16:04
Most replayed moment #2
3:173.6x the video's typical replay level
swa.link/archchat. Let's dive into prompts our technical debt, too. It's common and correct to say that all code is technical debt. Yep, all code is a problem that you have to maintain. I agree. Adding code is a necessary evil for developing new features. You almost
Said at 3:10
Most replayed moment #3
13:073.5x the video's typical replay level
but it also starts in a very minimal place. Unlike GitHub, which has been loading for the entire time I'm talking about this. Great work, Microsoft. Pi starts out with under 1,000 tokens of context. That's crazy. Some harnesses like Claude code start with like 10,000. With all the tools and features built
Said at 13:00
The graph counts replays. It does not show where viewers stopped watching.
Words
4,442
Runtime
20:29
Speaking pace
217wpm
Reading time
19min
217 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Technical debt's always been a massive problem for our industry. It's one that I've encountered more times than not across all of the different roles I've had, even building my own stuff. It's so easy for technical craft to build up. Thankfully, AI's here to save us, right? Well, sure. For many places, the technical debt is building up more as more people who don't know how to dev write are just throwing in prompts and begging to have their stuff merged. But on the other hand, I've actually had a lot of luck using AI to trim technical debt from projects where it wouldn't have been worth it
109 words, the words spoken in the first 30 seconds at 217 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 268 |
| Average words per sentence | 16.6 |
| Longest sentence | 98 words |
| Questions asked | 8 |
| Sentences containing a number | 24 |
Most used terms
Filler phrases
54 in total: like 42 · actually 6 · kind of 3 · basically 1 · literally 1 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Technical debt's always been a massive problem for our industry. It's one that I've encountered more times than not across all of the different roles I've had, even building my own stuff. It's so easy for technical craft to build up. Thankfully, AI's here to save us, right? Well, sure. For many places, the technical debt is building up more as more people who don't know how to dev write are just throwing in prompts and begging to have their stuff merged.
But on the other hand, I've actually had a lot of luck using AI to trim technical debt from projects where it wouldn't have been worth it before. I've talked about this a bunch randomly throughout other videos, but I want to go a bit deeper here. Not just on technical debt, but a new type, prompt technical debt. Crazy enough, in many ways, prompts can be technical debt, too. Shawn Gottesdiener already written multiple awesome articles that we've covered in these videos, and a feeling I'm going to love this one even more.
Technical debt's one of those things that hopefully, fingers crossed, will be kept around to fix. It might not be the most fun work. We might be going from builders to janitors, but it is absolutely a thing we have to be considerate of. And I'm excited to see how Shawn is thinking about it, in particular around the prompts themselves and the debt they represent. I already am seeing things in here I'm going to like, like the agent MD as debt.
I know for a fact that the agent MD in the T3 code repo is, at best, out of date, at worst, causing problems. Before we can get there, I am hoping to retire before our jobs become just slop cleaning up stuff. And at the very least, to pay my team. So, we're going to take a quick break for today's sponsor. I need to be real with y'all. I'm getting scared about the future of safety and security in software. I've talked about this a bunch in other videos, but I've never really talked about how to get the security right.
That's because it's different for everyone. But if I know one thing is true, you need to have the sources in your code. If it's not in your code base, you can't trust it. And that's why I'm really hyped for today's sponsor, Archjet. These guys have the primitives to make more secure software inside of your code directly. The easiest way to set up is to copy their prompt and paste it in your code base, then it will figure it out from there.
But if you look at the the you'll be just as impressed because it's really good. It was so good I ended up investing in the company because I was genuinely that hyped on what they're doing. They provide a set of packages that handle all of the identification systems that you need to know who a user is and if they're real or not. They have a rule section for setting things up like bot detection and you can even allow certain types of bots through if you want.
They have token buckets which is super useful for tracking how much a user's done a thing. You want to allow users to send five messages or maybe they can go to 20 pages before you start rate limiting them. Don't put that in your database. Put that on a platform built for it. You're not doing this by white-listing at the front before they hit your code. You're doing this in the actual route file in the next JS project here.
I love this example because you still just do the things you need to for all users. These things are free though. You're just concatenating some strings. But once you actually want to decide if we should do the call or not, you call aj.protect with the request and the other data you care about and then you get back a decision whether or not you should actually process this. And the decision has everything from the reason it happened which makes it easy to know, oh, is this a bot or are they out of messages or are they getting limited because of sensitive info or some other thing.
You return if they were denied, otherwise you let them go. In a world where everything's getting pwned, you should be confident in your code and your security. Get those things right at swa.link/archchat. Let's dive into prompts our technical debt, too. It's common and correct to say that all code is technical debt. Yep, all code is a problem that you have to maintain. I agree. Adding code is a necessary evil for developing new features.
You almost always have to do it, but each line of code adds to the complexity and maintenance burden of the system. All future changes to the system have to work with the existing code or at least avoid breaking it. Once systems accumulate enough code, they become impossible for a single person to understand. Instead of reading the code and understanding what it does, you must rely on guesses, theories, and heuristics.
Sensible engineers write as little code as possible. Yep, I have said this many times. I was really proud at Twitch that I think I ended with more code deleted than added and I was really trying for that. I love deleting code. It's my favorite thing. It's also funny that we're talking here about how no one person can keep all of the code in their head. I still remember the era where we thought AI would be able to do this.
I'm going to make a call out that feels a little bad, but hear me out. Surely you guys know of Michael Truel, the CEO of Cursor. He's an incredible dude, very smart, runs this company great. He's a huge part of why Cursor's been so successful. He didn't have much industry experience before starting Cursor. I'd argue he had effectively none alongside his co-founders as well. They didn't have much experience building things at scale.
I still vividly remember a conversation that we had had as well as some interviews he has done about this point because he seemed to think not only were there engineers who understood the whole codebase, but there was also the possibility of AI doing it, too. He firmly believed context windows would get bigger and bigger and the role of Cursor would be to have the best system to load the whole codebase into context. Then Claude Code came up and was like, "Yo, what if I just grep for shit?" And did a way better job.
We even thought AI was going to do this better and then realized it isn't. And now AI works kind of the same way we do. It has short-term memory loss and it needs help finding things in your codebase. It's gotten good at finding those things through an insane amount of compute being used at our album, but neither humans nor AI as far as we understand, can keep track of a million-line codebase confidently. They need to understand how the parts go together, but not the details of every line of code across that type of massive codebase.
And I find this is one of the hardest things for earlier career devs to make the jump for and even for startup founders that are quite successful. They're not used to the idea that the codebase is bigger than their own understanding. And even the most successful people have been guilty of this some amount. Back to the article, specifically the sensible engineers write as little code as possible. This is the thing I worked really hard to do and I was very proud of.
Many large projects now have a set of code-based specific prompt files, agents.md, cloud.md, the same files but in subdirectories, as well as skills. If you're building a program that uses AI, you have separate prompts for capabilities and for each tool, as well as a whole set of system prompts. Even better, we have dynamically constructed system prompts in tools like T3 chat where the system prompt is adjusted based on certain parameters you select based on what tools you turn on and off, things like search, as well as which model and model provider you're using because certain models, cough, Gemini, cough cough, need a lot of help not being [ __ ] So, we have some very complex, annoyingly so, code that is effectively just a string concatenation system to generate the right system prompt to steer the model roughly where we want it to go.
Obnoxious, but like borderline necessary. And now that I'm thinking about it, there's a lot of technical debt there. Prompts are important. Minor tweaks to an LLM's prompt can unlock significant performance improvements. If the same model feels different across Codex, Cursor, Open Code, and Co-pilot, it's almost certainly due to subtle differences in prompting. I know this is one of the things that Cursor's put a lot of time into.
The Cursor harness meaningfully improves the performance of something like Opus when compared to using it inside of Claude Code. It some benches measure it as high as like a 10 to 30% quality of performance improvement when you use Opus in Cursor instead. A lot of that is just the system prompt. And I know that Cursor does crazy things like AB testing system prompt changes, crazy benchmarks they build internally where they slightly adjust the prompt for different models to see how it behaves.
I know, for example, when Gemini 3 Pro dropped, using that inside of official Google stuff sucked and Cursor was weirdly good with it. And I learned that one engineer had stayed up really late the night before launch testing everything he could to try and force Gemini 3 Pro to behave. And I still remember the day I tried it, seeing it in the thinking traces, where the model had like five steps of thinking, where it was just talking itself out of using tools unnecessarily.
That was because they had to add a specific blurb to un-Gemini-ify the Gemini models. On one hand, system prompts aren't this like thing we have to carefully protect like we pretended they were not long ago. I can't tell you how many reports we've gotten from users of T3 chat that like, "Look, I stole your system prompt by asking this model this question." Cool, I don't [ __ ] care. Maybe give me some feedback so we can do a better job at it.
There is still a lot of work in the engineering of these system prompts though, because the difference between two can make a meaningful difference in the performance that you get. The author calls out that AI companies do spend a lot of time testing and tweaking their prompts. So, it makes sense why engineers would spend a lot of time tweaking their agent MD files as well. I'd even call switching tools or workflows to be a form of prompting.
Yeah. Which tool you're using definitely affects these things. If I start wrapping my agents in a route loop, put in a new skill file, or install an MCP server, that's still a change to my prompts even though I'm not the one who wrote it. Yep. And this is something I find people don't understand when they're using tools like Codex. In Codex, they have a plugin system where you can go and install plugins. When you've installed a plugin, that plugin is now one of the things in the system prompt that the model knows about.
I can even ask to prove it. What plugins and skills do you have access to? Here are all of the plugins that I have installed and it knows about. It knows about them because they're all in the system prompt. It also has this a big pile of available skills. The front-end design skill still finds its way in even though I have deleted it many different times. I don't know what is causing it to keep reappearing, but it is.
It's actually a good audit cuz a lot of these things I don't want. So, I'm going to go clean these all up later, actually. Yeah. This is a real problem to consider, and many models will just use tools cuz they're there, even if you don't want them to. MCP servers in particular are really guilty of this. I've seen people who just went and blindly clicked install on every MCP that sounded cool, and now whenever they start a new agent, half the context is already taken up, and they're really confused and upset, and think AI is not very smart.
Look, I installed everything, and the AI is still dumb. Good luck. It just doesn't work that way. Sean calls out that he thinks it's a bad idea to spend a ton of time tweaking a bespoke agent coding setup. Interesting. Shots fired at our friend Ben in his crazy pie setup that he keeps iterating on constantly. I fall in this camp. I like using things as close to stock as possible, and I'm just thinking about where they're even running.
So, why does he feel this way, given that prompt adjustments can deliver so much value? That's because prompt adjustments are model specific. Earlier, he said that AI companies spend a lot of time tweaking their prompts. In fact, they spend that amount of time for each new model release. I experienced this personally. There were bad things happening when I used Codex with 54 during the early testing that were largely because they hadn't adjusted the tool descriptions and system prompts properly enough for the new model, which resulted in it doing things like searching way more than it should have been, and polluting its context with nonsense once it found bad results.
They adjusted the descriptions to fix this. And as crazy as it sounds that a prompt that worked great for 54 won't necessarily work as well for 55, absolutely is true, and I've experienced this a ton. That's kind of why I thought 54 to 55 should have been a different name or number to indicate how big a jump it was. Not cuz it's so much better necessarily, but it's so different. Like 53 to 54, I didn't have to change how I prompt much.
From 54 to 55, I've had to rethink how I use the model entirely. I like the phrasing that you have to learn how to hold the model every time. I absolutely agree. Even between 55 low and X high, it feels entirely different. In other words, a set of prompts that you carefully crafted in January this year might be out of date or actively harmful by February. Worse still, you might not even notice. Model capabilities are already so hard to pin down unless you're running every problem through various different models and tools.
And even weak AI systems are surprisingly good at some problems. You might just think, "Huh, the new Anthropic model isn't as impressive as the hype." or "Wow, Claude code got worse recently." Yeah, absolutely this. I've experienced this too. And the labs have experienced it as well, like I just said. Like this new model in all of their measurements was killing and in their custom tools internally was killing. When I tried it in something a little older customized to my uses, it bombed.
Funny enough, this is actually one of the reasons I really like Pi because I didn't customize it. I use Pi in this boring a way as possible because I want to see how the models work with nothing provided. Given the most minimal possible setup, what does it do well and poorly? And then I add things to solve specific problems I have when I have them. And I think this minimalist way of building, the the Unix philosophy so to speak, is one of the best ways to avoid these types of problems.
When you get a new model, start it on the smallest possible tool with nothing included and then add things as you need them. I've learned that whenever I mention Pi, people don't know what I'm talking about half the time, so I will point you at it. It's the Pi project. It's now hosted on Arendelle Works on GitHub. It's by Bad Logic aka Mario. It is a set of tools, but primarily a coding agent. The magic of the Pi CLI is that it is super customizable and you can tell it to add features and make changes and it will, but it also starts in a very minimal place.
Unlike GitHub, which has been loading for the entire time I'm talking about this. Great work, Microsoft. Pi starts out with under 1,000 tokens of context. That's crazy. Some harnesses like Claude code start with like 10,000. With all the tools and features built into Claude code, the system prompt is now 65,000 tokens. With everything to say, "Well, it's only 12K." but that's insane. That's before you even start doing things.
Pi is really cool that it lets you build and customize it these ways is great. As I said before, I like it for not customizing it because it stays insanely minimal. The sinister nature of this debt is what makes it so scary though because tech debt's pretty apparent whenever you go to add a new feature or make a change to your code base, you feel the tech debt every second. But the subtlety of these regressions is much more painful.
In this sense, prompts are a worse form of technical debt than code. When technical debt blows up, it usually causes errors or a tangible slowdown as you try to understand the code. Prompts will decay silently. Yeah, very big deal. Also, even janky code tends to be relatively stable when untouched, but every single model upgrade could turn a functional prompt into a non-functional one. Just for example, let's look at the T3 code repo which has an agent MD that has not been updated for 2 months.
It still says, "This repo is a very early whip. Proposing sweeping changes that improve long-term maintainability is encouraged." This line probably causes models to do things it shouldn't. We also call it the T3 code is currently Codex first. That hasn't been the case for a while. This is an outdated file that needs to be updated. The issue is that we haven't had much reason to touch it because we're busy shipping features and changing this correctly would ideally involve us testing the changes to see if the model behaves better or not.
I did get a question from chat which is, "Isn't the sweeping changes line an intentional lie?" It kind of was initially. I leave things like this in often specifically to try and get the models to be willing to push back more and suggest big overhauls that you might not expect a model to suggest. This framing of it is not necessary anymore and is quite possibly damaging. The other project I hinted at earlier that I'm working on is Lakebed.
It's a new way of doing full-stack app development and the agent MD for this project is less a traditional agent MD where I tell it where things are and more an essay, like a letter to the model about how I want it to behave and how to think about the project, so that doesn't need as much context on what it is every time. It just understands. But, this is different from the read me cuz the read me is for users, not for agents.
Just like the way you have to think about these things is crazy, and it's it's a different type of engineering. It is still engineering in the sense that you're like trying to find the right way to get these pieces together and to get this technology to behave, but it's different in the sense that it is nowhere near as deterministic, and the failures could be very quiet and hard to notice. I also love this call out from Joel here that so MD is the best thing that came out of open claw.
Absolutely agree. The idea of like describing the why and how and not just the what and where to the model is a very good idea, and I totally agree. I can't wait till we get to patch MD though. That's a plan up and cooking for a while that I've not had time to do. And while prompts decay silently, even janky code will tend to be relatively stable if it's left untouched. But, every model upgrade can turn a functional prompt into something entirely non-functional.
Could you simply decide to not upgrade models? Some people are trying this, but the pace of improvement is fast enough that it isn't really practical. A delicately prompted agentic harness built around GBD41 is always going to underperform a bare-bones harness built around 47. This might seem like a sensible strategy at some point in the future when the rate of model improvement slows down. At this point, I don't believe that'll ever happen.
I've been proven wrong enough times. There's also the point that it could just be that models are so capable that we don't need the extra intelligence for normal end tasks, but I don't believe it's a good strategy today. I agree with the author here. In Shawn's view, most people should just be picking an AI coding tool maintained by a third-party company like Claude Code, Codex, Cursor, Co-pilot, T3 Code, etc., and leave it as unconfigured as possible so they can piggyback on the work of the teams of engineers who are evaluating and tweaking prompts with each new model.
This is the same reason I built T3 Code the way I did. Rather than make our own harness, our own system prompts, our own everything, we wanted to take advantage of the fact that companies like Anthropic and OpenAI and Cursor are putting a lot of time into massaging their tools to get them right, and we just wanted a better UI on top. This also means you should do things like avoid unnecessary MCPs and skills unless absolutely necessary and keep them off by default.
At least this way, if one of the teams gets something badly wrong, users will notice eventually and start complaining. Yep. How many times has Claude code regressed cuz they put something stupid in the system prompt? It's hilarious how often that happens. I've demonstrated that in many videos. When you write agent.md files, try to avoid behavior steering like the now outdated think step-by-step, you're a skilled engineer, or if you get a task right, I'll tip you $200.
Don't forget make no mistakes. Keep them limited to specific concrete facts about the project. Don't let models fill your agent.md with pages of barely reviewed text for the same reason that you wouldn't let them fill your code base with pages of barely reviewed code. Yeah. And I would also agree like even more so, bloating your agent.md is bad because it makes all future code bad and it makes it harder to clean things up.
He ends with a great sentence here. Write your prompts yourself and delete them whenever you get a chance. Absolutely agree. If you're using AI to generate all of your prompts, if you're using AI to generate markdown files that you use as prompts that are sitting in your code base hoping that they provide value, if you use like the Claude code /init command which I talked a lot of [ __ ] on on my agent.md and Claude.md video, you are making your models dumber.
And if you leave that in when a model upgrade happens, you're leaving something stupid from a previous generation holding back the new one. I see a comment from chat here, which is I just want the stupid thing to read my mind. The less I have to explain myself the better. This is true vibe coding. Well, I have good news. That is basically what was proposed here. Most people should just be picking an AI coding tool maintained by a third-party company and leave it as unconfigured as possible so they could piggyback off the work of the engineers who are evaluating and tweaking prompts with each new model.
Do that. That's what I do. I spend as little time as possible customizing these things. I have a like three-line global agent MD that is just like the tech I prefer building with because I don't want to tell the model every time I'm admitting a new project what I like. And even that gets in the way sometimes and I'm considering getting rid of it. Most of my projects don't have an agent MD or Claude MD for quite a while in unless I'm doing something big enough and bold enough that I'm tired of steering it and I just want to keep it going in a specific direction.
Generally speaking, you should just be prompting. And if it doesn't do what you want, you should prompt it to do what you want. And you'll start to build the muscle across all these different models, across all these different tools on what this tool needs to be told to do the thing you want it to do. And if you take anything from this, you should go do an audit of your markdown files that are being used for things like this.
Do an audit of your system prompts and the other things that get passed to agents to do work in your code bases and see if you made the mistake we made here where your agent MD hasn't been touched for literally 2 months in a code base that has been entirely rewritten multiple times in that same 2-month window. If you're watching a video like this one, you probably already have meaningful technical debt across your prompts.
It's time to go deal with it. In efforts of reducing context blow, I'm going to kill it now. Thank you as always. Peace nerds.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.