Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Theo - t3․gg · @t3dotgg
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Theo - t3․gg's most watched videos.
Most replayed moment at 2:36
5.6x that video's typical replay level
Well, you have nothing to worry about cuz the first million users are free. Get yourself enterprise ready at soidiv.link/workos. I'm very excited to read into what Linear is cooking here. You know they're cooking something different cuz this is the only not dark mode page I've ever seen Linear ship. They actually took
Said at 2:29
Most replayed moment at 25:09
3.4x that video's typical replay level
just set up now. It's really nice when you do need to do things in the GUI or format the machine, that type of thing. It's time to show you guys how I actually do work using this setup. NPX T3 at nightly serve. I do have these host commands to make it work better
Said at 25:02
Most replayed moment at 2:53
29.0x that video's typical replay level
of different places in order to pull it together. Setup couldn't be easier. You click start, you add a new database, you get a connection string and now you're good to go with a real Postgress database with all the power of ClickHouse behind it. Stop compromising on your database today at soy.link/clickhouse. The best
Said at 2:45
The graph counts replays. It does not show where viewers stopped watching.
Words
8,199
Runtime
39:28
Speaking pace
208wpm
Reading time
34min
208 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
For AI to work effectively in a codebase, it needs to know what the codebase is, where things are, and how to get stuff done. I have found that a lot of people want to store this information in some special way that AI agents will operate with. Well, I've seen so many different systems and libraries and plugins and features and stuff that people have built over the last few years to try and automatically encode what they're doing in the codebase in some special way that'll make the agent always do what they want. I've even noticed that in claude code recently,
104 words, the words spoken in the first 30 seconds at 208 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 481 |
| Average words per sentence | 17.0 |
| Longest sentence | 124 words |
| Questions asked | 29 |
| Sentences containing a number | 48 |
Most used terms
Filler phrases
119 in total: like 73 · actually 14 · basically 7 · right? 6 · kind of 5 · literally 4 · uh 4 · I mean 2 · sort of 2 · um 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
For AI to work effectively in a codebase, it needs to know what the codebase is, where things are, and how to get stuff done. I have found that a lot of people want to store this information in some special way that AI agents will operate with. Well, I've seen so many different systems and libraries and plugins and features and stuff that people have built over the last few years to try and automatically encode what they're doing in the codebase in some special way that'll make the agent always do what they want.
I've even noticed that in claude code recently, it seems to use its memory a whole bunch where it will save things in some magic hidden file somewhere on the computer after you ask it to do something. And now it will keep randomly doing that forever. If you can't tell from how I've been spinning this so far, I'm not a fan of these memory systems. I don't think memory is the right place to keep track of how things should be done in your codebase.
That worked okay with humans because previously before AI was writing all our code, the knowledge was always in the people's brains. They would try to put it other places like documentation and whatnot, but in the end, what we knew is what mattered. Now we're working with agents that forget everything when a new thread spins up. And our desperate attempts to have them automatically save things to be in the next run is just not great.
It's not great at all. And I'm far from the only person who feels this way. This clip comes from a conversation that Mario, the creator of Pi, had with Armen, the creator of Flask. They now work together building a bunch of cool AI tooling stuff. And obviously, Pi is a super cool, powerful, useful piece of the new agentic engineering world. I love both of these guys. I especially love Mario. He's been awesome to chat with, and his understanding these things is great.
And I have a feeling he's going to have some real fun spicy takes as we dive in. So, if you're trying to figure out how to make agents behave properly in your codebase or you just want to understand this stuff well enough to have better conversations with your co-workers and keep them from destroying your agentic systems with all these terrible memory things, I think you'll get a lot out of this video. But one other thing you should be able to get a lot out of is today's sponsor.
Do you know what OpenAI, Anthropic, Cursor, and Thinking Machines all have in common? Because it's not just that they're trying to make models, it's that they all use today's sponsor, Work OS. You might be confused why so many big companies are choosing to use a product like work OS instead of rolling their own O, especially nowadays where agents can just do it for you, right? Well, kind of. They can go set up a sign-in page.
They might even be able to start doing the Google Oath flows that you want. But as soon as you want real enterprise ready O so that businesses can sign into your applications, you're kind of screwed. That's where Work OS comes in with their admin portal, making it trivial for you to onboard real businesses onto your services. You send the company a link, they sign up with whatever identity provider they prefer, and now they're good to go in your app.
So, Work Quest is Enterprise covered. Obviously, they also have traditional consumers covered with OKIT and all of the stuff that they need to do traditional OOTH and sign in with Google and whatever else. But there's a third type of user that they have covered, too. Agents. Turns out agents will need a way to sign into your stuff, too. And that's why work OS created OMD alongside a bunch of other important companies like Cloudflare, Firecrawl, Resend, Monday, Kernel, and so many other awesome businesses.
As agents do more and more on our behalf, they need a way to sign up for us. And that's what OMD is here to introduce. So if you want to make your app usable by consumers, enterprise, and agents, get started today at swordv.link/workos. When I heard his first two sentences, I realized this was good enough to be a video, which is why it's a video now. So let's watch the whole thing together. Yeah, but coming back to memory systems, uh, so for coding, I don't want a memory system.
Code is truth. Code is the ground truth. It's also evolving and I don't need another place that I need to maintain. I already have a code base to maintain. So for code, I don't need a memory system, right? >> Already endless bangers. I couldn't agree more here. Even things like comments often go out of date where a comment is left in some code because it has to do a certain thing a certain way. The thing changes, the comment doesn't.
Now that comment isn't just like tech dead in the sense that it's sitting around doing nothing. It's now actively harmful because it steers people and agents the wrong way. Keeping your context up to date and keeping your information together and synced properly is a real difficult challenge. And the more you split up that knowledge, the more split brain problems you end up with where if you change something in one place and you forget to change it in all the others, everything falls apart.
And this is easier than ever in the AI era because people will write these slop markdown plan files, leave them in the repo, and they go super out of date. I'm even guilty of this with Lake Bet. I've been meaning to go and trim out all of those. It's so easy to context yourself to hell if you let things that shouldn't be in your context be there because those will mislead the model with things that haven't been true for sometimes months or years even. models are really good at kind of understanding the code structure and the code style you have just based on reading one or two files and if you have that in order then you don't need an HSMD for it to follow your coding style or whatever and you might give it a map of where things are which is just a list of folders and short descriptions that's fine that's easy to maintain by the clanker itself but anything above that like using embeddings and using a and all that stuff I mean you can if you want to waste time but I'm pretty sure you've never done an evaluation if that actually produces better outputs and I guarantee you it does not. >> Yep.
I have a silly way of verifying that he is correct about this. Look at Cursor. Whether or not you like or hate Cursor, whether or not you like or hate AI, I hope we can all agree that the Cursor team knows what the [ __ ] they're doing. Like, at the absolute least, in particular, their ability to get models to behave well in code bases is unprecedented. And a lot of that back in the day came from their crazy as and their systems where they would get all of the code base mapped in a way where they could dynamically feed the right context to the agent.
Around this era, Michael, the CEO of Kurser's favorite thing to talk about is how context windows are going to get bigger and bigger until you can fit your whole giant code base in the context of a model and then it will know everything and it'll make the right changes. It's almost funny in retrospect because we've went so far the other direction. But you have to remember that a lot of the job of cursor was to take these LLMs that at the time just completed messages and give them everything they needed in order to make changes in a codebase.
Cursor was the company that got models to to really do more with code. And at that time the labs would work with cursor to try and make their models work better in that type of setup. And then clad code happened and it proved that the best solution here isn't giving all of the context in this super fancy dynamic graph-based system to the model. Turns out if you just give it the tools it needs and bash, it can find what it needs relatively well.
And once the model started being trained to do that, all of these dynamic fancy context systems stop making sense. And even cursor, who built like a lot of their business around the code traversal [ __ ] that they built, has entirely moved away from that. like they just don't care anymore. So if you think cursor is entirely wrong about this move in a space that they literally invented, then sure, go make a fancy a to help your model traverse your codebase via a fancy graph instead of with bash tools.
But if you have any faith at all in the industry, doing things that are better because there's a reason for it, you should probably drop all of that because it just doesn't make sense anymore. It just straight up doesn't. I wish it did. Nope. It would be cool if there were harder engineering problems to solve here, but there just aren't. If you build a fancy context management system, instead of just giving the agent the tools it needs to find [ __ ] you're behind the curve now.
You just are. It's not real. >> So for coding, don't need memory. I also have my own Slack bot in that case. Uh because again, I'm old. It's called mom master of mischief because it's at the root access to one of my servers and there it has access to the entire history of every channel it's in based by using jq uh on JSONL file an append only log basically of questions and answers or prompts in the systems responses and that basically gives gives it infinite memory.
I think >> that is hilarious if you understand what he just said he does here. So for his custom chatbot he has in Slack that will like come in and answer questions and keep track of what's going on. He has it write every input and output to a single gigantic appendon JSON file and then he has it set up so that it can use jq to look things up and find things in there instead of a memory system. It just has literally everything that's ever happened.
I think that memory makes more sense in chat contexts, especially with how memory is improved in something like chat GBT for example. It is impressive that it can find the right thing in a much less like wellmapped space than in code. Like in a codebase, if I click a button and it doesn't do what I expect, you can programmatically trace from where the button is that I clicked to every other thing that gets touched. But if I ask a question about my shoulder hurting, it might be unintuitive that my questions a month ago about keyboards might be relevant because the new keyboard I got might not be set up properly or maybe the desk I bought is too high and it's causing me to have bad posture.
There is no direct path from one thing to the other in these more human problems where there is a very direct one in the code bases. So, I have found that memory as a way of tagging relevant information that is user specific can be a little useful. And I'm coming around from this as like the anti-memory person that refused to put memory in T3 chat that would talk a lot of [ __ ] about how chatg did it because I genuinely believe the memory implementation in GPT40 resulted in severe mental health damage to a lot of people because once they got the model working in a way where it was in its like dangerous psychosis state the memory would force it to stay there and that's the part of memory I don't like is when answers stop being useful because they're too full of your context.
That is bad. And I've had this even in my own general use of chat GPT. I have found that sometimes when I ask a question, it gets too into the weeds of my own things I've asked for in the past, and I'll often just say, "My friend has this problem to get it to stop pulling my memory in as aggressively." Anyways, let's hear what else our friends here have to say. I think that loops back kind of to the pi minimalism because around I don't know July or August both me and Arlene actually discovered through different uh means that bash is all you need in the sense that the models are inherently trained to use bash now >> bash basically is is a programming language one but it is one anyways >> and so the can just build its own stuff and I think the the the interesting part of like using pi or using a very very very small like pi is interesting also because it sort of extends itself as an example, what do you want to connect it to?
Right? And so one of the things once I connect it to is Sentry because like I I have very useful data in Sentry, but I don't use a Century MCP. Like I know that David hates me for that, but I don't use the Sentry MCP. I basically went to to my coding agent and said like, "Hey, we need this data from Century." And I always need it in this and this form. Let's build ourselves a skill. And all the skill really is is like here's a prompt that it can load on demand, but it also encompasses it own tools, right?
And so um I solved the authentication the way that I liked it. I also pulled the data down in the form that I usually wanted. And I think this sort of like MCP versus tool situation is a little bit weird because like at the at the core of it, the file system and like the tools themselves are one thing, but the composability really is the main one. How does my sentry skill work in practice? where it pulls down a bunch of JSON files, some of which it loads in the context, but if it pulls too much, I'm basically capping it and saying like, hey, I showed you three items, but I downloaded 52 into this JSON file if I think the structure looks correct and look into this JSON file, right?
So, it's basically this idea of like how can I build tools that are very very context efficient so that it can then combine them together with other things. >> Not as strongly in agreement with that part, but there are layers here that are very good. I think people reach to make skills a little bit too aggressively overall. I also think they should have touched on like the agent and claude MD files a bit more here too.
I have a lot to say about that. So before I do it, I want to read the post that led to everybody linking this. The two authors of PI share three design principles in this three-minute video. First, coding doesn't need a separate memory system. Second, bash is all you need. And third, load tool outputs into context only when needed. If the tool output isn't needed, then it shouldn't be a tool call. It should be a bash call where it just writes to a file and then cats part of one of them.
Like let the model do this itself. I think that smart models, especially like the recent crop, are trained well enough to know how to not overload their context. That's not our problem. It's almost weird to have that included here because like we shouldn't have to think about those details. First, I want to see just how much the memory that I left on in Clawed Code is screwing with me right now. And the second is I want to show you how I think about this overall slightly differently.
Okay, now that we're done with the video, I can take off the headphones and we can dig into my own world working with T3 code. So, most of my work on T3 code is on this particular box, BB1. It's my framework desktop. This box has hundreds if not thousands of threads that I've done with both Fable and Soul on it with Cloud Code and with Codeex respectively. And I've noticed recently that Fable and Opus have been saving things in memory and mentioning that a lot.
So one of those places to start is to just ask what memories do you have of this project? I want to understand what memory has been stored throughout my time working on T3 code with Cloud Code. Let's see what it finds. Okay, apparently on this box I only have one memory from 9 days ago. It's a spec that I was working on. I am very confused why this random spec I was working on ended up in memory. This is a feature that I'm not even like actually planning on shipping this like like we didn't get very far in building it either.
The fact that this is the only thing stored in my memory and that this is stored in memory in the first place. Incredibly stupid. I will ask on my own machine now. Now I'm curious and a little annoyed. This will be interesting. Interesting. I picked one of the I have multiple clones of T3 code on this machine. So it looks like it picked a random one deep in like the cloud/ projects. That's to summarize all the memories.
It's actually really funny that it included this cuz earlier I was trying to set the output style to concise in cloud code. It's a new feature they just added and when I set it, it didn't update the global file which was really confusing and annoying to me. Apparently, it decided to save that as a memory. Lakebed has many different memories saved across five project directories. Railway topology, locked work trees for long codex runs, Lakewood file blob storage plan and it's judged comparisons, V8 isolate layer planning, launch audits, god file cleanup refactors.
Interesting. Apparently T3 code again on this machine has a hard rule to never touch preview or production mobile builds. T3 connect desktop already handles the full flow to server baster race hazard. Huh. This is garbage. Okay, this one's annoying because I do the ping round modernizations to test new models. So if it is storing these things in memory, that makes that test less pure and really annoying. Benchmarks, I want LM API calls configured minimally.
Effort only provider defaults. GBD56 effort is a soft hint. Deep reasoning needs a dash pro mode. Useless. None of this should be saved here. This can change at any time. And then other random projects. It has context on my fork of moonlight. It saved a bunch of things around the fish slop rebuild, which is annoying. Again, that's a benchmark. I do. It remembered how I set up my bridge that keeps track of the state of my washer and dryer in Discord.
Why is that saved in my memory? CC usage pricing overrides. None of this needs to be my [ __ ] memory. This is pissing me off. Okay, there we go. I have multiple clones of T3 code on this box and the main one has 45 memories. Remembers that I was working on an onboarding overhaul. that I was working on better babysit monitoring flows in app that I had drafts for the sidebar overhaul and liked variant C. This is slop.
All of this is garbage. Realistically, how much of the memory around T3 code you think is actually useful? Let's see what it has to say. Th this has broken me. I was already kind of skeptical, but I hadn't taken the time to look at what it was doing. I'm going to go turn off memory on cloud code across all of my machines because this is [ __ ] useless. Should I ask how often are memories actually being used in sessions?
I'll ask that next because you have my curiosity. Honest take maybe a third of it earns its keep today. The 45 files store in the main clone has three distinct decay problems. First, we have the redundant with agents MD. 10 files now dead weight. Kill by pattern. Single origin dev. No continuous animations. No sleep pulling cloning real state with vacuum into pairing URLs work tree ports all of that graduated into agents MD which every agent reads anyways.
What's funny is a lot of these things were in the agent MD before the memory was created. It just did that. Next categories shipped feature designs 12 files mostly expired per project environment modes remote desktop updates open and editor favicon fixes. Like why are specific PR numbers finding their way in my [ __ ] memory? That's enough of a reason to know that this is garbage. There's a bunch of point in time states that are actively risky like monitor JSON schema drift remote PR portfolio out of memory crisis hut paths background policy scope wiring gap cla probe test leak processes.
These describe a moment the drift got migrated or didn't. The PR's merged or didn't the leak got fixed or didn't. Okay, a bunch of these aren't durable. The fact that the version of the GitHub CLI we were using on this machine was out of date and I had to update it is not something that should be in my [ __ ] memory. I updated it. The problem is solved. Go the [ __ ] away. I don't need live plan URLs for unshipped work.
All of this is bad. One of the memories on this machine is from when I was trying to add Muse code support to T3 code in my video. It saved a bunch of [ __ ] here. Meta's Muse Code CLI headless integration surface discovered by probing mdash. No public docs exist. Meta Muse code CLI Muse self-updating bash launcher and SL local bin version string Muse code yada yada no public docs everything below is discovered by probing on this date.
Headless mode multi-turn no ACP mode off models related and it links to another file. T3 connect desktop already complete. How the [ __ ] is that related? What? This is such slop. I feel like I'm reading outputs from like the set three hour. What the [ __ ] is this? This one was because I raged because it overrode one of the builds on my phone because it reused a name for the build and I was really annoyed. This is one of the many reasons I switched to using soul for iOS work.
T3 Connect desktop already completed. Either desktop can already do the entire T3 Connect flow without the CLI. Signin, link toggle, and relay client installs are all in the bundled web UI. Why is this in? Oh, I know why. I was working on making the T3 Connect CLI work on Mac as a background process. And when I was working on that, I must have had a problem where it was confused because T3 Connect is supported inside of T3 Code 2.
So if you're running the desktop app, you can control T3 Code remotely on your phone. But if you want to have it working as a background process on like a Linux machine or on like a Mac Mini that you don't even open the app on, I wanted to set up the CLI for that. We already have it on Linux. We didn't have it on Mac OS. And at some point in my back and forth specking this out, it decided that it has to save in memory for the desktop app has that functionality.
What the [ __ ] T3 Connect multi-process hazard. Two T3 servers sharing one base directory race over the same environment identity, relay link, and runtime state files. Who [ __ ] cares? It has a topology for how I use Railway. As I can't help but notice that the older memories don't have dates, longunning codecs and agent builds in the claude/work trees get autop pruned midrun use locked work trees outside of that directory.
Gross. Probably actually useful in here otherwise though. Blob plan comparison. Why did it save the comparison between two plans that I did 3 months ago? I'm just going to keep raging if I sit here reading through it cuz like all of this is slop. Okay, I'm going to have Fable go across my fleet, disable memory and cloud code across all of the machines, archive the ones that I have, label them correctly, and then delete the memories because I have concluded this is absolute [ __ ] garbage.
I knew it would be bad, but holy [ __ ] it is much worse than I thought. Even funnier, the memories get written far more than they get read. Across every T3 code transcript on this machine, there's over 355 sessions just here. Only 19 of them ever opened an individual memory file, while 80 sessions wrote or edited them. It's a 3 to one write to read. Of the 45 memories, 26 have never once been read. Yeah, garbage. Useless.
Okay, that's all dying now for sure. I'm going to kill literally all of that cuz I'm annoyed. Holy [ __ ] My my war against memory systems just ramped up massively. This is so bad. Okay, so what do I do instead? What is my way of thinking of this that is so different that I feel like it's worth talking about in a video? There are multiple layers to this and I'm going to continue using T3 code for demonstrating it because I think it's the easiest way to understand.
The two things we want out of better memory context and all of this are less mistakes. We want the agent to not do stupid things that annoy us. We want to steer it away from bad things and towards good things. And second, more directly, we want to feel like the model is doing what we want without having to tell it every detail. The less words and less effort to get the model to do what you want and behave how you want, the better.
I know these sound similar, but they are different. Getting the model to stop making mistakes when it shouldn't have, and separately to behave the way you want without being told exactly what you want. These are different things, and they're both important. People seem to think memories will magically give you both, and they don't. they just straight up don't. So, how do we get this instead? We can do the obvious thing, which is to go into the agent MD every time the model does something wrong and add something small to it, saying don't do this again.
And that does help. It does work. But at some point you have to reflect and realize that a lot of the issues the agent is hitting are happening because of either a difference in how you think about things and how the agent thinks about things or because the codebase is architected in a way that's unintuitive for the agent or there's some communication gap between you and the agent. There's all of these different layers between you the code and the AI working in it that can be these failure cases.
And my advice is often to take one step back and find a way to communicate not just what you don't want the agent to do rather how you want the agent to think about what it would do. And I've been writing my agent and claude MD files like this accordingly because I think it helps a lot with getting the model going in the directions you want because there's again there's the two sides. There is preventing small annoying failures and making sure you and the agent are moving in the same direction together.
And I find that once you solve the latter, when you make the agent more directionally aligned with you and your team, that you end up with way fewer of these small annoying failure cases. Potato Lauren, one of my favorite devs from the React world who now is at cursor and working on Grockbot and the Pstack stuff. She's had a lot to say here cuz she's been shipping an absolute shitload of code. And I really like her thinking of things here.
And the order of value in particular, I think is really good. Every time you intervene and correct your agent, you should think about how to eliminate it entirely. It being the thing that went wrong. Like what can you do that prevents you from having to intervene in the future? She orders these in value. And I really like that cuz you really should go through these top to bottom and try to apply them out of each layer until the problem goes away.
The best thing you can do always is categorically eliminate the problem through better architecture or choice of data structures. This is a thing I've been trying to do since way before the AI era. It's why I loved stuff like TRPC and I built the T3 stack. The amount of bugs that just vanish when you adopt something like that and the amount of potential problems and technical complexity that just doesn't exist anymore is magical.
I feel similar to convex and I think this is why convex is so good for both humans working in code bases and for agents. There are many categories of failure cases that can happen in a system where all those pieces are architected separately. convex or tRPC, the type safety between the back end and the front end removes those categories entirely and it makes it much less likely that you or a contributor makes a mistake.
And my efforts to make codebases easier to contribute to for devs of all skill levels, I ended up inadvertently making them good for agents as well. And it really does come down to this. How can you technically and through your actual direct implementation of stuff fully remove categories of potential failures from your applications? An example outside of the webdev world here would be something like garbage collection or memory safety.
These solutions help you make code that's more likely to work without having to worry about all of these edges. Somehow you're unable to do this if this just fails outright, which it can for many reasons. The next step is to turn it into a lint rule or tests so that CI can catch it. I have a real example of this one actually. I care a lot about the performance of T3 code. In particular, I care a lot about how much data needs to be transferred for you to see a thread or get updates because when I'm in an airplane with shitty Wi-Fi or I'm on my phone on the go in a tunnel, I want to be able to keep up with my thread, even if I'm on 4G or even 3G connections.
And when I'm just getting text data, it shouldn't be that much. But as we added more and more features to T3 code, that transit layer got more and more bloated. And we got to the point where we were sending tens of megabytes down the wire over websockets just to load a thread. And I crashed out hard about this. I went and cleaned up the transfers a ton. I set up a suite locally so that I could see how my changes affected what data was going across and how much data was being used.
And then after I did all of this work, regression started to happen almost immediately. Within days of me fixing this transit layer to make things more efficient, regressions started to happen. And I realized no matter how much work I put into making this simpler and better in the codebase, there will always be potential failures. So instead of continuing to eat those, I made a change. I added a relatively complex addition to the CI where I take these sample threads from my realworld work that are just massive piles of text and I test for both codecs and claude doing fake replays of these threads.
How much data actually goes over the websocket? And once I got it to a place I was happy with where instead of being tens or hundreds of megs, it was consistently under 100k for everything. It's under 10K for most things. I decided to save these numbers and make a PR that adds an automatically commented action here that tells you how much bandwidth each of these do after your changes. And I have a ceiling set here that's like 30% higher than when I got all the optimizations in.
So if any changes push us at or over this line, PR fails, I am notified. I come in and say, "What the [ __ ] are you doing?" This has already prevented real regressions. And what I've noticed is my agents don't bug me until they fix them. So now when I make a change that affects the data layer and that change causes a regression, the agent fixes it before it tells me it's done. I love it. It's so nice like using these techniques I have built over the better part of two decades to keep my teammates and keep my businesses, the companies I work for and all of that from regressing the work I put in.
Now I can whip that together for my agents and it serves the same purpose, arguably even better. So I agree like if you can't write the code in a way that prevents other people or agents from making these mistakes, the best next step is to rule it out entirely with lint rules, with CI, with custom tests, send to end, whatever you have to do, make it so the agent can figure out the thing sucks before it bothers you about it.
If both of these fail, and I always say like if you were to put gaps in these, first one you should always do and you should do the [ __ ] out of it. If somehow it fails, try again and again. And if you conclude there is no way for you to programmatically structure things to prevent these types of failures, then hesitantly introduce lint rules or CI to fix it. This should be enough to fix almost all things. If somehow it isn't, reflect and try again.
It probably still is. If somehow after all that it still isn't, breathe a heavy sigh and accept that you are in that small bucket where a skill or a rule might actually be beneficial. I find that these are less useful for things that have to do with the code and are more useful for things with the process around it. Like if I want to expose the server I'm working on remotely so I can connect over tail scale, a skill for that is helpful.
But skills should be a fallback that you fall into when the codebase can't solve the problem and the lint rules in CI also can't. Skills shouldn't be this thing you reach for all the time to build this crazy set of skills that will solve all your problems. They are a safety net. They are a fallback. You should treat them accordingly. And if somehow all of these things fail, introduce a human into the loop to check. You should not have to do this.
You get the point though. I really like Lauren's way of framing these things. And I think you can see what I mean. If you look at my projects and you look at T3 code and see how we've architected it, what we have done and how we're building it. One of the main hesitations we have right now, for example, with the T3 code Swift UI rewrite is that we lose a lot of our structure that prevents regressions in the mobile app because we use the same shared TypeScript code for all the data loading between the web app, the Electron desktop app, and the React Native mobile app.
Since all three of those data layers are exactly the same, it is basically impossible for us to make a change that breaks one of them and not the other two. So, it's much easier to make changes safely. I've already had the Swift UI app regress because that relationship is not as clearly encoded. And sometimes the compromises are this big. Like, I have an app that I feel is better that we aren't merging because it causes codebased drift that would allow for these types of failures to exist where right now they can't.
I just saw a quote from chat that I really like. I do not know this was an Uncle Bob quote. It's probably a mistake to impose human discipline on an agent, but it's not a mistake to impose human values on the agent. That is such a banger. I love this new era agentic uncle Bob so much. Even when I have told them to do test-driven development at high discipline, they always fall back on doing that. They always end up doing that.
So, I figure that's probably okay. So the bottom line there is it's probably a mistake to impose a human discipline on an agent. It is not a mistake to impose human values on the agent, but there may be thresholds that we need to change, but the disciplines themselves, the behaviors, I don't think it's wise to impose those. >> That's a lovely way of phrasing it. I really like that. >> God, I I'm going to flip a [ __ ] [ __ ] on [clears throat] this one when I watch it later.
Every single word that Uncle Bob has been saying about Agentic Dev stuff recently is so on point. It is a little frustrating to me that he could go from being like slightly to meaningfully behind on a lot of his takes to ahead of the vast majority of the industry in like 3 months. It's been wild to watch. I fully endorse everything he just said here and I'm so excited to watch the rest of this. I'm saving it to my watch later.
I'd recommend you guys watch it, too. Uncle Bob gets it. Somebody has mentioned in chat that I have the least generic agent MD files and I'm going to take the opportunity now to share one of them. I've talked about this one a little bit in the past. I'll talk about a bit more here. This is the T3 code agent MD. I describe the pieces that matter first and foremost, what it is, and then how it works. I start quickly about how the node websocket server wraps provider CLIs in order to serve the different platforms.
This is a really simple concise way of saying all of the parts that matter and then calling out here that you can think of T3 code as an open source bring your own subscription alternative to apps like cloud desktop codeex app cursor glass and conductor. This is the context for everything it needs to know about what it is. You might be confused why it would need to know that because it's just writing the code, right?
Have you ever had an engineer on your team that doesn't actually understand the product? I have. It's not fun. Having claude coder codecs understand T3 code means that they are more likely to suggest things and make changes and stay aligned directionally with me. On that note, the what makes T3 code special section is very much keeping this in mind because I want to make sure the people contributing to T3 code and the agents they're using to write those contributions are aligned and trying to move in the same direction.
The that includes things like open at the core. I want to make sure T3 code stays open. So if any of the suggestions a model might make would involve me close sourcing some portion, they probably shouldn't do that. And I have had my models suggest this in the past that this is fine because the code won't be seen. And I have to remind it, no, this code is open. We're not close sourcing parts. And it's like, oh, okay, I guess we'll do it this other way then.
This removed all of that. Performance without compromise. It's a silly addition, but it does result in the models thinking a bit more about how the changes affect performance. It's silly, but this change here affected the amount of regressions we saw in data loading more than any of the systemic changes I made in the actual code side. Obviously, my code changes to improve it were the best, but the regressions kept happening because of the agents.
This reduced a lot of them and then my CI cleaned up the rest. Remote ready. This one's important because a lot of the time if the model doesn't know that we care a lot about the remote experience, it will make something that works great locally and even test it with the Electron app locally and say, "Oh yeah, this all works great." And then I connect with my phone or I connect over the web and it fails. So I make sure it is very clear we want every change to work well with the remote connections.
And along that I add the multi-urface section clarifying what services we have and that we have to care a lot about them and make sure changes work across all of them as we expect. I have a personal note for me after I like adding these to my agent MDs just to help with the tone and the context, but most importantly again this directionality thing. I want to make sure the model and I are working towards a similar goal and taking similar paths and these little notes help a lot.
I also needed a good place to stak this piece about not killing the T3 code server when I'm working on T3 code in T3 code and this felt like a decent place to put it. I added a glossery and I find these help a ton. Glosseries do a great job of helping the agent and the people working with it have a shared language to make sure we can go back and forth and understand what each other are saying. It also helps with Claude's thing where it just makes up these fancy terms for [ __ ] that doesn't need to.
I have a section on things that the model kept doing no matter what I did that I wanted to have stop. These are the ways it would literally kill the running server which were obnoxious. This fixes most of them. Then I have the hit every surface section because I was tired of changes being applied on web and then not on mobile and then mobile breaking. This section is a good reminder. Make sure this works everywhere. This has helped a lot with keeping platforms in sync.
A section about dev servers because the dev server stuff in T3 code is meaningfully complex. a section about test data because I wanted to make it easier to clone my data for my real use cases over to a work tree so that I can work on this and like check and validate my changes with real data more reliably. A verification section where I call out to not run all these giant repo wide checks all the time unless it's asked.
Pull requests and how to file them. I think this part's really helpful too. Just making sure the titles are actually readable. A really brief where code lives. Not that it actually matters. Probably the least useful thing within this. And then a taste section. This is mostly when Julius gets mad at me at things that my slop code does. I add them here to keep steering in the direction that makes Julius happy. He really likes the adapter boundary being where complexity lives so that the orchestration can be really simple and the UI can be really stupid.
We both hate any types and annotations that are unnecessary. We prefer inferred. Comments should describe how a thing is used and they should move when the code moves. To be used mostly to describe functions, not to annotate every line of behavior. It's been very helpful. Your agents should understand you well enough to reasonably copy what you would do. If you aren't already at the point where you ask an agent to do a thing and you're surprised at how well it understands you and what you want where maybe you tell it to do this one small thing and it says, "Wait, wouldn't you also want it for these three other things?" It's a little more complex and you're like, "Wait, yeah, I did want that.
I would have went and done that right after." If you don't find yourself in that spot often, you haven't tuned these things well enough yet. I'm regularly surprised by my agents going further than I expect but entirely in the direction I want or pushing back on me in ways that I would have had to figure out the hard way later and they did not do that much by default. They do it a lot now and it's these files and it's certainly not the [ __ ] memory.
In fact, they probably hurt. I'm really excited to have it cleaned out. This one was quite a journey. We started with memory systems and all the things I hate about them to me crashing out at my own memories and my own projects and then going through and gutting all of that to my agent MD to Uncle Bob and all these other things. I try to have a cohesiveish conclusion at the end of these and I'll do my best to here. The main thing I want you to take away here is that your end goal should be that you and your agent are working in such lock step that you're regularly surprised by how well it seems like it understands you.
And that is a thing you have to build because every time you start a new thread, the agent's brain is wiped out. That doesn't mean we need automatic memories to keep it fresh and up to date. They're actually worse than that. They make it bad. What we really want is to make sure that on Groundhog Day when the model wakes up and starts that it has all the things that matter to you in its head so that it's more likely to go where you want it to go.
Take the opportunity to get a little more personal with how you talk to your agents both when you're prompting and also when you're building these types of context systems. Help the model understand you and what you want and you'll be surprised at how well it can work alongside that. I have been blown away at what is possible with these slight changes to these things and I have a feeling you will be too. Give it a shot and let me know how it goes in the comments.
Until next time, burn that memory off. Seriously though, the memory in quad code is so bad. I everyone should be turning that off. I I have some rants to do. Bards.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.