Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Matt Pocock · @mattpocockuk
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Matt Pocock's most watched videos.
Most replayed moment at 6:58
2.7x that video's typical replay level
grid keyboard. I don't remember that one. And so, what I did was I then worked through each of those tickets in a new session. The way I did that was I just called Wayfinder on that ticket name. I did it in a slightly fancier way where I actually have a handoff skill that automatically wrote me a prompt and
Said at 6:51
Most replayed moment at 5:53
2.5x that video's typical replay level
under the hood. Uh, let's get into it. I'm going to open up a new Claude session inside here and I'm going to run my improved code base architecture skill and we'll turn off auto mode. Auto mode does some funny things with these human in the loop style flows and so I don't want it on here. We can see it's going and
Said at 5:45
Most replayed moment at 4:48
3.3x that video's typical replay level
CI script that does type checking and runs the tests." So, now I'm going to create that issue, and we can now run our agent to see what happens. So, after that, it should be ready to be picked up. First, I'm going to add this little piece of code to my package.json here, which is just going to allow me to run a
Said at 4:40
The graph counts replays. It does not show where viewers stopped watching.
Words
12,288
Runtime
1:06:45
Speaking pace
184wpm
Reading time
51min
184 words per minute, just over the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
We're live. Okay, sweet. I think we're live, everybody. Right. What I'll do is I'm going to cut this down on YouTube a little bit later and then we'll have like a proper I'll give it like a proper intro and I'll, you know, introduce you and all that stuff. So, I'm going to hide you from the main chat and then I'll bring you back in. Does that sound good? >> Sounds great. >> All right, cool. Oh, and also, how do you pronounce your um nickname? Is it potato? No,
92 words, the words spoken in the first 30 seconds at 184 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 532 |
| Average words per sentence | 23.1 |
| Longest sentence | 311 words |
| Questions asked | 132 |
| Sentences containing a number | 18 |
Most used terms
Filler phrases
930 in total: like 256 · you know 186 · uh 160 · right? 93 · um 91 · actually 59 · sort of 47 · kind of 21 · basically 9 · I mean 7 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
We're live. Okay, sweet. I think we're live, everybody. Right. What I'll do is I'm going to cut this down on YouTube a little bit later and then we'll have like a proper I'll give it like a proper intro and I'll, you know, introduce you and all that stuff. So, I'm going to hide you from the main chat and then I'll bring you back in. Does that sound good? >> Sounds great. >> All right, cool. Oh, and also, how do you pronounce your um nickname?
Is it potato? No, it's just potato. >> Is it just potato? Okay. >> Yeah, it's Well, it's supposed to be the Japanese spelling. So, the long story behind my my Twitter handle or my ex handle is uh it was when it was from a time when I was playing video games >> and I wanted a cute nickname. So, I was like, "Oh, potato." Cuz I love food. Uh but then obviously the proper spelling was taken. So, I had to improvise. So, I came up with >> potato.
But I just say potato. >> Okay. Potato. is my business partner was obsessed that it had to be potato. It was potato. Yeah, Joel Hook. >> Yeah, I guess if you want it. >> Yeah, accurate. >> All right, let me intrigue you then. So, hello folks. I've got another treat for you today. Last time on this kind of podcasty thing, I suppose, we had Uncle Bob and we talked about software quality. We talked about agents. We talked about lots of cool stuff.
Now, we have uh an incredible guest, someone who I'm delighted to welcome on, who's been exploding on Twitter recently about software factories, um about increasing the quality of your work, about increasing your velocity and climbing the trust ladder with agents so that you can ship more and more and more. And it is potato. Welcome. Thank you so much for joining. >> Thanks for having me. Yeah, very excited to be here.
Yeah, big fan of yours >> and a huge fan of yours. I think people have been talking about this like it's like the meeting of the skill minds, the skill Mount Olympus or something because both of us have very popular skill libraries. Um I've not, as I was saying before we started, I've not used a ton of yours and like I want to get all of the juice out of your brain so that I can go and use it properly and use it better.
And I think where I want to start with this is you gave a talk um pretty recently like about 10 days ago and posted on X which went absolutely nuts is about how I shipped 2,500 PRs last month to production got about 3 million views or something on X and I watched it and I loved it and I recommended it and I kind of want to run this as almost like a Q&A of that talk basically of giving you because it just I just had tons of questions about it and I wanted to dive into it.
And I think where I want to start is you talk about a trust ladder with agents where you as you trust agents more, you can get them to do better and better things and or scale them to up to use more and more agents. So what is your story of how you climbed the trust ladder and how did that work when like you got SpaceX and >> started climbing more and more? So I think this the the journey sort of began even before I joined cursor uh which is now SpaceX AI.
Uh so the story is um after Meta so I I used to work at Meta on the React team. Uh I took a month off uh because I was feeling kind of burnt out and of course when what what do you do when you're burnt out? You go and start a new side project. Um and so I started a side project. you know, I was uh of course using AI to to write code. Uh but then I started to realize uh you know, I was spending like so many hours just micromanaging one agent, right?
And you know, at the time, this was back in February, maybe February, early February or January, you know, people were really obsessed with this idea of like orchestration. This was like, you know, before, you know, things like cursor, you know, like the agents window was had become popular. So people were still in like like 2 land you know in their terminal and they were all talking about okay here you know I built a custom orchestrator right and so of course I had I was a bit nerd sniped by that and you know as I was building my toy project uh I got nerd sniped by oh how do I make my AI coding setup more efficient and so you know I I kind of started the journey there where I just you know took a step back and realized yeah I was spending all this time micromanaging a single agent you know I was creating skills and I was like finding it quite difficult to measure the output or the the result the impact of the skill as well so I was kind of flying blind but I was you know iterating really fast um and um so that project eventually sort of became the basis of PAC even though I didn't know it at the time um and a lot of some tricks I had learned like building that early set of skills.
Actually, it's still open source if you want to if anybody wants to take a look. It's on my GitHub like potato/ noodle n o dl e. Um, and in there you will see some skills and a brain directory. And so I was really interested in this idea of how do I, you know, extract my own ability, if that makes sense, and give it to the agent, right? because I was I I realized that you know all I was trying to do was trying to teach the agent to write code more like me you know do do you do do workflows more like me.
So you know the skills were like an entry point to doing that. Um and then you know after I joined cursor uh I was starting to work on the agents window and uh it had a lot of performance issues. Uh it was it was it was pretty laggy. Uh and so since I had experience working in React, I was asked like, "Hey, do you want to come and help out uh with the agents window? Um and so the the beginning of the cursor journey was very manual.
Uh I was deep in like looking at like flame graphs and heap snapshots and trying to see like why exactly is the app so slow." Uh but then coming back to the same realization like you know I was sort of the bottleneck. I was doing everything manually. I was sort of the meat proxy in a way. I was the meat proxy between my agent and Chrome DevTools. Uh and I was like really annoyed by that. >> And what month of the year is that?
Let's say where are we in the timeline? >> Uh so I joined Cursor in March. So this is like early early April probably early April is when you know uh I joined and I didn't have any skills, right? I had I I sort of abandoned my personal skills because I didn't think they'd be relevant anymore. Uh but then working on the agents window uh and now working on grockbot uh I sort of realized that a lot of the lessons I had learned from those skill time building the the initial set of skills were very relevant especially around things like verification uh you know being very rigorous in your work um because I think agents from my experience even the the frontier ones tend to tend to take shortcuts.
Uh they tend to do the easy thing. Uh so uh a lot of the skills that I've built have been around how do I make the easy thing the right thing? You know, how do I make that the best thing? >> The idea of sort of distilling your expertise and turning what you do every day into processes, that's something that feels super familiar to me. That's exactly what I've been doing with the skills. And I suppose there's something in that which is a lot of people think domain expertise is getting less useful now as people uh start to rely more on AI where what do you think about that just as a sort of vibe check before we start talking >> I actually feel like domain expertise is is more important than ever you know uh I think I wrote this on my ex at some point but you know at times I sometimes think of you know AI as is like especially as the models get smarter and smarter and more capable and the frontier models are just getting so good like I love Opus 5.5 by the way um uh you know as the models get really really really good it almost becomes like the bottleneck is no longer the agent right it becomes your ability to express your intent and your goals in a clear way that the agent can understand and actually carry out and That's where I think you know like people with a lot of domain expertise are extremely have a have a huge advantage in my opinion especially if you're a little bit like you know tech technoc curious you know so I I think of people like you know like uh like a doctor or a lawyer or you know someone who who has a deep expertise in a particular non-engineering domain and if they're actually just a little bit techsavvy and they can figure out how to use agents they can actually build really really great products, right?
If they if they have a clear enough vision in their head and they can articulate it in a way that the agent can build it, you know, I think that that is really the the bottleneck these days is is like the transfer of your intent, right, and your vision to the agent. >> Yeah. I've been obsessed with language basically since agents um dropped. are just obsessed% thinking constantly about the the composition of words, how I can make things sharper, what um what might be hidden in the phrases that I'm using.
And it's and finding what I love is when you find a word that the agent then hooks on to and then goes, "Okay, I'm going to reinforce that word. I'm going to reuse that in my thinking traces." You know, I found that with um TDD was an early example of that. a lot of chat about TDD recently of like, you know, people say, should you use TDD with agents? Doesn't matter. What you're doing is you're getting the agent to think about TDD, getting it to write tests, getting it to prioritize things in a different way than it did before.
And that's why sort of grilling, I think, works effectively. Grilling is a >> Yeah. Yeah. It it draws those words out of you, right? or it or at least it helps the agent understand your thinking so that they can propose those words to you and you can pick up and say yes exactly that. >> Uh I've actually copied some of the the tips that you've shared as well where you know one of my favorite ones that you've shared recently or or not or like maybe in the past couple weeks is about uh reducing or eliminating tautological tests.
Like one of my pet peeves of agents is like all of the useless tests that they write. And so, you know, that was one thing where, you know, the word tautology, right, is is is I guess, you know, not many people necessarily know that if if especially if English isn't your first language, but there's a lot of meaning to that word. And it's like it's almost like compressed, right? Like you compress a lot of intent and meaning into words.
And so I I I totally agree with you like I think language I've always been interested in language actually uh like programming languages natural human languages and how they came to be and it's so interesting that now with agents it's sort of like this meeting of natural language with programming language but it's all it's all language out of the it's all communication. >> Totally. I did a drama degree right so you know I've been thinking about language and Shakespeare and stuff for a long time and so this all feels very familiar. >> Yeah.
Um, so okay, there's sort of before we get into like because I think the thing I want from you is like software factory stuff, right? Software factory is the big buzzword. Software factory is the thing that I'm thinking about too. I'm sort of releasing a course in that direction too. >> And it's this sort of scaling yourself up to un unrealistic numbers of PRs basically or PR numbers that sound ridiculous to people who don't understand how this works.
So where I want to get to is sort of from people who are doing kind of like one to five agents today up to, you know, hundreds of agents running at once and how that sort of functions. And so I'd love to hear about your metaphor of the Michelin Kitchen instead of the software factory because I think that says a bit about the way you think about this stuff. >> Yes. Yeah. I I I've I've never really liked the term software factory.
Not because you know it's not accurate but I think I think the a lot of people when they think factory right they don't necessarily equate that with quality or craft right things which are very important to me and a lot of people and technologists who work you know building products we care about the user experience we care about the things we're building so while so while I do think software factory is an apt term it also I guess maybe conjures up negative you know, maybe sometimes negative connotations.
So, Michelin Kitchen is the thing that I've sort of landed on where it's much more I feel like it's much more aspirational and uh I like the metaphor a lot cuz you know I like food. I'm called potato of course and I like cooking and I see a lot of parallels, right? Like with food, right? When you're cooking a meal for yourself, for example, it's both utilitarian, like you're trying to just feed yourself, right? And and survive. uh but it can actually be transformed into art, right?
And that's what what a Michelin starred chef or even just a chef or a cook can do with food is take something very ordinary and turn it into a delicious meal that you know takes you back to your childhood days or something like that. Um and so it almost like mirrors that trust letter that I talk about where uh you can sort of imagine your own journey as a home cook, right? Uh, as a home cook, you are doing all of the food, the cooking yourself.
You cut all the vegetables, you do all the prep work, you do all the cleanup, you know, you are the one man or one woman show really. Um, and it's an interesting thought experiment like, okay, if you were to cook a meal and then you add people, right, your your partner joined, your brother, your sister, and now suddenly you have your whole family in the kitchen. And I think most people would get very stressed by that, right?
The thought of, "Oh, so many people are just mcking around in my kitchen. They have no idea where all the utensils are. They're >> I have a max capacity of one person in the kitchen." Yeah, absolutely. >> So, I feel like that that's really apt because when you ask yourself that question of how do I go from being a solo cook, right, to having an army or even not not even an army but a few sue chefs, right, that that are helping me in the kitchen.
How do I think about dividing the work in a way that makes sense? You know, I'm not dividing work just for the sake of it, but in a way that actually makes the sum the to the better than you know the total of its parts. And so the Michelin kitchen metaphor to me like works really well in that regard because you know as a chef you're you know if you become a chef you're in a position where you're not necessarily cooking all the food yourself anymore but you are thinking you're almost like the tech lead right for the kitchen where uh you know chefs have to think about you know not just cooking but they have to basically organize the whole kitchen and they're like the CEO of the kitchen.
You have to think about when do you order ingredients, how do you store them, how do you prepare them, when do they have to be prepared. You know, it's a whole it's a whole job, right? That's not just cooking. Um, and I think that again it mirrors so much of how engineers write code today where you are not writing the code yourself anymore. You have agents, right? But you as the human are still responsible for the final outcome, right? your name still is associated with the work that you do, your reputation and you know so how you set up your kitchen right and how you set up your skills your environment your codebase I think are ultimately the new ingredients that go into um building product >> yeah I think what I love about your approach is the amount of focus that you put into the environment that the agent operates in right because I think a lot of people they think right the agent is good I'm probably not going to be able to make it better let's just trust what these magic model people have put into the harness and the model combination uh there's nothing I can really do like I can't mess about with claw codes internals or something or whatever you're using um but what I love about your approach and it's something I advocate for too is that you can change the environment the agent operates in right you can make changes in the codebase and also give it tools tools for verification as well and allow it to verify its own work.
So the thing I I loved about watching that talk is the amount of focus you put in verification and like that is the lever that you can start to generate trust. Can you talk about that and what that concretely looks like? Let's start like looking at practical ways that people can improve their own processes, their own kitchens. Yeah, I've I've I've said this a lot actually that you know even if you don't use PAC or you know your skills I think that the single most important skill that should be in your toolkit is verification because without verification and for for by the way for those watching who don't know what that means it's this idea that you can give you can sort of give your agent uh hands and eyes in a way that's the the analogy I like where the agent is able to run the code, right?
And actually uh interact with it like a normal human user would and also do things like you know debug it, you know, take traces and snapshots. Um and uh funnily enough like that was actually the first skill I built when I joined Cursor. uh that gave me a lot of that was that was the thing that actually started to let me ascend the trust ladder a little bit in a way that some of the other skills I had looked at or built had not really let me do because no matter how good you know some of the other skills were like the how skill the Y skill the unsop skill were I was still relying on me right as the proxy between my agent and the output so that you know if the agent can't actually see the result of its work there's no way it can actually iterate right and so this is where people start to talk about loops this idea of a loop and really I think the term loop you know seems kind of uh almost abstract like people like what what what is a loop what is an agent loop but really to me like the most important part of a loop that allows it to be a loop is the verification part because the agent is able to to verify its own work and uh you know that takes you out of the equation where now I can actually do something like so the very one of the very first use cases I had for verification was you know like the performance work that I was doing on cursors agent window and I want I wanted to get to a point where I could do something called hill climbing uh which is a term that I I think the labs uh talk about a lot which is this idea that you know you have some kind of rubric or a way to judge or score something and now because you have a loop you can have an agent continually try to make improvements to that.
Uh I think Carpathy, Andre Karpathy also famously released uh something called auto research that has a lot of these ideas. Um but yeah, verification I would say is probably the most important skill in PAC uh and many other you know tool sets. Uh and I think it's the most important thing to focus on. So a lot of the a lot of I spent a lot of time actually you know tuning the verification the create verification skill um and also internally the the we have so many verification skills now like every app that cursor has or spacexi has has a uh verification skill that is automaintained as well >> uh and it's become critical infrastructure for our team because everybody uses it >> and you went pretty far with that too right like you had a um in your talk I saw that you actually built a custom CLI for that too.
So what does that CLI do? Like how does it execute things and why did you I mean that's proof of how deep you're going right of how much you're pushing that. >> Yeah. So this is actually a tip I learned early on where um I guess you know back in January or or late last year the thing that people were concerned about was context window, right? That was the big the big topic at the time was how do I you know manage the context window because you know compaction summarization wasn't really that good yet and people were always people had this there was almost this meme in the community that you know once your agent summarized or compacted once it would become sort of stupid right for the rest of your session.
So there was a lot of thinking around like you know being very efficient with your context usage and so that was actually the inspiration for some of the uh the CLI work inside of the verification skills. I guess now it's less so about context because uh you know agents are much better or harnesses have gotten a lot better with summarization. Um, I still think there's some benefits to, you know, uh, having a clean context window.
Uh, so the CLI is really just more of a way for me to take the deterministic parts of what the skill does and encode that into a script or CLI to reduce to kind of take away the judgment that would otherwise unnecessarily be used because with judgment so I also think of you know agents and skills in sort of like it's like a gradient you have some parts of the work that are entirely judgment based right you know something that requires thought you know putting together multiple pieces of context thinking um and then you have the more deterministic parts like I don't know if you wanted to uh refactor some code right from one pattern to another that's very mechanical right you don't you don't need an agent to think about it and come up with it in a novel way each time right and so that that was really the inspiration for the CLI and you'll see this in a lot of the other skills that I've built is like I try to extract out the deterministic parts and turn that into code and just leave only the parts that actually require judgment to the agent.
So in a way I think of the skill as kind of like a wrapper, right? It's a wrapper with some light instructions around how to use these custom tools that are inside of the skill. Um but yeah, I don't think the CL is really that interesting in its own really. It's not like a novel piece of software. It's just something that interacts with like Playright and the Chrome DevTools protocol and calls a bunch of APIs. It's like it's just a bunch of glue. >> No, it's fascinating because it's a way of hiding information from the skill, right?
It's a way of conserving the skill, keeping the skill quite small, I imagine, and then you're able to delegate more of the complicated deterministic stuff into a script within the skill. So it's almost you're compressing information and making the agent do more consistent things more consistently. >> Yeah, >> that's fascinating. Yeah. And it it also helps I guess if you care about context window, it it does help because now the agent doesn't need to uh you know re reinvent uh things cuz uh one thing I had noticed early on when we didn't have a CLI was that uh well the the agent would try to verify his work but it would basically rebuild the world each time and then every agent did it differently.
And I was starting to notice like that's very inefficient, right? I was wasting it. It was actually not just about context usage, but also speed, right? Like cuz now my agent had to actually go off and write the scripts or the CLI and test it and you know and it doesn't work and the last agent did it and it worked but it discarded it. So it was just very obvious at that point like I should just turn this into a CLI and put that inside of the skill uh so that every agent that uses it now benefits from that same piece.
Um, but I I also think like, you know, it's a good push for people to think about is how much of your skills and rules could actually be deterministic. Um, that's like another core thing or or one of my core principles that I like to think about is yeah, how do I uh make very efficient use of determinism and non-determinism and, you know, let Asians shine at the non-deterministic parts, right? because that's what they're trained to do.
Um, and the other parts which are much more mechanical or, you know, straightforward can be just pure determinism. Um, and you'll see this as well for things like doing migrations. Um, which is another big thing that I've I've talked about is, you know, going from one technology to another, especially one that is better for agents, right? And a lot of how you can do that migration is I think through things like scripts and CLIs like the deterministic parts like code mods you know like crawling the astra syntax tree and transforming code literally mechanically right like a script does it for you instead of the agent totally makes sense.
I I mean I think what there's another thing there which is you're taking stuff away from the agent and you're kind of putting it in the environment too a little bit which is let's say you have a a thing that you notice the agent always gets wrong you want to make that um just impossible within the environment and that sort of comes down to code quality as well I mean I talk about a lot like having a what a good codebase means, right?
What is a good codebase? And there's a definition I like which is a a good codebase is a codebase that's easy to make changes in, right? Easy to um change stuff without things screwing up. And that means that you have a lot of guard rails that you have a lot of um the agent or the human is constrained to very narrow paths. And that's again something you talk about in your talk >> and you talk about this not only on the kind of >> sort of automated checks side of things.
So linting and time checking blah blah blah but also in the way you design abstractions and you guys even I think built a framework uh for your agent to work in too. I think what I'd love to hear is you obviously think of that as very important right and that's how important is that compared to other things you could be doing like building features or shipping work. Yeah, I think that's um I almost feel like the new job of the engineer is really to to spend time on the environment.
Um I almost actually wrote a tweet about this yesterday, but I but I didn't. But I think that I think that you you know if you if you haven't really spent time, you know, building trust in your agents and building skills and tools, you can get stuck in this mode where you're very low on that trust ladder, right? You don't have a lot of trust in your agents work. And so the only way to cope in that when you're in that situation is just to kind of lock in and micromanage your agents.
And that's very time consuming. And when you're stuck in that mode, you don't really have the luxury to think about, you know, uh higher level things, right? Like like making yourself more productive. In the same way that uh I guess analogy would be like if you've never taken the time to learn like your tools right as a developer when you were writing code yourself and you know you had you've never heard of VS Code, you've never heard of Vim, you only knew about Notepad uh and you had hadn't even heard about Git.
That's sort of the analogy. It's like you you haven't spent the time sharpening your own knives, right? And so, of course, if you have a dull knife, then everything's going to take a long time. Um, and you're going to you're just going to be and and especially if you know deadlines are looming, then you don't have the now you're stuck in this rut, right? Where where you you you don't have sharp knives, you don't have good tools, but you're under all this pressure to ship, right?
And so, all you can do is just focus on that. But I do think that, you know, if you can find yourself the time to actually spend time thinking about your setup, it's again going back to the cooking, you know, analogy. It's like uh, you know, if you, for example, if if cutting cutting the garlic is like super slow, right? There are garlic mashers, right? You can you can buy and you put it in the thing and you like squeeze it out, right?
It's super fast. Uh, machines and tools were invented for a reason, right? And so if you're operating a Michelin kitchen and your your your cooks had no tools, then of course everything is going to be extremely inefficient, very very, you know, every every every cook is going to make something up of their own. So I think the tools and the determinism to me are, you know, taking that part away and and just like you said about constraints as well.
It's the constraints are are to me as well like uh actually a slight tangent on that is uh I think we should talk about TypeScript because like we we actually both share like a background in Typescript where you know you obviously have done a lot of work with TypeScript and total TypeScript and you know you're a leader in that space and I uh had adopted TypeScript pretty early and I had given like a talk or two at Typescript conf uh many years ago and so one of the talk that I did actually was about type systems and constraining the constraining types.
Like one of my most favorite things about Typescript is actually type narrowing, right? This idea that you go from a very broad type, all right, that could be anything and then you through type guards and you know type narrowing and you know runtime checks, you can actually narrow the space and say like oh this isn't just a string this is a very special type of string. it's a constant, right? Like I but I I determine that through the type system and in a way it's like uh there's a lot of parallels I think to that with constraints in your codebase where is you're you're constraining the space, right?
If you if you think about category theory as well, you know, you're constraining the the number of possible types, right, that can can exist and you're saying there's only one type, right? And for for us like that framework that I'm building called Dune uh uh it's not an open source framework. It's the the way I describe it to people. It's it's kind of like a our internal Nex.js for our Electron apps. Uh but it comes with a lot of really really restrictive lit rules and the codebase is designed in a way that there's really only one way to do something.
So we make use a we make use of a lot of conventional patterns. So like features all go into a specific directory. Well, every feature has its own directory. As an example, you know, there's like a a thing that discovers features like through a registry and like crawling the codebase and stuff like that. But this conventional pattern and the lint rules make for an environment where it's actually very hard to write bad code.
And that sort of frees up the it both frees up your own mental uh you know capacity as well as the agent sort of doesn't have to think about that anymore where it's just like oh there's only there's I should just if I want to add a new feature it just goes in the feature the new feature directory and all the code goes in there and I'm not going to append to a god file right that was really actually the inspiration for those feature directories is the very first couple of versions of Grockbot were composed of like eight god files which were like at least 10,000 lines long if not longer and so I kind of had to break it up into smaller pieces uh but it was just observing you know actually that's another important part is observing how agents fail and then every time you see a mistake every time you see something that could be done better you think you step back and think how do I turn this into a lint rule how do I make it so that the codebase makes this impossible, right?
And it comes back to me for my, you know, my background learning TypeScript and types uh type systems is how do I constrain the space so that you know I know precisely what I'm working with and I think yeah there's a lot of parallels there. >> Totally makes sense. And Denny, I mean, it's funny that you mentioned Typescript and God files in the same sentence because Typescript famously has a 25,000line type god file. >> Although I don't know if they've rewritten that and go, they probably have, they um Okay, so environment is important.
You should watch your agent like a hawk to make sure that any mistakes it makes, you turn them into things in the environment. And the benefit of the environment is you're not overloading your agent, right, in terms of rules, in terms of things it has to remember. It's just in the environment and so it stumbles into the rules and exactly um >> you know bounces off them and hits them at the right moment. >> So, okay, we still haven't talked about the 2,500 PRs.
Where do those come from? Like how do you you've built your trust ladder, you've worked on your environment, and you understand, okay, um I now want to scale up. So what are the mechanics of that scaling? Are you um initiating 2500 like chats per month? That can't be right. So there must be are there any kind of automated triggers that trigger stuff in your repo? Like how do you get the software factory kind of triggering work by itself? >> Right.
Um, I'll definitely say that the prerequisite to, you know, something like a very high volume of of pull requests, um, is the environment, you know, the the stuff we just talked about where I definitely would not have been able to do this if I had not spent the time, you know, thinking about the kitchen, right, and the knives and the tools for my agents. And so in a way I I think of this as I've spent the time building one kitchen and one restaurant and now I I'm in a position where I don't actually have to be there anymore because the environment you know that the same analogy right of restaurants right >> exactly yeah exactly it's like you're Gordon Ramsay and you know you you've taught your executive chef like all the tricks of the of coming up with great menu.
Uh, and like the kitchen is set up really well. Everything's just perfect. And you're now in a position where you can open your second your third restaurant. And I guess I I sort of see each project that I work on, like each big chat is sort of like a restaurant, right? And and I'm I'm I have multiple of them operating at the same time. And I'm sort of like helicoptering between them. Sometimes some more than others depending on how in the loop I am.
But yeah, definitely I think there's there there are external triggers and context that those projects don't have that for a long time I was the proxy for that. So, uh, the best example I have is like, you know, you have a project that's working on a feature, uh, or you're trying to fix a bug and you're getting bug reports, but the bug reports are going to things like Slack or Linear or X, right? And these are external systems that aren't connected to your inner loop.
So, I like to talk about this outer loop and the inner loop. Uh I don't know if I'm using the definition correctly but to me my inner loop is like basically my engineers my agent engineers working on the code to an building towards an intent or snapshot of my intent right and the thing about that is that the snapshot can go stale right new information comes to light that I then have to be the proxy of and you know transfer that context to my agent.
So, you know, if you if you don't have these triggers pulling information back into the interloop, then you sort of have to play that role where you're you're off, you know, in Slack or X or or whatever and you're gathering context, right? You're getting context about bug reports, about feature requests, about, you know, something someone said about, you know, our backend infrastructure has some limitation. you know, all that information you have to f that across to your agent.
So that's where I think like tools like Grockbot are really good because they help you automate the outer loop as well. And when you connect those two loops, it's very very powerful because now all of a sudden your agents have the ability to get context for their for themselves, right? If for example uh you know either through just as a simple example like maybe you have the Slack MCP right or you have uh your own harness right that you've built a Slack subscription into for a particular Slack channel.
Now all of a sudden you can tell your agents okay subscribe to the Slack channel. every time there's a uh you know bug report about something go off and triage that thing right go reproduce the issue right using the verification skills that we've already spent time building and all of those other skills that we've set up so that I have a lot of trust right I have a lot of trust that these agents can actually go off and understand the bug you know uh verify that the bug actually still exists on main and it wasn't something about you maybe the user's setup or their data or maybe I don't know they didn't install a dependency or something like that like uh basically I think uh creating that yeah creating those two loops and connecting them is really a very important part of the job these days um especially if you are thinking about how to scale yourself.
So, a big theme here is really just like always thinking about like what where am I the bottleneck in this process? Why do my agents need me, you know, to answer this question? I I always like to think about that. And so I try to think about how do I actually get the agent to answer its own question, right? But not by hallucinating, not by guessing, but actually real data. And you know, a lot of people talk about this idea of a company brain, right?
Or a context graph. I feel like those terms are unnecessarily complex uh or even abstract. To me, it's just about um how do I take information that my agent needs that I would otherwise have to go and pass it myself and just teach it how to do it, right? And that removes me from the equation. And so, how I arrive at 2,00 or however many PRs is the fact that I have all these loops set up, right? And so, uh, it allows me to open chain restaurants, right?
I can I can really parallelize myself. So, yeah, I'm not sitting there creating 2,500 chats, right? Of course, it's really like these projects are um actually cursor has a new feature called projects, which are these like coordinator agents. Um, and so the coordination co coordinator agents are really good at sort of delegating and not doing work of their own, but they manage and supervise like almost a list of tasks and they spawn sub agents to go and do them.
And so I'm just constantly feeding context or teaching the agents how to get their own context and then they're going off and doing the work for me. Uh, and really the big the last thing I'll say to this is like the big unlock for me for getting to 2,000 PRs is starting from the question and working backwards of how do I get to the point where my agent can merge its own code? Because the obvious thing people ask me when when they when I tell them, "Oh, I shipped 2,000 and 2,500 pull requests last month." They'll be like, "How did you review that?" Right?
That that's a lot of PRs to review. Like your team must hate you. Do do you mind if we go there in a second because a good question about that? >> Yeah. Yeah. Yeah. >> I want to like this analogy is great. I want to like deepen it a bit which is before if you're like manually initiating all those chats it's like you're bringing the orders to your chefs manually, right? Whereas if you've got an agent sort of like doing the expo then you're able to sort of run it yourself. >> What is what does that concretely look like then?
You've got these sort of Grock bots that are um subscribing to channels, pulling in Slack messages, and you, it sounds like, have a couple of coordinator agents or like chief of staff agents that like monitor that or something like when you look at your computer to manage your agents. What does it look like? >> Yeah. So, so uh this is I guess somewhat confusing but we're working on you know simplifying and unifying but so uh there's graphbot uh which or you know you can use other tools of course as well but I I largely think of these tools as like your outer loop these are tools like you know graphbot that have connectors right these are connectors I guess they a lot of people call them personal agents um but they're connectors to things like your email your calendar slack uh plaid I don't know like all these different services and they are a great source of pulling context in to your work.
So the same way that a human like you know if I were if I was a manager and I was leading a team of engineers um you know like when I used to work at Netflix one of the biggest things that managers would talk about was this idea of context not control which funnily enough you know has so much uh has so much uh carry over to the agents world uh of you know you you know you you of course can drive to an outcome you want by control right by by micromanaging, but what you want is to provide context instead, right?
Like teach the agent, teach your engineers how to be self-sufficient and then you don't have to micromanage them. >> Um, and so I see a lot of parallels there. Uh, but yeah, graphbot, so concretely, I have some grabbots that look at my Slack channels, look at my X, uh, or my emails, uh, or or linear, and they're just constantly they have routines that subscribe. So they're constantly watching and I have I I'll tell them things like you know uh I'll watch for issues with uh bugs in the graphbot desktop app as an example.
Uh and whenever you find that send it to my cursor project. So one of the really cool things about grabbot is it connects to cursor. So cursor has uh like I I just mentioned this new feature called projects. And a project is really a uh again like a you get a coordinator agent that's in the cloud. It has its own computer and all it really does is like it's a manager of agents. It's like your executive chef, right? Your your chief of staff.
It doesn't do the work itself. It delegates and orchestrates and manages the work of other sub agents to you know that report to your chief your chief uh of staff. And it basically is responsible for driving the work forward and managing things and uh passing context to them. >> So if you get a sudden burst of issues, let's say you get 30 issues at once in one payload or something or very quickly the coordinator agent can figure it out and delegate. >> Yeah, exactly.
It gets like uh you know 30 the 30 or so payloads and spawns a sub agent or a single coordinator agent. It can actually do a bunch of different topologies of agents and it will sort of figure out the best way to uh you know efficiently distribute the tasks to your team of agents. Um so I use uh cursor projects a lot um and I also use grapot a lot and cursor projects are my inner loop and grabbot is my outer loop. Grabbot takes all the context, external context, gives it to the projects because it can actually just send messages to those projects, right?
You don't even have to open cursor. You can just tell your grabbot, okay, create a project, right, for these series of tasks. They're all related, right? Maybe as an example, you know, you've had a uh a big burst of issues that are all about performance, right? Your app is slow uh and they're all connected, right? Maybe some of them even have a similar fix, right? But and you can certainly go off and just spawn one agent per task, but then you've lost that sort of thread between them, right?
And and you may duplicate work or you may not really think about the higher level problem. You know, sometimes when you you you solve bugs, you know, it helps to have multiple bug reports that are are slightly different because it helps you really, you know, zoom out and see actually, you know, the problem when I looked at this one report, I thought the bug was here, but actually when when I see the other multitude of bugs is actually up here, >> right? >> Yeah.
Got you. So that that's why you have so many agents in that loop then, right? Because it's not just you have um like you have a bug report comes in, you spawn a single agent to look at that bug report. that a that single agent will be duplicating work with other um other agents, right? Because if there are multiple bug reports coming in through the same thing, they're going to be duplicated work. >> Yeah. >> Fascinating. >> That's really fascinating.
Okay. And so this just this endless series of triggers um coming from real users reporting real reports um builds up this sort of and accelerates the factory sort of adds more orders in. Other than bug reports, are there any other sources that you use for like um accelerating for pushing these PRs? >> Uh well, funnily enough, it's some of it comes from uh reading the code too. So, I guess I have sort of uh well, so to clarify that, you know, the 2,500 PRs, they're not obviously like 2,500 features, right? they are a lot of the work actually is spent on gardening like another term that I really love.
Uh where so I guess this is more important when you have a big team of engineers human engineers that you work with where and also this goes back a little bit to what I was talking about with the environment. You know, setting up a really good environment that doesn't just help you and your agents, but everybody on your team, right? Think of a new hire who doesn't have a lot of context on all of your engineering practices joining your your team.
And if you have a really good environment, they can be productive from day one, right? They can they don't have to like, you know, make open a bunch of lowquality PRs. They can start, you know, they can start just turning out really good code. Um, and uh, yeah, I think I I sort of lost my train of thought. >> I've got a I've got a followup, which is what's what are the mechanics of like triggering >> like how when when do you trigger a a agent to go and look at the code, right?
Because some people might say, "Oh, let's just do that every hour or something or like on a chron job or what's >> Oh, yeah, yeah, yeah. Yeah. Uh I saw some of your recent tweets as well about like you know the some of the tweets you've been doing which are great for setting up your routines. Uh I have some routines like that as well. Um so uh one of them is uh like looking through just another simple example is you know React has a lot of foot guns.
Um, so, uh, as as I'm sure you're aware. And so I have an agent that's just constantly looking for band patterns. And the interesting thing about that one is that I don't actually tell it to fix the issue first. I tell it to append it to a document. And then every couple of days I look at it and I see actually these are all the same thing, you know, and so that gives me, you know, you almost want like a buffer, a cue.
Sometimes that's actually more effective than just spawning off a couple of like a lot of sub agents to fix every single thing because when you are in kind of pure execution mode and just trying to like you know f uh you know execute on the orders that are coming in very fast you sometimes miss the big picture. So sometimes having a buffer forces you to think about the big picture because you you you have these artifacts and things that you can look at as a human um and sort of use your own human judgment to or I guess you can use an agent to do that as well.
But you give the agent and yourself a way to identify patterns, right, that you might otherwise miss if you're just only solving each bug at a time. And that's also really the benefit of having something like a chief of staff agent is uh it can see the forest right uh in addition to actually doing the execution. >> Fascinating. That's I mean my brain is exploding a bit there with the sort of chief of staff at the software factory.
I think I might have to change some of the course that I'm filming next week. Uh, all right. Let's talk about let's talk about review, right? Because this is the reply that you get, you know, is >> did you read did you taste all 2500 of those dishes as they swept past you? >> And I assume the answer is a variety is a version of no. >> Yeah, I think you you you don't want to be in a position where you're not tasting the food ever again. uh but you also you know for scale you cannot be tasting every single dish that comes out of your kitchen especially if you have multiple restaurants.
So it's it becomes more about sampling right and thinking about the processes in the same way that you know if I guess maybe this is where the the the factory analogy is a bit more apt is you know as a quality supervisor on a factory you can't look at every single item you sample right you take you you you go in there every day and you look at the quality of the pull requests you look at the code that the agents are writing and you scrutinize it very rigorously and you think about all the inefficiencies, the bad patterns that the agents are doing and then you think about how to course correct the environment, right?
Not not that single agent. Uh because if maybe if it if it was a one-off incident, it's fine. You know that maybe there's nothing to fix there. But if you actually notice that multiple agents are are having the same issue, right? They're taking the same shortcut. they're they're propagating the same workaround everywhere. Uh that's a sign that you should go off and think about how to uh amend your kitchen or your factory, right?
Like thinking about your skills, your constraints, your lints, your type systems um and setting or adjusting it so that that problem doesn't happen again. And when you do that enough times, then you get to a place where the codebase is again like the environment is so constrained and so it guides you so well that you can just you can just step away, right? That's the dream. And I I'll definitely say it um it's very hard to get to this point.
I don't want to sell this as like you know something that you can just do easily by using PAC like it takes a lot of time and effort to think about your code and where you see your agents failing and thinking very thoughtfully intentionally and setting out guard rails and constraints so that they do the right thing by default >> and you're not like if to go back to the software factory analogy this isn't a dark factory right this the Lights are on, right? >> It kind of is.
Yeah, actually. >> Is it? >> Yeah. Well, it's dark in the sense that so um it's dark in the sense that Well, I think if my agents are merging their own pull requests, it's sort of become dark where I go to sleep. My agents now work 24/7. Uh I have I have the equivalent of like more than 10 chiefs of staff, right? Each working on a different area. Like for example, I have one that's working on performance of the Grockbot desktop app.
I have one that's working on uh fixing bugs that users report. I have one that's exploring rewriting it in a different language just for fun, you know, like what if what if you know, just reimagining what what it would be if it was like a native app. It's just a toy. Um but the idea is like yeah, I uh I when you spend the time setting up your environment, I've gotten to a point where I review the pull request after it's landed, right?
I I tell my agents full autopilot is is something that you can do in in PAC and that will trigger off this very intense rigorous verification loop where it'll spawn a bunch of verifier agents for every pull request and it will fuzz right fuzzing meaning that it will actually run the application. It's going to click around and try to use it like a real human. look for regressions, look for bugs in your implementation and um it will try to find issues with the thing and then it will fix it itself.
It'll do that again and eventually get the PR to a state where it can land. Uh so it does it does it is quite token intensive. You can tune this of course. Uh so you know instead of like 10 verifier agents you might do like one right or you just tell the agent to verify it doesn't work. But yeah, the key thing is the verification part is really the key piece that gives me a lot of confidence that I guess verification plus the environment, right?
It's these the combination of these two things that allow me to step away and say agents go off and merge your thing. I'll review it in the morning by looking at my commit history >> and if I see problems, I go and course correct, >> right? And I'll go and revert or modify, add new lint rules and whatever. Um, so it does it does take time to get to that point, but once you get it, oh, it's so it feels so magical. Uh, I I I tell people like I'm sleeping so much better now because, you know, it took it the very first day I turned on the sort of dark factory.
It was very scary because I was like, "Ooh, what if I call the SE, right? What if I break something overnight?" >> Uh, and it took a lot of >> it took a lot of uh bravery, I think, to do that, but >> somehow I did it. And yeah, now I'm in a place where my agents are are merging their own code while I >> It sounds like >> I think it's dark in that sense. >> Yes, it's dark sometimes, right? You do sometimes. >> That's true.
That's true. >> Because I think of a dark factory is like almost like if you take the original definition of Kapathy's vibe coding, right, which is the code almost doesn't exist. You forget the code might be a thing. I think your approach is totally different from that, which is that code and the environment is essential. And if the code in the environment are bad, then you will get bad outputs. Garbage in, garbage out.
So I I I think this is a this is a different thing. It's like, you know, the I don't know, maybe there's a dimmer switch or something, right? Like, you know, some parts of dark, some parts were light. >> This is why the maybe the the kitchen is a better analogy. Yeah, >> the restaurant because, you know, even as a as a as a as a restaurant restaurant, you still might go to your restaurants every now and then to take take a peek in, taste the food, right? >> Yeah. >> Uh I think that's >> the idea of sampling instead of blocking, I think, is really important. >> Yeah. >> I think what would you say to people who are in I guess you're obviously in a pretty security conscious environment where you're working, very security conscious. >> Mhm.
Um maybe there are folks working in like um medical applications or law or finance or something. I think of it like some PRs are kind of like two-way doors which is you can merge it and then revert it, right? It's cheap go back through. But there are some PRs that are one-way doors, right? That will cause data loss of some kind that will do something that can't be easily walked back. How do you deal with situations where most of your PRs, let's say, are one-way doors?
Like, is this something you just wouldn't recommend or like what do you think? >> Yeah, I think that's a really good question. I think that it all comes back to me to the quality of the verification that you're able to um get out of your agents. And I think for domains where the work is verifiable, this is easier, right? and the the oneway doors become two-way doors in a sense. But I guess I don't know if you're working on something that is like extremely is very hard to verify programmatically then I think yeah you're definitely in a position where it's very hard to get to that point.
Um so I do think like yeah verifiability of the domain is an important aspect to be able to do this. Um and software engineering is just one of those things where it's quite verifiable in in a lot of cases maybe not totally um you know like other domains like mathematics I think are another example of not all of it of course but some aspects of mathematics can be verifiable if you write a proof for example um and so yeah I think it's a great question that I don't really have the answer to and I think that this is something the industry and us as engineers will have to figure out is you know my sort of uh hope and prediction for the future is that we'll see more and more interesting new agentoriented programming languages and one of the most fascinating ones that I've seen so far is this one called bend bend d bend um and that language is one where it kind of marries programming with proofs right there used to be a time, you know, where you actually had to write your proofs in a different language.
And proofs, by the way, for those uh who who aren't familiar is this idea of uh that you can sort of formally verify that some code is correct mathematically, right? Especially if you've written your code in a very functional programming way. uh but for the longest time you had to do that in a separate language like lean or ta plus or uh I'm blanking on some of the other other examples but uh like languages like that where you would construct the mathematical proof and then use a solver essentially to det that that you've covered all the cases you don't have like a race condition or or whatever.
So yeah, I think trying to sum up the question, I think yeah, if you are in a position where you can figure out how your agents can truly verify the work in a way that gives you confidence, you can actually, you know, uh have the PRs merge cuz if it compiles, right, if it if the proofs show you that it's correct, then why wouldn't you just merge it? Um but of course, yeah, not all domains are verifiable. Yeah, it's a tough one.
Um, okay. I think we've got to think about wrapping up because we are nearly on the hour. Have you Have you got something after this? I mean, I've got something before I give my son dinner, but >> I I can go a bit longer after you. >> Okay, let's let's go five minutes longer then. Um I think I just want to have one more question which is I think I want to ask how you see Pstack and how you see skills in general like in terms of we talked about this before we went on air which is like people think of as like my skills versus your skills and how do you combine frameworks together?
How do you use Pstack with my stuff? like what should you take from each one? And because I think I see skills as sort of just derived from process basically like they're just processes turned into words and I would love to know how you recommend people take Pstack and take my stuff as well and turn it into their own processes. I think you shared a tip actually today that I thought was actually very relevant which is this idea that you go off and look at your previous transcripts right and you sort of mine for information of you know your own look through your own your own prompts right to the agents where you correct them where you have to constantly intervene and uh you know take that higher level learning and turn that into a a reusable skill right so that agents stop repeating that mistake I think that uh PAC and your skills are very complimentary.
I I totally agree with you that they're like a skill is really much just process. I mean it's just at the end of the day a skill is just English or or language. It's just markdown. Um and I think you can you can definitely weave them combine them in a way that makes sense to you. But I do think that uh everyone should have their own set of knives, right? I keep going back to the the the cooking analogy, but it's so apt because like you know every chef when they go to a different job, right?
When they go to a different restaurant, they carry they bring their knives with them. The tools go with them, right? And so trust to me is really about trust in your own tools. And when you spend the time sharpening them and understanding them really, really well, you can do great things. And everybody's skills and tool set is going to look different. You know, someone might find a lot of success combining, you know, like your grill me with dogs uh or wayfinder skill with some of the execution skills in Pstack as an example.
Some people might use more of your skills. Some people might use more of my skills. I think at at end of the day, it really just comes back to how much do you trust, you know, me and Matt, right? Okay, if you if you trust us both, of course, use our skills, but I also encourage you to, you know, look at your own transcripts. Um um tell your agent to look through, you know, some of the all of the patterns that you've used, the the times you've had to intervene, you know, suggest turning them into lint rules or new skills, right? the the past chats I I often say is like a a treasure trove of context because the you know it's like the process materialized right like it's the real process not an abstract idea in your head it's the actual thing right and you can actually see how it happened in practice and extract so much information from that and there's so much so that I actually turn I have a skill in pac called recall which is exactly that um where this was a pattern where you know I was working in I was working on a similar problem.
So specifically I was working on virtualization for the cursor application and there were a lot of bugs and so you know every time I started a new chat I was like ah this there's so much good context from the last one though you know I want to bring it over to the new chat how do I do that and that's where the transcript came uh you know looking at the past transcript came about and then recall was just a way for me to collapse collapse and compress that workflow into a skill. so that I didn't have to just say I didn't have to write a long essay every time go look at all these chats right and you know blah blah blah so I I largely think of skills especially as agents get more capable as really encoding workflows you know skills from last year were really more about like almost like implementation details like here are the exact script commands you know you should use right I think with the latest models you can just delete those parts and just really focus on the workflow and tell it it's more like a skill becomes more like a series of steps, a series of real process.
Uh, and I think over time we'll see that skills get smaller and smaller, you know, more compact. Um, and yeah, they're very compatible. Or you can, you know, if you want, why not read our skills, right, and com and combine them in your of your own, right? Combine Wfinder with potato mode and make your own custom mode, right? Like skills are the the thing I love about skills are that is are that they're so malleable. You can do anything you want.
It's just language. >> Absolutely. There's nothing magical in them, right? They're just words. And >> Exactly. >> If if there is any magic in them, it's just the words chosen and the phrases used and the thinking that's been done to turn those like take abstract process and turn them into language. And once that thinking has been done, then it's just there. It's available. that's on the surface and you just nick it. Um, Lauren, thank you so much.
This has been glorious. >> Yeah, this has been super fun. I really enjoyed talking to you. Hopefully, we can do it again. >> I'd love to do it again. >> I'd love to do it again. Absolutely. Um, yeah, we'll check in in uh >> through Yeah. Monday. Let's do it. >> Let's do it. >> Part two. >> Well, thanks so much. I'm going to close the stream here. Laura and I will uh chat a little bit and stay here. But thank you guys so much for watching.
The glorious.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.