Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 11:58
4.6x that video's typical replay level
issues. Uh I also invented OS certification. I just close the tracker whenever I want, so I have my life back. So, does this work? Yes, sort of. >> [laughter] >> Which leads me to act three, slow the down. Everything's broken.
Said at 11:52
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
The graph counts replays. It does not show where viewers stopped watching.
Words
2,596
Runtime
19:33
Speaking pace
133wpm
Reading time
11min
133 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] A friend told me about an engineer at another company who ran a recurring war room just to delete dead experiments. Every two weeks, the team would stop, gather, and clean up flags that everyone agreed should be gone. That's the bestase version of this problem. Someone burning political capital to make maintenance happen by hand. That's the part that I still can't get past. My
67 words, the words spoken in the first 30 seconds at 133 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 261 |
| Average words per sentence | 9.9 |
| Longest sentence | 37 words |
| Questions asked | 14 |
| Sentences containing a number | 9 |
Most used terms
Filler phrases
11 in total: actually 6 · like 3 · kind of 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] A friend told me about an engineer at another company who ran a recurring war room just to delete dead experiments. Every two weeks, the team would stop, gather, and clean up flags that everyone agreed should be gone. That's the bestase version of this problem. Someone burning political capital to make maintenance happen by hand. That's the part that I still can't get past. My name is Andrew Robbittor and I'm an Android engineer at Reddit.
This is a talk about what's actually stopping us as an industry from scaling that work responsibly. Spin at the gate until green. the engineering primitives behind self-driving code bases. First, what I'm not going to do, I'm not going to stand up here and tell you that AI can write code. You're at the AI engineer world's fair. Half of you are running coding agents right now. So, let's skip the part where I sell you on AI writing code because you laugh me out of the room.
For me, my only input is the chat box. I describe it builds. I review. Your mileage may vary. The point was never how the code gets written. It's about whether you ship good software, a product that people can actually trust. And every gate in that system is a human's call. I've watched this everywhere. You've got your own version of it. the war room that I opened with someone burning a week of hackathon freedom on lint cleanup.
As an industry, we've built tooling that can change the shape of the work, and we still spend our scarcest resource on work like this by hand. Human creativity, human attention wasted. The job changed and most of us didn't even notice. We're harness engineers now. Our work is the system that helps produce and verify the code, the constraints, the gates, the skills, the verification. Think about how you got good for a second.
Mentorship, code review, the water cooler chats, the incident that paged you at 3:00 a.m. You absorbed all of it without noticing, just by being there. Humans absorb judgment implicitly. Agents require judgment explicitly. An agent boots with a blank context window and no memory every session. Every piece of judgment you picked up by osmosis, it has to be made explicit or the agent simply never has it. Quick test. Swap in a smarter model.
You get a slightly better answer. Now take away the tests, the gates, the review and the whole thing falls over. So the model was never really the bottleneck. The judgment around the model is the model is raw talent while the judgment is the organization. And this judgment already exists. It lives in scars, in post-mortems, in your most experienced reviewers. it just doesn't scale because it's trapped in one head. So step one is making it accessible to the whole org where one reviewer's taste can guide a thousand changes they'll never touch.
Ask experienced engineers to clean up a feature flag and they they're going to think about a whole string of questions first. A flag is never just a flag. Is the rollout frozen? What about the flag right next to it? Which team owns it? That bundle of questions is judgment and it's the most valuable thing in the building. Right now, it's trapped. It surfaces once in a code review, then it's gone, and the next person relearns it the hard way.
A skill is institutional judgment made executable. It's that bundle written down so that the agent refires it instead of guessing. A newer engineer can work with 10-year instincts at their elbow. In 1986, Marvin Minsky named this a knowledge line or a kine for short. It's the configuration of mind that solved a problem reused on the next one. So, a skill is a kine for a codebase. And skills aren't documentation. Agents drown in facts.
What they lack is which facts matter and which decisions are dangerous. Documentation preserves facts while skills preserve judgment. Skills are judgment that you reuse. But often times work such as feature development or a refactor outlives a single session and the context window wipes clean every time. So you need the other half. A record of the work itself. The plan, the decisions, what you already tried, even the surprises along the way.
That's a worklog. A plan is a prediction. A worklog is a record that starts with one. So new session, fresh agent, zero memory. You type one word, continue. It reads the work log and picks up at milestone 7 out of nine. No re-exaining and the work log is the context. This talk is living proof. I built it across a couple of sessions using a worklog. I came back to a half-finish presentation and a fresh agent read the work log and resumed exactly where I had left off.
That's the only reason why you're seeing a finished talk and not me apologizing for one. And for my personal side projects, this isn't willpower. A git hook blocks any commit that doesn't update the worklog. This way, the memory keeps itself. Skills encode how to do something. Personas encode whose eyes look at it. I got this from a Carpathy tweet. There's no you when you prompt a model. It doesn't have an opinion to give.
So don't ask it for one, ask for a perspective instead. What would be a good group of people to explore XYZ? What would they say? [snorts] On my personal side projects, it's just me. No security team, no designer, no one to bounce ideas off of. So I have the model review my code as a security lead as a UX reacher as a UX researcher as Mchavelli when I want to know something when I want to know how something gets abused.
I even built a design panel of opposing philosophies my designs have to survive. Then I read the reports and act on them. Each persona loads a kind of judgment that I don't personally have on demand. Encode a domain's taste once and anyone can borrow eyes that they don't have. The engineer with no designer gets design judgment. The newer engineer gets a security reviewer's instincts. Push it to the whole org and every team can now encode its own taste for its own agents.
The judgment that used to live in a handful of heads is suddenly everywhere at once. skills, work logs, personas, judgment out of our heads and into the repo, portable to a machine and to people that don't have it. But the second that this judgment lives in three different places, you've got a governance question. How does one team's judgment reach everyone else? Simpler than it sounds. Every gate that you already trust, whether it's CI or types or lint or the design system, is just institutional judgment pulled out of someone's head and made mechanical.
We just never called it that. So, you don't mandate it. You make it legible and you let it spread. All right. So, now your judgment exists. It's encoded and it's governed. Great. But none of it matters if you can't actually trust what the agent did. So what actually separates a tool that helps you write code from one you can trust to run on its own? Verification. Verification is all that stands between you and software that you can trust.
Everything before this slide is knowledge. Everything after this slide is trust. And trust is a human's call twice. Once on how it's built, once on the proof that it works. Verification is a ladder. Builds and tests. You've got those screenshot tests with a model reasoning over them can catch contrast and overlap a human a human ski. Then video the feature actually running. than telemetry in production. The more rungs you can run without a human, the more you can hand off.
You earn autonomy one rung at a time. And the agent doesn't have to be right the first time. It just has to know when it's wrong and try again. Generate, test, fail, regenerate. Nobody yolos code to prod and neither should your agents. We all check our work. spin at the gate until it comes back green. I do this on my side projects. I give the agent its own QA. It maps the feature, drives the app like a real user, and records itself walking every flow.
The agent that wrote the code will always tell you that it works. So, you make it hand back a recording of the thing actually running. Producing that recording forces it to make the thing work. You can't fake a passing run. And the recording is real proof. I watch it today. Down the line, you could have an LLM watch it first and send it back. Fix this. Redo that. But a human still makes the final call to merge. The requirement forces the work.
The recording proves it. I got tired of watching old feature flags pile up. So I built an agent for cleaning up stale feature flags. It runs in my local workflow and it doesn't merge any code. The models the easy part. The judgment in front of it is the whole game. Every flag gets scored the way an experienced reviewer would analyze it. How many modules does it touch? Is it multivariant? Does it share a component? Then it checks the experiment data.
Is the roll out frozen? Is there a sample ratio mismatch? Is there a variant already rolled out to 100%. Only the safe mechanical cleanups ever reach the model. Before I trusted this workflow, I back tested the scoring against months of cleanup history. Then I ran it live. Seven for seven PRs with green CI. Now it runs daily on my laptop and hands work back for me to review. $1.26 a pull request, a backlog of around 520 flags a year, runs for under $700.
If you had engineers working on this by hand, at least $26,000 that you're looking at. This is the work that nobody was ever going to do. I've seen feature flags in the codebase that are three, four years old, and some are even older. The model wrote the code and I wrote the judgment. One hard lesson from my side projects. When you let an agent author the code, your guard rails are your code review. A machine doesn't feel a polite comment.
It only respects a hard gate. I built a pre-commit hook so my agents couldn't write to main. Then I asked Codeex, would the hook even stop it? It told me flat out repo hooks are not sufficient protection against me. Its patch tool writes underneath the hook. The agent told me that my gate was worthless. So I moved the gate down to the OS level. Later I asked it to valid I asked it to list the valid reasons to unlock and it quietly slipped emergency recovery into the allow list.
A self authorizing exception that nobody asked for. It even admitted it. CODEC said that's the kind of escape hatch that starts as safety and ends as self- authorized nonsense. The pit of success assumes people take the easy path. Agents do not. They will build ladders to climb out of the pit of success. They find every escape hatch that you leave. And if you leave none, they'll invent one. So gate the choke point. Make the bypass operator only.
Never hand the agent a reason that it can grant itself. A gate with an escape hatch isn't a gate. Once judgment is verifiable, it's executable. A flag agent, a dependency agent, an accessibility agent. Each one is narrow. Each with its own encoded judgment pointed at a different chore. Remember Minsky from the Kline? That was just a footnote in his in his idea. He called it the society of mind. A swarm of small, simple specialists that compose into something that looks effortless at the top.
No master brain anywhere. That's the fleet. specialists, each spinning at its own gate, and a self-driving code base emerges from the swarm. And that flag agent, that's the war room from the opening, turned into a specialist. The work somebody organized a whole meeting around is just one member of this agent society. Every one of them doing the same thing, spinning at the gate until green. Three stages. In the loop, you prompt.
On the loop, you orchestrate. You review the outcomes. Off the loop, a trigger fires it. Craw an event. No human has to kick off that first draft, but humans still own the approval and the merge process. The unlock there is a cloud runtime the model vendor may or may not ship for you. Everything that I've shown you, skills, personas, the gates, you wrote it down once, but the codebase doesn't hold still. It moves every single day.
A skill written for last month's architecture isn't out of date. It's wrong. And that's worse than nothing because the agents trust it. Encoded judgment rots. If the skills don't move with the code, then the whole thing falls over. So you keep it alive in two ways. When an agent fails, you don't just fix the bug. You run the postmortem and fold the lesson back into the skill so that failure can't happen twice. Failure becomes a constraint.
But nobody is hand auditing 300 skills for drift. So, you give the agent a bedtime, a scheduled pass where it reads its own skills, finds the stale ones and the contradictions, and opens up drafts for review. The system sleeps, cleans house, and wakes up sharper, but that's a whole separate talk on its own. A self-driving codebase is a system that keeps its own judgment current in addition to writing code. Everything that I've shown so far fits somewhere along this judgment pipeline.
And every mechanism in this talk, skills, personas, work logs, and the gates, the fleet, all of it was doing one job. Managing judgment. Every session, the model wakes up like Drew Barrymore in 51st Dates. Brilliant. and no memory of yesterday. No idea what it's built or what it's even for. The collection of primitives in this talk is the notebook that we hand it. The skills, the work logs, the judgment that let it do real work.
We don't ship PRs that pass CI because the model is a genius. They pass because we built a harness that won't let it be wrong. Every gate in it is a human's call. That's harness engineering. Scale the judgment, not the model. So here's what I'm asking you to do. Don't just go write a skill. That's only that's only one implementation. Find the judgment your team always asks the same person about and make it explicit somewhere that an agent can reach it.
It could be a persona, a lint check or a commit hook. That's the job. Now less product engineering and more harness engineering and the method demonstrates itself. The flag agent produces PRs that pass CI. This talk resumed from a worklog. We started here. Humans absorb judgment implicitly while agents require it explicitly. Self-driving code bases will exist the moment we get that judgment out of our heads and into the codebase.
Then we point the agents at the gate and let them spin until it's green. Thank you. >> [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.