Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 11:58
4.6x that video's typical replay level
issues. Uh I also invented OS certification. I just close the tracker whenever I want, so I have my life back. So, does this work? Yes, sort of. >> [laughter] >> Which leads me to act three, slow the down. Everything's broken.
Said at 11:52
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
The graph counts replays. It does not show where viewers stopped watching.
Words
4,406
Runtime
24:40
Speaking pace
179wpm
Reading time
18min
179 words per minute, just under the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Hello everybody. Thank you for following on a session about the death of the code review with a session about the death of the code review. Uh who knows howululing decisions get made, but uh Swix decided it would be funny for those two to be backtoback. Uh, hi. Uh, I'm Lori. I'm head of developer relations at Arise AI. Uh, some of you may remember me from when I used to co-found npm, Inc. These days, I think about AI and how to test it. Uh, I'm here to
90 words, the words spoken in the first 30 seconds at 179 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 267 |
| Average words per sentence | 16.5 |
| Longest sentence | 109 words |
| Questions asked | 12 |
| Sentences containing a number | 27 |
Most used terms
Filler phrases
265 in total: uh 197 · um 50 · actually 10 · like 4 · sort of 2 · kind of 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Hello everybody. Thank you for following on a session about the death of the code review with a session about the death of the code review. Uh who knows howululing decisions get made, but uh Swix decided it would be funny for those two to be backtoback. Uh, hi. Uh, I'm Lori. I'm head of developer relations at Arise AI. Uh, some of you may remember me from when I used to co-found npm, Inc. These days, I think about AI and how to test it.
Uh, I'm here to talk about a problem that everyone is seeing right now. The rise of AI agents has dramatically increased how fast developers can produce code. Uh, but the speed at which humans can review code. Uh, yeah, I am having a talk right now. Thank you, notion. very good reminder. Uh the team the speed at which humans can review code has stayed exactly the same. Um and that is creating a new bottleneck that everybody is feeling.
Engineering teams across the industry are feeling it. Uh so what can we do about it? Some options are uh we could skip human review entirely. Uh some folks are trying that. Can we automate reviews reliably? Some folks are trying that too. Uh what we're doing going to do today, what I'm going to do for you is I'm going to look at what the industry is really doing right now uh and try to figure out what you can do uh today when you leave the room.
So let's start with two numbers. Uh recently three economists tracked more than 100,000 GitHub developers uh matched them against telemetry that showed exactly when each one started using AI. Uh the developers who turned on autonomous agents wrote 741% more code but only 30% more software shipped. Uh that is the scale of the problem right there. They were writing code nearly eight times faster. Uh but their actual ship software only rose by a third.
Uh and the authors of the study are blunt that review was the bottleneck. Um the problem is that the route to production still runs through humans. uh and the human steps uh in particular code review uh choke everything downstream. Um producing code is suddenly a whole lot cheaper. Knowing whether to trust it is still very expensive, especially if you have uh sensitive code with a blast radius. Um how can we bring that second number down and thus bring this the amount of software that we ship back into line with the amount that we can code?
Uh let's zoom out first. Let's be clear that the problem is real. generation is no longer the bottleneck. Uh Stripe in Anthropics launch materials for Fable this year reported migrating a 50 million line Ruby codebase in a single day. Uh which is work that they'd estimated would take over two months for a team. Um Bun, which is now part of reported that they migrated over a million lines of Zigg uh to Rust in six days.
Uh and of course all of us in this room are feeling this on a smaller scale. uh our agents are generating whole apps and we're just sort of hitting the merge button and feeling guilty about the fact that we haven't read these code diffs and we're just sort of hoping that it works. Uh we can't possibly keep on going like that. Um so the obvious answer and the answer some people are trying is you just review more. You take people whose job was previously to review code as well as write code and just say review code all the time.
This is your job now. Uh and that doesn't work. uh partly because that's really boring and those people burn out really fast. Um but also the numbers say that we can't um the best study we have on this was done uh two decades ago at Cisco um over 10 months they took 2 and a half thousand reviews they took 3.2 2 million lines of code. Uh, and the study says that reviewers stop finding defects effectively if they try to read more than 400 lines of code in one sitting and their effectiveness completely falls off a cliff.
Uh, if they try to review more than 450 lines of code in an hour. Uh, if you do the math on that mean that means that a 10,000line agent pull request at that pace would take three or four working days to get a real human review. that is one pull request from one agent. 10,000 lines is uh absolutely you know absolutely an expectable number of of lines of code uh for an agent that is doing a lot of work for you. Uh and a developer can now run a dozen agents at once.
Um although I'm tend to be suspicious of the people who do um so we can't just review harder and we're all feeling that too. The people who are trying to review harder are burning themselves out. So what is the alternative? Uh some people have decided to just stop reading the code entirely. Uh Peter Steinberger who is the creator of OpenClaw uh says that you shouldn't be prompting coding coding agents anymore. You should be designing the loops that prompt your agents.
Uh Andre Karpathy who is one of OpenAI's founding engineers has made the same argument about taking yourself out of the loop. Uh because the human in the loop is holding the system back. Uh but this goes beyond bold claims on Twitter. Uh in February this year, OpenAI published an account of building an internal product uh with in their words no manually written code. They started with a completely empty repository uh and agents wrote everything.
Five months later they had about a million lines of code about 1500 merged pull requests uh pull requests uh and they had done used three engineers to do this. Uh even the scaffolding of the agent to review the agent and stuff like that was also written by agents. What they said was humans may review pull requests but they are not required to. Uh we've pushed almost all review effort towards being handled agent to agent.
I think the interesting thing about the OpenAI experiment is that they did not tell us what the product did and they did not review an open they did not release an open source product saying and this is how we did it and this is how you should do it which suggests to me that there are still holes in that strategy. Um but it is one possibility. If open AAI says that you can do it, maybe you can do it. Um but one poss so agent open AI says that's one thing you can do.
The bet there is that loop design substitutes for inspection. Uh whether you can build a loop that is good enough uh to review all of your code for you depends entirely on what the loop can see. Uh but that is an enormous caveat. What exactly can the loop see and is it reliable enough to do code reviews? For years, the industry's proxy for uh reliable for quality has been whether the tests passed uh because that is what the benchmarks measured.
In March this year, uh a research group called Meter, who you've probably heard about before, uh tested that proxy directly. They hired four active maintainers of open- source projects uh from the same open source projects that Swebench tests against and they got them to look at uh PRs that had already passed SWEBench is greater. So SweetBench said this pull request is is good enough to merge and they got the open source reviewers to look at exactly the same PR and say is it really good enough to merge and it was only good enough to merge about half of the time.
Uh, and the failures weren't about correctness because obviously the PRs were passing all of the tests. They were about code quality and they were about changes that quietly broke other code, things external to the test suite. Uh, one caveat that meter mentions is that the agents got no chance to iterate on feedback. So, a human uh, contributor obviously being told that their PR isn't going to get merged get another go at it.
They can take another swing. Uh but that caveat doesn't really isn't really relevant to our purposes because that's just putting a human in the loop and what we're trying to see is whether or not we can take the human out of the loop entirely. Uh Cognition, the makers of Devon, uh built a benchmark around the maintainer's actual question, which is would you merge this? Uh they called it frontier code and they launched it in June.
Um more than 20 maintainers built 150 tasks from their own repositories, each one over 40 hours of expert work. Uh Frontierbench grades behavioral correctness, regression, safety, scope, discipline, test quality, and maintainability. This is a human review rubric uh made machine checkable. Um, and what they found was that Fable 5 before it got pulled and then unpulled, uh, scores 88% on Swebench Pro, but only 29% on the hardest slice of, uh, Frontier code.
So, the same model on the same surface with the same job does 51 points less uh, less well if you're asking not does it pass the tests, but whether or not uh, you would actually merge the result. Um, and this isn't one model having a bad day. on the same set of tests uh GPT 5.5 scores under 6%. So the strongest models we have are nowhere near uh passing human review reliably. So somebody has to write down what mergeible actually means.
That is it's clear that we haven't done that yet. Uh the moment they do something will happen. Sarah Guo who is a prominent investor in AI uh wrote about it and said that solving for that solving for that benchmark a benchmark of genuine mergeability uh will be a critical turning point in the development of coding models um in the essay where she was talking about this she made clear why models got so good so fast at generating code and that is because a compiler is a free verifier a test suite is a free verifier and anything that you can verify cheaply you can train against until you beat it and that is what has been happening uh at the major model trainers.
Uh so a benchmark of mergeability if we could if we could make one wouldn't only be a measurement it would immediately become a training signal for the frontier models. Uh which means maintainability scope discipline regression safe safety the moment somebody comes up with reliable tests for those the big models will be trained on them. Uh and whoever writes today's review standard is writing next year's default model behavior.
Uh we have some prior art here knowing we have seen this happen already. Uh which is that in 2024 OpenAI trained a model called critic GPT to catch bugs in model written code. Uh human reviewers working on it beat the model alone uh and the human alone. Um, so OpenAI built a model to review the models in order to clean the training and it created a training signal that's now built into the models. It made the models better almost instantly.
But mergeability standards are still in the future. Right now we do not have a good rubric that captures everything that a human decides uh to uh consider when they are deciding whether something is mergeable. And I promised you something that you can do today. Um so who is actually running automated review right now? Uh one company is GitHub. GitHub is uh copilot's reviewer has done 60 million reviews and now accounts for more than one in five uh code reviews on all of GitHub.
So machine review of pull request is the mainstream default uh on the world's largest code host. That is more than an experiment. That is a large production uh a large production deployment. Uh, Curser is also doing an enormous amount of code review. Cursor uh, has published its reviewer's architecture, so we know how they're doing this. Um, and the details tell you what the job really is. The first version of their reviewer uh, would run eight review passes over each diff uh, and then shuffle the order of the reviewers uh, to review the code again and again uh, because the order in which it did those reviews changed the outcome of those reviews. what they were doing was filtering for false positives because a reviewer, an automated reviewer that flags something as bad when it's actually good, uh, is going to get ignored.
Um, and it's not just cursor who's doing this. There's research about it. Uh, a team at PKing University tested the same idea independently. Run several review passes, see keep what they agree on. Uh, and they found that it review raised review quality by up to 44%. Uh the multipass trick keeps getting rediscovered because uh false positives are the thing that kill a reviewer and it actually works. So then cursor rebuilt uh their code reviewer.
Um so that the model reasons over the diff, calls tools and decides where to dig. Um and my favorite detail from their rebuild uh was that they had to tell the model to be more suspicious of the code. Uh the model tended to look at the code and say, "Well, that looks good to me, so ship it." Which is exactly what a human would do in that situation. uh they had to tell it don't trust the code by default. Assume there is something wrong with it.
Um and somewhere in that sentence is a whole is a whole talk about what a good review actually is. Uh it is about being suspicious by default. Um the next thing that's happening is that review is starting to fuse with repair. Cursor's reviewer now spawns a fix agent from its own findings. So it doesn't just flag the bug. uh it writes a patch uh and it hands you back uh a diff to approve. Uh you approving that diff is obviously another code is another human in the loop where we don't want humans in the loop.
Um the next thing cursor wants to do is the reviewer running the code to prove its own bug report is real. Um the line between reviewing and rewriting is getting very thin indeed. Um and GitHub and cursor are by no means the only vendors getting into this game. It is uh getting very crowded in there. There's a ton of companies uh in the field. Uh Code Rabbit is the largest dedicated reviewer has now reviewed over 13 million poll requests.
Uh Graptile builds a graph of your whole reposi of your whole repository so that the reviewer can see how a change lands in distant code. Uh graphite builds its evaluation set from which of its own suggestions developers accept or reject. Uh but the thing to notice uh about what every one of them is doing is that all of them are using the same metric as their definition of success which is is a human accepting my answer.
Uh cursor calls it the resolution rate and it has driven it from 52% to over 70%. Um the reviewers are trained on human accept judgment every day at scale. They are training their harnesses to get better at the definition of good as defined by human acceptance. That is a preview of what the models are going to do except these companies have already shipped it. So the status quo is you can generate code automatically, you can review the code automatically, but humans still need to review the the reviews.
Can you skip that part entirely? Can you get the human entirely out of the loop? Uh there are two prominent projects that I've that have tried so far that I've heard of. Uh in February this year, Anthropics Nicholas Carini had 16 agents build a C compiler from scratch in Rust. uh it was able to compile the Linux kernel uh across about 2,000 sessions with no human in the loop. But if you read that experiment closely, uh it's true that there was no human in the loop.
Uh no human approving the code as it was written, but there was absolutely a human on the loop. The system uh that reviewed the code, the system that checked whether the code was doing what it was supposed to do, all of that uh was written by a human. Um the test harness, the feedback systems, all of that stuff. Um it was automated testing with tests that took a human to write them. Um Carini's own warning when he wrote about it was that it is easy to watch the tests pass and assume that the job is done uh and that it rarely is.
Another very widely publicized experiment which I al which I already mentioned briefly in passing was when bun ported its entire uh runtime from zig to rust. That was about a million lines of code by agents in six days. Obviously no human read that diff. Um the gate for that experiment was the existing test suite. 99.8% of the test suite passed. Um so the test did real work. Uh a fun fact about that experiment is that the PR uh where they uh where the agents decided that the they were done with all of the Zig code and they deleted all of the Zig code in one giant PR was flagged by another robot as this is AI slop.
You can't possibly delete all of your code. Uh which I thought was a fun aside. Um, but there are some big caveats on that bun experiment. Somebody looked carefully at the ported code and it has 13,044 unsafe blocks. Uh, in a comparable human written uh, uh, Rush codebase of that size, uh, you would see about 74. So, three orders of magnitude more unsafe memory unsafe blocks. An unsafe block is a place where the author asserts rather than proves uh, that memory is being handled correctly.
So the test suite can certify behavior at the public interface. Um it cannot certify 13,000 assertions that the test suite was never designed to look for. So they took humans out of the loop for sure. Uh but there's now no knowing what is lurking under the surface of their rust as a result. So let's go back to OpenAI who have pushed this the first list because they show what skipping human review costs. They didn't delete review, they moved it. uh codeex reviews its own changes uh then calls in more agents to review those reviews um in a loop until every agent reviewer is satisfied.
OpenAI made Codeex bootable at every single change so that Codeex can actually run a copy of Codeex, look at it, look at the UI, see if the bug is being fixed. Uh and they expose the whole logging stack to the agent. Um, and a line from their writeup says exactly what the Cisco study said, which is that when something failed, the fix was almost never to try harder. It didn't work sometimes. Uh, and one detail that you'll recognize from your own week is that for a while on this project, they had to spend every Friday cleaning up AI slop.
Uh, by hand as humans. Um, so that eventually didn't scale and so they trained agents to look for AI slop and get rid of the AI slop. uh review didn't disappear is the lesson of this of of this experiment. Uh it got rebuilt as a system and that system is built by humans. So can you skip the human entirely? Uh someone who tried it for real and then changed his mind is Dexter Horthy. Uh he spent six months telling people not to review the code.
He famously did that at AIE last year. Um he told people just ship, let the agent do its thing. Uh and in March this year on stage uh he took that back. Uh he said I was wrong. Please please read the code. We tried not reading the code for like six months. It did not end well. We had to rip out and replace large parts of that system. That is not a benchmark. That is somebody who ran this experiment for real. Tried it with real code on a real system uh and retracted his remarks.
Uh OpenAI is running uh lived with the results um and reversed in public. Uh so Sarah Quo uh in the same writeup that I mentioned earlier talked about why that happens. Passing the test never told you that the change was the right change. It never told you that this module exists because there are three external users of this module. Uh it never told you that there's this cron job that nobody will admit to writing that relies on that module existing.
Um there is context outside of the test suite that the tests don't uh can't find. So in everything that we've looked at so far, the human uh checkpoint is continuing to survive. It is moving around uh but it is surviving in predictable places. One is where correctness isn't cheaply checkable. Another is where uh the blast radius is large. In security conscious environments, every time you tell people, oh, we can just get rid of human review, you get an immediate no.
Uh wherever someone has to put their name on the result. Uh but what's happening is the human role isn't disappearing. It is moving up the stack, possibly several levels up the stack from inspecting the code directly to designing and tuning the systems that inspect the code and designing the definition of good. Which raises the obvious question, who reviews those systems? If you've built a system that does your reviews for you, how do you do the meta review of your reviewer?
Um and the answer won't surprise you. It is humans. Uh, Anthropic ships an automated security reviewer and in its readme it has a huge caveat uh, which is this action is not hardened against prompt injection attacks uh, and should only be re used to review trusted PRs. So, Anthropics code reviewer uh, can be talked out of its findings by the very thing that it is reviewing. Um and that is a finding that has been reproduced in study in March this year reinforced uh that vulnerable code dressed up in an innocent commit message uh fooled an autonomous review agent in 88% of attempts.
Uh the same attempts uh sent to a human reviewer uh passed only 35% of the time. So you you take the human out of the loop. Uh you don't just lose a reviewer, you lose the thing that was hard to fool. So automated reviews fall for confidently framed bad code. And confidently framed bad code is exactly the kind of code that agents are very good at producing. They're very like this is good and I am ready. Uh and they uh are wrong sometimes.
Um and the other problem is the field doesn't even agree yet on how to grade these graders. Uh benchmark scaffolds have been caught leaking answers. Researchers disagree on how to review quality at all. um which leaves one reviewer that you can't automate and you can't skip which is production. Uh once the premerge review is all machines watching what the code actually does becomes the last reviewer standing. Um once the code ships the test results stops being the interesting thing the trajectory does what the system actually did step by step uh when it ran against the real world.
Uh I'm not going to give a pitch for a rise here. there are enough pitches for a rise at this conference. Uh but if you're shipping automatically reviewed code, then systems in production that review what it actually does in the wild become indispensable. Um so after all that evidence, after all of that review of what the world is doing about automated code reviews, where I land is that code review isn't completely dead, but it is changing an enormous amount.
Uh it is being rebuilt as an engineered system. Humans are moving from being the engine that drives a code review uh reading code line by line uh to its pilots. Uh and given the numbers that we opened with that is probably the right trade. Every layer of that system, the benchmark, the classifier, the rubric, the test suite, the eval is itself unreed until somebody decides that that is their job. Um and the teams that win the next few years won't be the ones that generate the most code.
They'll be the ones who can say with evidence why they trust what they shipped. Uh and I open by promising you that I would give you with something practical that you could walk away with something that you can do today. Uh and that is to stop reviewing PRs. Um it is the wrong level of abstraction for 2026. Your human judgment is extremely valuable, but it can be made to scale much further than it is scaling right now.
Uh pour your precious time into building uh a reliable review harness. codify your definitions of good, your company context, your domain knowledge, and then uh crank up the agents to work with that. You can go much faster than one-third faster uh if you concentrate your efforts higher up the stack of reviewers, rules, and evals. Uh I hope this look at what the industry is doing today has helped you what to helped you decide what you should do and what to expect next.
Uh and thank you so much for your time and attention.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.