Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 11:58
4.6x that video's typical replay level
issues. Uh I also invented OS certification. I just close the tracker whenever I want, so I have my life back. So, does this work? Yes, sort of. >> [laughter] >> Which leads me to act three, slow the down. Everything's broken.
Said at 11:52
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
The graph counts replays. It does not show where viewers stopped watching.
Words
2,460
Runtime
12:40
Speaking pace
194wpm
Reading time
10min
194 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
How's everyone doing? Awesome. Uh, my name is Du. I'm one of the co-founders of a company called Gravile. And at Grapile, we're working on AI agents that validate poll requests with full context of the codebase. The way Graval does this is every time there's a pull request, we have a swarm of agents and they go and look at every file that's changed, every file that's related to the files that changed to figure out if there's bugs. And it also spins up your code in a sandbox, installs the dependencies, spins up local host,
97 words, the words spoken in the first 30 seconds at 194 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 135 |
| Average words per sentence | 18.2 |
| Longest sentence | 66 words |
| Questions asked | 12 |
| Sentences containing a number | 20 |
Most used terms
Filler phrases
33 in total: like 13 · actually 10 · sort of 5 · literally 2 · kind of 1 · uh 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
How's everyone doing? Awesome. Uh, my name is Du. I'm one of the co-founders of a company called Gravile. And at Grapile, we're working on AI agents that validate poll requests with full context of the codebase. The way Graval does this is every time there's a pull request, we have a swarm of agents and they go and look at every file that's changed, every file that's related to the files that changed to figure out if there's bugs.
And it also spins up your code in a sandbox, installs the dependencies, spins up local host, clicks around to try to break things, mocks inputs, and everything else to figure out if the code is broken. But I'm going to talk about something a little bit different, which is fully autonomous coding agents. So I first moved to San Francisco three years ago to work on AI coding. And this is because GBT 3.5 which came out in 2022 was the first model that seemed like it was really good at programming at the time.
Code complete tab complete was like the main paradigm for coding. The most exciting thing that was happening was you had tab complete from cursor and you had co-pilots code complete and that was like the way in which AI coding works. In 2024, for the first time, multifile editing started to work. Cursor was the first to do this and then a couple of other products followed. But now, for the first time, AI could simultaneously edit multiple files at once.
But things got really interesting in 2025 because for the first time, we had autonomous agents. They could be given a task and they could go and fulfill the task and create entire pull requests all at once. And this came to a precipice in December of last year when coding agents became literally completely autonomous. And anyone that's working in AI coding knows that December when the new models came out was sort of this like watershed moment in the history of AI coding because these products for the first time were like truly autonomous in how they programmed.
And all over Twitter were all these crazy stories of people that were having these agents spin off and open a 100 pull requests a day. They were experimenting with polyphasic sleep so they could stay awake to re-trigger their agents from time to time. And as someone that started programming before agents, I was very skeptical that real companies could program this way. So I was really really curious, are these fully vi PRs really actually good?
Is this something that's sort of a Twitter hype and you can have independent developers and and maybe like small startups with no real customers do this? But anyone with real customers with an actually commercially viable code base, could they still use these end-to-end coding agents? And at Grubile, we're fortunate to work with some very large companies. We work with Nvidia, with Coinbase, with Scale, with Data Dog, with American Express.
And we got really excited to see if these coding agents were being used really at these large companies and if they were any good at doing real world coding. So, as an amateur data scientist, I decided to go and dive into our data. We review more than a million poll requests a month and we tend to get a lot of interesting data on what's good and bad about these poll requests. We do this for thousands and thousands of companies and they're usually enterprises or at least companies with serious products with real customers.
And so we figured actually studying this data would yield some really interesting results. Was a surprisingly hard to figure out from all the poll requests that GR was reviewing which of the poll requests was actually generated by AI. Surprisingly difficult to do this. The first thing that I tried was I started looking at the GitHub author field. GitHub every commit has an author field and so went through the authors and turns out that less than 1% of pull requests had codeex or clawed or cursor as the author.
But it wasn't intuitive to me that only but 1% of all code was fully AI generated. It just seemed intuitive that that number would be much higher. So I started to look for other signals. Now thankfully Claude and Codex and some other products leave a PR description in the footer. In the PR description, they put something along the lines of co-authored by claude or co-authored by cursor and so on. And this gave me a little bit more signal, a little bit more data on which PRs were ostensibly largely, if not completely AI generated.
The third thing I looked at was branch name prefixes. If you use Codex, you know that Codex names as branches. And it seemed reasonable that if someone was using these coding agents in such a sense that they were writing the branch names or writing the PR descriptions that the PRs were largely AI generated. That seemed like a reasonable assumption to me. So now I had a good set of signals to indicate whether a poll request was in fact completely or largely AI generated or not.
And I came to the result that about a quarter of all the poll requests that Graal was reviewing in any given month were completely or at least largely generated by AI. Then I backtracked this data to the last 12 months and it turns out that this number is growing really fast. In fact, early last year, fewer than 1% of pull requests had any evidence of being completely AI generated. So this number is going very very fast as model performance is getting better.
What's also interesting is that you can't tell when new models came out on this chart. It seemed like the progress is very continuous. These things are just generally diffusing into the economy at a generally high pace. So then came the next question. Everyone's vibe coding. Everyone's producing these PRs that are antiendentic. Not human plus AI but literally AI. The question is are these PRs any good? First question, what does it mean for PR to be good?
That seemed like actually a very important question to answer before we got into judging these AI generated PRs. And so I took a few methods to try this. The first one I tried was revert rates. If a pull request reverted, it was probably pretty bad. And so that seemed like a pretty sincere sort of guess at what a bad PR could be. And so I started tracking, okay, which PRs were in fact reverted. GitHub conveniently names its branches with a revert dash PR number and PR name.
And so I was able to track the rate at which these things were being reverted. And I found some interesting data. Codex PR is reverted about one out of every thousand poll requests. Devon once every three and a half times every thousand poll requests. Humans right in the middle at about two and a half. So there doesn't seem to be very big difference between the rate at which poll requests were reverted from people versus agents in my study.
Now I was very skeptical of this and I figured okay this is probably because humans are making agents do the easier work the simpler well scoped pull requests and the more complex work was being done by humans and of course complex work had a higher propensity to be reverted because the poll requests were more complicated had larger surface areas risks and so on. So I decided to measure the average size of PR and whether the revert rates were sort of aligned for both of those and I found really interesting data.
Turns out there's actually very little correlation, if any, between the size of PRs of humans were getting reverted versus size of PRs from agents that were being reverted. So this does not seem like actually very strong evidence that the human PRs are any better than the agent PRs. The second set of signals I started looking at was Graile comments. Now, now Grapple is reviewing all these poll requests and Graile finds P 0's, P1's, P2s across all these changes and it seemed reasonable that if Grapile was finding more bugs in the code that the code was most likely worse.
So then I started tracking the counts of P 0, P1's and P2s in all these changes. And interestingly, once again, there was not that much of a difference. In fact, three out of the four agents that we tested performed better than humans in terms of the rate at which they were producing P 0. they were producing fewer P 0 than humans were. Similarly the case with P1's and P2s. Broadly speaking, human generated PRs were about equal in quality to agent generated PRs based on this data. have the agent go and look at those comments, address them, and make a new commit on the pull request branch.
That is the most common way to use guptile. And so it seems reasonable that if the pull requests were higher quality, it would require fewer iterations before they would be merged. And so I started tracking the number of iterations, the number of review rounds between when the poll request was opened and when it was merged. Sure enough, very little difference. Devon's PRs 2.1, Codex PR is 2.45, four, five number of review cycles to merge and humans right in the middle.
Once again, very little if any statistical difference between human generated and AI generated pull request in terms of how many iterations before they were ready to merge. So now that there wasn't any strong evidence that agent generated PRs were worse than human generated PRs, I got curious if there were qualitative differences. Maybe the types of ways in which agents failed were different from the types of ways that humans failed.
And so then I looked at the corpus of reptiles comments. It makes on average about four comments per pull request. So we had this corpus of several million comments across the last several months. And so I started scanning them for specific phrases and words. For instance, SQL injection or N plus1 query. And I plotted the frequency with which these terms occur in graphile comments for these various agents. And so I made this chart of patterns of failure.
To interpret this chart, you can assume that 1x is the human propensity for producing that type of error across the entire chart. And you can kind of see that there's actually quite a lot of variation in the types of failures that these agents seem to produce. For instance, Claude is one and a half times more likely to produce a SQL injection error than humans. Devon is about half as likely as humans to produce a off bypass issue.
I found it very interesting that there was this much variation in how these agents were performing and how different their failure modes were from humans. So it turns out that in spite of my initial skepticism around the enterprise usability of endto-end coding agents, the evidence seems to suggest that they're here and they probably can contribute in real meaningful ways to enterprise coding environments. And so we started to think a little bit more about what code review would look like in such a world.
Here's an interesting stat. Today, Graal is used by several tens of thousands of engineers every single week to review all of their code. So, we also know generally speaking the number of pull requests that each of these people are writing. The median grubile user writes 50 pull requests a month. So, call it about two plus per workday. The 90th percentile writes 500 poll requests per month. That is a drastic difference between the median and the P90.
The P99 is in the thousands of pull requests. That is very interesting in its own sense because that means that the people on the margins are actually producing poll requests at the rate at which they're coming up with new ideas. And beyond being interesting, it also makes you wonder well there existing systems for validating this code which is manual code review of course testing maybe you have a QA firm that you work with naturally can scale to that same degree.
And at Grav, we decided to take sort of a first principles view at what really good validation could look like. Instead of saying that we wanted to automate QA or automate testing or automate code review, we took a step back and said, what would we need to do for anyone to be able to merge hundreds of pull requests a month in an enterprise environment where it really matters that the code is correct? What would need to happen between when the code was expressed into a pull request and when it was merged and deployed safely?
We figured we only actually had to answer three questions. The first one is does this change violate the user contract? The second one, does it increase the propensity of a future violation of the user contract, whatever the user contract might be for that application? And third, does it fulfill the intent that the author described? Did the poll request do the thing that the author wanted to do? And so we started approaching this problem from base ground level and said, okay, agents can probably figure out if something's going to violate the user contract and detect bugs.
If you let it spin up the code in a sandbox, have it install the dependencies, mock the inputs, run the browser agents, you can probably start to discover most of the issues that could occur, you get a pretty high degree of confidence on merge. Today, almost a fifth of all the poll request or gravile reviews are merged without any human review or without any human testing, which I find very interesting and that is a number that we care a lot about and want to bring higher and higher, of course, within the guard rails of producing really high quality code.
Thank you so much. My name is D, one of the co-founders of Graptile. We have a booth here which you should come to. And if you're interested in trying Graptile, you can find us at gretile.com and try it today for free. And we'd love to hear your feedback. Thank you so much.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.