Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 13:49
2.3x that video's typical replay level
Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well, you know? It will usually give you some context about what happened.
Said at 13:43
The graph counts replays. It does not show where viewers stopped watching.
Words
3,782
Runtime
22:32
Speaking pace
168wpm
Reading time
16min
168 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> All right. Hey everybody. Uh, yeah, yeah, yeah, yeah. Okay. I have about 18 minutes. Um, I have a slide deck here that that Claude put together for me, but Claude still can't do very good slide decks, so I apologize for that. Just just just focus on me, right? Right here. This is like the most fun thing to do every now. If I if I start like running out of time, can can you like call time for me? All
84 words, the words spoken in the first 30 seconds at 168 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 333 |
| Average words per sentence | 11.4 |
| Longest sentence | 102 words |
| Questions asked | 102 |
| Sentences containing a number | 22 |
Most used terms
Filler phrases
180 in total: like 55 · right? 49 · um 19 · uh 15 · you know 14 · actually 12 · kind of 7 · I mean 5 · sort of 4.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> All right. Hey everybody. Uh, yeah, yeah, yeah, yeah. Okay. I have about 18 minutes. Um, I have a slide deck here that that Claude put together for me, but Claude still can't do very good slide decks, so I apologize for that. Just just just focus on me, right? Right here. This is like the most fun thing to do every now. If I if I start like running out of time, can can you like call time for me? All right, perfect.
So, I'm Steve Yege. I am here on behalf of Sneak today. I am not getting paid for this talk or anything like that. They were going to nice enough to buy me a ticket to the conference, which is cool, but I'm mostly here because I wanted to hang out with all of you, and I don't bite, and feel free to come and say hi afterwards. I I wandered the halls amazed at all the brand logos. Um, I'm up here because like the title of my talk is like "Agentic Security", but the real title of my talk is "Be Scared".
See, I Come on. Come on. Level with me. Who here is scared? That's good, but you're all in the security talks. Of course you are. The problem is they aren't out there. Right? I I went to Commonwealth I went to so many places last year, but I one of these big banks I was at, I was talking to their chief security architect in December, and I was doing a Q&A, and talking about vibe coding, and everyone was like asking me questions, and I was you know, knocking them out of the park, and he stands up real quiet at the end, and he goes, "Yeah, so he goes, "If everyone's shipping code at the same at sorry, at 10 times faster and the defect rate stays the same, the security defect, right?
The the vulnerability rate. Then that doesn't that mean that the defects surface goes up by 10x? And I it hit me so hard. I I sank down to my knees. I was just like what are we going to do about this? Because it's a really really important point. Cuz the subtle implied question is not if the defect rate stays the same, the defect rate's going to get worse. A lot worse with AIs writing the code. And so I didn't have an answer for him.
I actually have a partial answer for you today and I'll share it with you, but it's only a partial answer. The real answer is you have to be scared of what's coming. All right? And it's becau it's bec it's it's it's Do I have a slide about this? Yeah. So it's not just that you're just like putting in more of the same defects cross-site scripting and blah blah blah. You're still putting those in. Fable wrote a an XSS vulnerability during the short time that I had with it.
Not its fault. We'll talk about it in a minute. >> [snorts] >> But those are the old ones. All right, we are we all know how to fix those. There's new vulnerability types and new attack surfaces coming and they're here and many of them are incredibly well polished. Like what's an example? You guys know about slop squatting? Right? Where the AI hallucinates a package name. Let's say that you want a graph database and so you're like, "Okay, I'm going to use this graph database." And the AI goes, "Oh yeah, I know.
It's you know, it's graphy 123." And it goes off to the package manager and downloads graphy 123. And it builds and it runs, and the tests pass, and it looks right, but what it downloaded was a backdoor. Because that graphy123 wasn't a real package, but somebody noticed that the LLMs are hallucinating its name, and they uploaded one that does exactly the same thing as the one it thought it was getting, plus a vulnerability.
Bonus. Yeah? Is that scary? Yeah, it should be. It should be. How do you even detect that that's happening? So, so we're entering a world where everything you write every bit of code that you generate is going to have to get far more security scrutiny than it's ever had before. Okay? It's just you're not It's hard for me to convey how scared I am. All right? But let's start with where it gets generated. Now, when I worked at Google, I worked really close with the TAP team.
They did tests. The test automation platform, yeah? And they uh So, they ran all of the unit tests and integration tests at Google. They had a massive fleet, you know, and and they learned stuff about how bugs work. And how bugs work is they have a lot they have like a life cycle where if you see the bug right away, you'll fix it. And the longer the time goes for when you see a warning or some sort of issue, the longer that passes Okay?
It's got this sort of half-life of urgency, and all of a sudden, eh, it ain't really biting anyone anymore, okay? This is how we handle all of our bugs. At at Google, they recognized that this is such a human nature phenomenon that they worked really hard to to move the reporting of bugs as you were typing. Because that's when you're most likely to fix it. If you make a bug, and it goes, "By the way, there's a divide by zero here or there's a backdoor or a vulnerability or whatever." You'll You'll fix it right there.
But, if it gets to code review time, you're like, "Is it really worth it?" Right? >> [snorts] >> And the problem, folks, is that that works for all classes of bugs except for security. There's no half-life on it biting you. It's not like, "Oh, because users aren't getting bothered by this security vulnerability that it's not a problem over time." The problem compounds over time. Yeah? So, you have to treat this class of vulnerabilities the way that Google treated their top vulnerabilities at the time, which is to surface them at the developer's fingertips.
What if the developer doesn't have any fingers? I don't know how how many fingers the LLMs have. >> [snorts] >> Um then you need to surface it to the LLMs. But, wait. You say, "Wait. Wait. Wait. Wait. Wait. Wait. Fable's really smart." Or at least it seemed that way for the 2 days I got to use it. Uh can't Fable just write secure code? Right? I mean, come on. Come on. You You all know secure I mean, you're all here in this room, right?
You all know security is an arms race. One that never ends. One that's going up exponentially with Moore's law. One that's going to get real uncomfortable when quantum comes along. Thank goodness that's like 5 to 7 years away. Uh my buddy says 15, so maybe somewhere in between. But, in the meantime, right? LLMs are a real problem. Yeah? So, how do you surface How do you surface the vulnerabilities? Well, first I want to figure out how to do it myself.
I have a game I've been working on for 30 years. I just had Fable do a security hardening pass during the time I had it. Went through and it did all my cloud hardening, and it found a bunch of credentials, and and all a bunch of stuff. And it [snorts] started to give me these vibes like, yep, yep, Harding pass looking pretty good. So, I ran sneak, right? And I I don't have the numbers here. Yeah, I didn't I didn't include the numbers cuz I'm done.
But, um uh it found 241 vulnerabilities, right? Just a ton that Fable hadn't even thought to look for, right? Uh and it's because look, so I I I've told people about the rule of five. When you do things with LLMs, often you have to get them to do up to four to five reviews of the work that they did before it's like actually ready to ship. And it's because their cognitive process is very similar to ours and it goes through a draft and then a revision and then polish and editing until it's, you know, it's finally ready to go.
It's like painting a wall. Some things you don't just do all at one pass. You do them in multiple passes, right? So, um security is one of the So, what I found I wrote a book on vibe coding last year. I did did more vibe coding anything than anyone, you know, two two years ago. And and what I found was um that they're really good at doing one thing at a time. Even the really good models like Fable, right? Just because of this multi-pass painting a wall phenomenon.
You got to give them one task at a time, which means you can't give them security at the same time as you give them correctness. They'll do a half-ass job of both. And you don't want a half-ass job of either of those, it turns out. So, you do it in two passes, right? Now, [snorts] it five It's been five months now, but five months ago I wrote an essay called Software Survival 3.0 where I talked about what software has to do to survive when LLMs can synthesize it all.
I don't know if any of you all saw that, but the basic the basic gist of it is that LLMs can synthesize any software that they want, but they're very lazy in a good way, right? Lazy in like they don't want to spend tokens if they don't have to, because that's money and power and bad for the planet and bad for your wallet and so on, right? And so they use tools to help them whenever it can save tokens, right? So tying it all together, if the LLMs are doing the coding and they're happy to use tools to help them offload cognition, you see where this is going?
Give them sneak. Give them chain guard. And I still think there's a missing piece in this picture that I'll tell you about at the end. Um chain I don't know if you all know about chain guard. Chain guard uh Chain guard is a supply chain that you sign up for and they give you images that have been pre-vetted to not have vulnerabilities and they update them. So it's your inputs, okay? And then sneak handles everything else.
The code that you write, the code that the LLM writes, the dependencies that you're pulling in from slot squatting, right? The innocent stuff. It can find My understanding is that sneak can actually find vulnerabilities that are proprietary, that only they know about because they're ahead of the CVE registry. I There's some truth to that. When I ran them on my code base, it didn't find any vulnerabilities that weren't already public CVEs, but damn, it was easy to [clears throat] use, right?
I think that a tool like sneak is going to give your LLMs superpowers, okay? Cuz what you do is you add it as a pass to the prompt that you give them for whatever they're doing. And say one last thing to look at and have them run your security analysis, all of the tools, get the open source ones, get the sneak one, get the chain guard one, get the all of them and have them check each other's work, too. Right? If you want to get really serious about this on lunch time, right?
But secure your supply chain because Five Eyes is warning us. Yeah? My poor my poor game. By the way, please please I just told you that my game has 241 vulnerabilities. Don't go hack my game. Give me a couple days to fix the bugs. Yeah? We all good here? Good. All right, but if you do hack it, you're going to bother like five players, all right? >> [snorts] >> Okay. They're very loyal. Um look, Five Eyes, which is like a bunch of right governing it's it's big countries that are have have their eye on on the cybersecurity, you know, landscape.
They just announced that it is now months, not years, until it starts happening. It. You all know what it is, right? It is when open source models catch up to Mythos. Does anybody here believe open source models are going to catch up to Mythos? Interesting that it's about 50/50. Anyone got a time frame in mind? Who said December? That's pretty accurate. Cheater. Yeah, it's about 7 months. So, actually it's shrinking, so it's probably about 6 months now.
Yeah. And Mythos is real real good at hacking your systems. And and just just remember you can't trust it to automatically write good code any more than you can trust it to write elegant code by default. That's a separate concern. It's a separate pass. You can't expect it to write performant code by default. That's another pass. You see what I'm saying? You can't expect it to necessarily write the code according to your company coding standards, okay?
These are all passes that go through your code and I just want you to remember that security should be your first one and your last one. Okay? Give it extra. Okay, and the last thing I wanted to talk to you about first of all, go do all this, and second of all, the last thing is really true true truthfully, okay? Dial it in here, folks. Go to your families offline, like in person, and get your your code words refreshed.
Cuz another kind of scam that's coming along is you get a call from a family member who's in distress, and they need money, and it's very convincing, and there's a video of them, and you're going to need a way to distinguish them from from AI. Okay? It's months away, and some of your families are going to be slow to catch on to this stuff, but bank accounts will be drained. I heard that Congress was given secret demos of draining bank accounts.
I've been scared of this for close to 2 years. I heard one talk from a security researcher almost 2 years ago at ETLS Las Vegas, and he stood up in front of the crowd, and he said, "You're all not scared enough of what's coming." It'll affect you personally, not just your company. Okay? So, that's my message to you. It's not a message of hope and positivity today, >> [laughter] >> but it is a it is a message that that's that should be clear crystal clear is that there are tools, open-source tools, free tools, commercial tools, okay?
Techniques, practices, okay? That you can use right now to get started on fighting in this arms race and protecting yourself. And that's all I've got today. Thank you. >> [applause] [applause] >> Questions? We ready to go? >> Sweet. Um just super simple. What has surprised you recently? In the world of AI coding. >> What has surprised me in the world of AI coding? Well, I'm not really super representative. I spend a lot of my time trying to predict the future by like hammering agents really, really hard.
Yeah. Um so um Uh you know, one of the surprises and I shouldn't have been surprised, but one of the surprises is that AI is moving faster than the world is moving. Uh tech is moving faster than society can move. And the surprises show up when friends, smart friends, resist uh you know, the inevitability of AI and they call it psychosis or they right? Or they poo-poo it and they say, "Well, it'll never be actually smart." Or whatever.
They they can't see the curve. Right? And uh and that surprises me. Uh maybe it shouldn't. Um it's for it shows a sort of tunnel vision. I think people have a tendency to look about 3 months back and about 3 months forward and be like, "Oh, it looks pretty flat." Right? But but and so that that surprises me that people aren't honestly that people aren't more scared and that and by the same by the flip side that people aren't more excited by it.
Right? Cuz you know once you actually, you know, once once you get it, I mean you you don't even want to be here. You How many of you are running Claude code right now? Most of you, right? It's really fun. So, you know, I mean like that's a surprise, too, that the world is pushing back so hard on that, right? We're in an awkward phase. We'll get through it. Any other questions? Woah. All right. Well, you pick. Feel free to bail also.
You don't have to stay. >> I big fan. Um, what's the like most impressive uh thing you've seen Gastown do? And like how much human intervention was involved or steering? >> Oh, Gastown? Yeah. Gastown. Gastown was a lot of fun in January. Um, yeah, uh Gastown is a Beads machine and I still use Beads. And I'm working I want to I was talking to Angie, I want to donate Beads to the Agentyc Foundation, you know. We're going to we're going to put multiple backends on it.
Beads is a task tracker, right? Beads is how you do Boris Cherney loops. You know how Boris is like you shouldn't be prompting your agent? Has anybody here actually like successfully How often do you get Claude to actually run all night for you? Like for real run all night. A few of you, right? Like this is the next frontier, I think, of actually getting agents to run like for a long long time unsupervised. You can do it with Beads by queuing up enough enough work and having them claim and all that.
And there are some other some other systems that'll do that. It was really fun when Gastown did this for me automatically once. I filed a whole bunch of Beads and they disappeared and I was like, "Oh, no, another bug. My Beads disappeared." And what had actually happened was that >> [snorts] >> one of the agents just found them and just implemented everything, right? So I was like, "Whoa, I really like swarms now." Yeah.
Fun times. Does anybody here regularly work with more than 10 coding agents at once? You see like not many, right? The world is still in the we're still kind of like prompting and using a few here and there, right? It's going to accelerate really fast next year. Other questions? You. You can just say it. >> Yeah. Um, your talk focused a lot on vulnerabilities from agents writing code. I'm curious what are your thoughts about agents taking actions >> Yeah, so that was the third dimension that I really wanted to talk about.
It's just I don't really have time, but like my friends over at Tessell, I'm advising them. They're actually doing this, right? There's just this whole space of who's looking up who's looking over your agent's shoulders. I had this conversation just now. Like I tell everyone, everyone's just starting to stand up agents like 24/7. Like processing queues, responding to events, like agents that actually do stuff, right? 24/7 and I I I encourage people to think adversarially.
I I think of adversarial groups of agents tasked with doing that queue management cuz one agent will always eventually screw it up. Right? So, you got to have those supervisors. And so, there's whole systems emerging here. Kind of can go out and go look at all of your and say and start to like you know, do do do that hardening stuff. Like do they really need all those credentials on that service account? Really, right?
Only for this one action, maybe we can like separate this one out. That kind of thing, right? This is a brand new frontier, but it's one, ironically, that even though there's kind of an almost nothing out there, there's some experimental stuff, you still have to be thinking about it right now and designing a solution in-house right now, right? Because otherwise your engineers are going to spin up or your non-engineers are going to spin up a bunch of agents with way too many permissions and then, right?
As soon as a bear munches into the igloo, everyone's dead, right? That's the old security analogy. I don't know if they still use that one anymore. >> [clears throat] >> Yes. >> Hey Steve, nice to see you. Um I'm curious in terms of like this the best practices that you've seen uh that you use personally or or maybe you haven't tried, but um particularly with regards to prompt injection. So, >> Yeah, so I was supposed to talk that during my right during during my speech here, I I I wanted to mention it. >> [snorts] >> There are a whole bunch of attacks happening on the training and on the prompting side, right?
So, on training and inference. >> [snorts] >> And so, the bad guys find ways to sneak in stuff, right? Um and it's like like like the simplest version is the new XSRF where like the user puts in some some text and then some bad actor puts in some extra text saying disregard everything and do the following, right? And then they just get more sophisticated from there. Um I mean, I don't have any good answers for you other than like this is real.
It's kind of an education problem at this point. You need to get everybody thinking about it, right? And then you I feel like there are like new security roles about ready to emerge inside of companies, agentic security that it's kind of an extension of what they're already doing. Who's already doing this? Right? Yeah, so you you've already got people that are going out and looking after the the sort of security of your agents that are in that are deployed in the Yeah, so you're way ahead of everyone.
I'd love to come talk to you later. Ha. This is all brand new stuff, yeah? Cool. >> [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.