Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 6:57
5.9x that video's typical replay level
do light mode. It's I It's not my nature, but sometimes. That's better, yeah? Okay. So we have we have a model and we're trying an old LG Sorry. We We shouldn't have seen that. No, we'll
Said at 6:50
The graph counts replays. It does not show where viewers stopped watching.
Words
3,066
Runtime
15:05
Speaking pace
203wpm
Reading time
13min
203 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> Cool. Hey everyone. Thanks for joining the session. Sorry if my sound's a little croaky. I've done three talks back-to-back. So, one on Wednesday, one yesterday keynote, and then workshop style session today. So, I'm Vincent. I'm going to be talking about malleable evals um from static AI measuring uh to adaptive systems. Now, let's jump into who I am, what I do. Uh call myself the friendly can car. I use AI, use technology. I'm always on the edge. Um for those of you that haven't seen my keynote, I do um yeah, I just live on the edge
102 words, the words spoken in the first 30 seconds at 203 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 160 |
| Average words per sentence | 19.2 |
| Longest sentence | 63 words |
| Questions asked | 23 |
| Sentences containing a number | 13 |
Most used terms
Filler phrases
240 in total: like 130 · um 31 · kind of 30 · you know 14 · sort of 11 · uh 9 · actually 8 · right? 5 · I mean 1 · basically 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> Cool. Hey everyone. Thanks for joining the session. Sorry if my sound's a little croaky. I've done three talks back-to-back. So, one on Wednesday, one yesterday keynote, and then workshop style session today. So, I'm Vincent. I'm going to be talking about malleable evals um from static AI measuring uh to adaptive systems. Now, let's jump into who I am, what I do. Uh call myself the friendly can car. I use AI, use technology.
I'm always on the edge. Um for those of you that haven't seen my keynote, I do um yeah, I just live on the edge and and just do some fun stuff. So, this is me using VR goggles in um like back in 2013 when like people hadn't even heard of VR. It came with a warning label, said only use it for 5 minutes. I used it for 3 hours, then I vomited for 3 hours after that. So, measurement, anything we do in technology, anything on the edge is going to be janky, it's going to be weird.
And that's kind of fun, in my opinion. Now, whenever we talk about evals to people and and a little bit of pretext like my my role at uh Comet um work in evals, I do an eval research. I work with universities. We benchmark and run evals for large set of companies and organizations, everything from like Uber to Netflix to to banks even in the UK. Um but the the thing that's been going on right now is that hey, like there's kind of joke that like evals are a little bit dead.
Um and it's a little bit of a joke, but there's a little bit of truth to it as well. And I'm going to hopefully like kind of walk you through the mindset shift and hopefully explain a little bit less about evals, but like what's actually happening um in the sort of agentic AI space, and then how do we then translate that back to evals. So, when we think about like software engineering as a practice, when we thinking about like how do we measure things, we kind of look at it from the sense that, you know, we're going to start with like this thing is meant to do something.
Um so, we we'd start with a set of examples and write some unit tests. Um we might do like a manual regression suite, which is like, hey, when we do A and B, sometimes C happens, and C is unfavorable. Like, let's not do that. Let's not make people vomit when they put their VR goggles on. Um we could do things like CI/CD pipelines, like make sure that like the thing ships out and works the way it's meant to and intended to.
But mostly like we do things in engineering known as chaos engineering and observability. Um for those of you that are in unfamiliar with the term chaos engineering, it's basically where you're like doing all kinds of random stuff and just breaking it and just having fun with the technology and just seeing where you can stretch it and where you can go. Now, when we apply this to AI and data science space as you we traditionally know it in the last little while, 2025 included, uh we do things like static benchmarks.
Like we have these like evaluations. It's like, oh, um I'll give an example. It's like my how how compliant is my AI in like risk? I'm going to ask it a bunch of questions and make sure it doesn't talk about, you know, selling me some financial services cuz that's a big no-no. Um we will then handcraft like a set of questions and examples and sit there and like tune this thing up, make sure it's absolutely perfect um before we deploy the the AI system or the the models, we'll do some like some sort of offline evaluation where we're just kind of cycling through those tests.
But we're missing that sort of chaos engineering space. We're missing that, you know, like what comes next and how do we mess up with it and how do we know where we can stretch this thing? And I think that's like a an honest gap that we see in the space. And that's why we're just so hyper-fixated on benchmarks and evaluations. If you go to any AI conference in the academic space, all people talk about is like benchmarks.
I created a benchmark for like um at adding numbers and what LLMs think about it. It's like, well, great, but like how is this actually helping me? So, then you end up with like this huge humongous set of like data sets to try and somewhat explain what is happening with your agent until something goes wrong. And it's a matter of time before something goes wrong, and it will, and you're kind of back to the drawing board and trying to figure out what's going on.
And the reason for that is that our AI applications are not static, but we're treating them like they're static software. Um yes, when we ship software, we might change unit tests. They're a little bit quicker to do, but realistically speaking, even software is becoming malleable. Um so, flip to my keynote I gave uh yesterday where I'm one of the core contributors of something called Open Claw, the harness changes itself.
Like the harness will shift. Like you want to create skills, you want to do other things. Like it will adapt, right? So, that adaption that we're seeing inside of things where software's being shipped at lightning speed, how does your benchmarks keep up with that? Like how does your benchmarks adapt to that space? Um this is one of many papers that are out there. Um I don't remember when this one was published, but this concept of like adaptive testing for LLM evals, this concept is like somewhat revolutionary, maybe, but like what happens if our benchmarks would change with our with our applications?
Uh I didn't write this, but, you know, great that someone did. But it's just kind of poses the question that like why are stat benchmarks static? Like why don't we test in a more sort of adaptive manner? So, this could be great. This is more like selectively testing and just being a bit more smart about how we test. But it's still like, you know, taking us in that journey. I think it's like a mindset shift. Now, rewind to like what we're seeing in the AI space for a minute.
Um we had prompt engineering if we like focus purely on LLM uh space. Um we had this prompt engineering world where it's like, hey, I'm going to like doom scroll, uh wordsmith instructions. I'm going to just like bash random words into an AI and hope it improves. So, if I'm building this like banking app or creative app, I'm like going to stick all kinds of random words and see what comes out the other end and makes it creative.
It's a little bit akin, and I'm not trying to downplay medicine in any way, it's like, hey, I'm going to make medication for like uh I don't know, liver disease, and turns out it cures pain. Oh, okay, these are painkillers now. That's great. And the same thing we're doing with like prompt engineering, we're just like bashing words into it and hope it changes. And for some reason this died in like 2023, but people still do it.
It's kind of in- in- in- intense. Um and then we kind of went into this like world of context engineering, and I think this started making evals a little bit more relevant because it was a little bit more complicated, you know, there was steps involved. There was like data coming in, and search was like a thing. And we're starting to steer the agents in a direction with things like rag and tool calling. And the beautiful part is there's like, well, okay, I'm a I'm an organization of this big agentic system, maybe I can break this agent up into its parts.
Maybe I can go, oh, I have this MCP tool that does like some sales agent thing. I can test that thing is doing what it's meant to. Right? I can just go off and like be sure that that thing is happening. So, this come on this this process of like, you know, tool calling and and kind of breaking this larger agentic piece up into its its some of its parts made evaluation somewhat steerable and a little bit more understand, but it still didn't kind of hit the the hit on it on its head.
Um but then in 2025, like where are we going next? I mean, if you look around, we we can see that code is cheap. Um that's not a changing thing like tokens are available. Debatable if you think tokens are cheap or not, but, you know, they're they're there, which makes tokens cheap. And then tokens become fast food. Essentially, we can consume more tokens, therefore we generate more. We we the velocity of creating software and applications increases.
Um and models become really good. I think this is the thing that I think I think a lot of people just have not yet comprehended that a lot of the AI applications that are now running, the models can do absolutely amazing things. Um I've been working a lot of like optimization problems, and we can take these models that are somewhat like seen as like these generic systems and be able to do like amazing things like solve ARC-I2, which is like puzzles.
And if I recommend anyone who's like interested in evals, look at the ARC-I2 puzzles. I've tried the ARC-I3. Some of these puzzles are like really hard for humans to solve, but like machines can like pattern recognize. An LLM can actually pattern recognize and and start to solve those. So, what that brings us to is like intent engineering. Um this kind of concept of like machines can self-optimize based on intent. Right?
And we're seeing this with the harnesses that we're seeing coming out where, you know, we've got this with like Open Claw, but we're also getting this with like other types of harnesses inside of Claude and Codex where it's trying to understand you, and it's trying to adapt to you, and give you a better experience. Now, the problem with this is that when we have intentful machines, the evaluations become even more complicated because it's like, how do I know my experience is different from your experience and different from someone else's experience?
Like how do we start to build testing around this sort of methodology and understanding? And I think the complicated part of this is that it it just kind of exacerbates the the the kind of need for evaluation even more. Um there was this kind of joke that I was saying earlier that people say, oh, evaluations are dead, they're going to go away, observability's dead, they're going to go away. But realistically, now more than ever, people want to know what's happening inside of these agentic applications within their different layers because then it gives them some understanding of what's what's going on.
Like we we use words like, oh, these these agents are insecure or um we're not sure what's happening. So, how do we actually how do we actually turn that into something meaningful? So, going back to my earlier slides on kind of this concept, let me just let me just recap where I was for everyone. You know, we said there was this like we had these static benchmarks, we hand created evaluations, we would do like these offline evaluations and we had this big gap.
What we're actually moving towards is like this intent-based outcomes. So, if you if we think about like there's this this concept of like intent engineering that that I'm I'm mentioning, like how do we actually map that to something? So, instead of saying, you know, 1 + 1 = 2 or user asking this very specific question and this is the answer and this is what we're going to compare towards. It's like how do we deal with how do we define ambiguity in agent?
How do we define personality inside of an agent? And and how does that look like for an organization? And some of the research is showing things like, oh, we can build rubric. We can do it like how we how we, you know, evaluate art pictures and and things like that in in in schools. We can self-curate suites from traces, as in not me, but the agent can, you know, once we start tracing these applications. Let's just say 80% of the time it's the same stuff that's happened my agent.
But now suddenly my customer base has changed. And because my customers have changed, they're going to start looking at things they're going to start asking questions differently. Things are going to start changing inside of my agent. But why are we not measuring this? Like why are we not taking these traces and feeding them into agents and going something has changed and then telling that to the user or telling that to the owners of these agents and changing the the the suites, the tests.
We can do online always-on evaluation optimizations. So, like to that point, once we start looking at the traces once we have agents doing the evals not static benchmarks, we can have this like as an always-on sort of service. And then lastly, we can do this sort of like telemetry in the loop. I have a paper that's been written on this, which is essentially when we when we're writing software applications or MCPs or anything like that, agentic systems, if the harness is aware of the telemetry it's aware of like what's breaking, it's aware of like how much it's costing and you can set some conditions around it, it can kind of self-correct itself.
So, we're starting to see this with harnesses where, you know, it's had an error, it's had an issue, it's going to fix itself, it's going to continue on. So, I think this is a kind of a case of like instead of trying to predict what's gone wrong, like how can we be more smart about using that data back into the agent to be able to kind of make it heal itself to some degree. So, I'm kind of calling this like a calcification problem, like the or eval calcification.
I'm I'm still stewing it in my head. Sounds like a really nice paper title. But this idea that it's just going to like become harder and harder unless we can get some kind of smart about it. And I think one of the kind of concepts I want you to kind of stew on and think about. This is like one of the auto research outputs that you can do if you haven't tried it, like Kapathy's auto research. There's really basic sort of auto optimization using Python.
You set a goal, you set a target and it kind of tunes itself, it tweaks itself. You could do this with absolutely anything, right? You could do this with like, I don't know, what's the best mix to to to what's the tastiest barbecue or the cheapest barbecue mix that you you want to make it could be anything, right? You just set a reward signal. But the the the the key here is that your users are going to have a a point of intent that you want to sort of optimize towards.
And then how do we sort of get the machine to like correct itself and and kind of look towards that as an eval. So, then our evals don't become the data set or the starting point, our evals become like what is the end state that we want to get to and then we just let the machines do the work. We have evaluations where it's just the agent and we're just defining the end state. So, just one last thing to sort of bake into your minds before I'm going to finish a little early so we can do questions or anything else, it's fine.
Is that you can imagine this space where you go like 80% is like the static stuff, like it's been defined in an intentful manner, but that 20% is always going to keep changing and it's that 20% that's going to mess up your business. It's going to be someone who's going to come and ask a weird question or use your agent in a really strange way and it's going to be absolute hell for you. So, how do you kind of create agents to kind of manage and maintain that 20% and keep an eye on it and then adapt and change your evals?
So, I think people need to start looking at the evals not as this like static data set thing, but actually as like code as like software or as like a a living agent, not as a point in time, but as like a self-optimizing growing solution. Um that's more or less it. I was going to present more of like an in-depth demonstration of this where we've applied it at Comet, but like the the end state is not quite finished yet and it will be over the next coming weeks.
But I wanted to kind of give you guys something that's like not a sales pitch and something that you could kind of conceptually map to a to the problem space in your own worlds as well. So, if you're working with software, you're working with AI agents, I think you need to start realizing that like the agents will start to shift by themselves. The problem space, the data sets will change and you also need to kind of treat this problem with an agentic mindset as well.
Thank you. >> [applause] [applause] [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.