Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 13:49
2.3x that video's typical replay level
Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well, you know? It will usually give you some context about what happened.
Said at 13:43
The graph counts replays. It does not show where viewers stopped watching.
Words
4,087
Runtime
20:02
Speaking pace
204wpm
Reading time
17min
204 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
All right, hi everybody. Um let me just try to stand straight so I don't have to crouch over. Um thank you all for coming. This is our talk on evaluation at scale for everybody. Um I hope everyone's in the right room and if you are, thank you for coming. We were expecting like 20 people so this is like way more than what we expected. So all right, who are we? So I'm Nick, I'm a product manager in Kaggle Benchmarks uh and I basically run and build our Benchmarks platform alongside a couple of our engineers and I also
102 words, the words spoken in the first 30 seconds at 204 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 222 |
| Average words per sentence | 18.4 |
| Longest sentence | 86 words |
| Questions asked | 19 |
| Sentences containing a number | 14 |
Most used terms
Filler phrases
263 in total: um 88 · like 75 · uh 51 · you know 24 · actually 10 · kind of 10 · basically 5.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
All right, hi everybody. Um let me just try to stand straight so I don't have to crouch over. Um thank you all for coming. This is our talk on evaluation at scale for everybody. Um I hope everyone's in the right room and if you are, thank you for coming. We were expecting like 20 people so this is like way more than what we expected. So all right, who are we? So I'm Nick, I'm a product manager in Kaggle Benchmarks uh and I basically run and build our Benchmarks platform alongside a couple of our engineers and I also focus on our AGentIC Eval solutions.
I'm originally from Singapore but I live in the San Francisco Bay Area and so I flew in to do this this talk and attend all the great talks out in this conference today. >> And hi, I'm Michael. I'm a software engineer on Kaggle. I've been working at Google for about a third of the time that I've been alive and Kaggle for about half of that. So uh yeah, but mostly working on evaluations and benchmarks for Kaggle at the moment. >> All right.
Um has anyone here heard of Kaggle? Put your hands up if you have. Okay, great. Um so a lot of people know us for competitions um but we don't just use that. We're the world's largest AI ML community of 30-plus million users and we've been working a lot in the GenAI Eval space over the course of the past 2 years um and we think there's lots of interesting problems in the industry that not many people are trying to solve and we feel like we're positioned well to solve them and we want to share more of the work that we've been doing and also invite contributions if you want to get involved in the space.
Um so we have a simple agenda for today. First is AI Eval today are kind of broken and we'll talk about why. And then step two is we're trying to solve it, not saying we're the all cure and we have everything um kind of sorted out. We'll talk about what we're trying to do, the challenges we're running into, and also maybe like it might inspire some of you uh in terms of how you think you might be able to help contribute to this very important problem that we're trying to solve for.
With that, I'll jump into the first section. Um so, first problem, evals are scattered, decentralized, and get stale fast. I don't know how many of you have tried to keep track of like AI benchmarks, but basically like 10 plus of them drop every single day, and the best way to find out what they are is go to archive and spend like hours scrolling through them, reading every paper. That doesn't make sense. Um we don't think it makes sense.
I can't even do it even though it's my full-time job. And I think what happens after that um paper gets published is that, you know, you see some of the leaderboards in the papers, and what happens after that? They just get stale. The authors move on to the next best benchmark because they just want to publish lots lots of papers, no fault of their own, but these leaderboards no longer become relevant as time goes on.
The second issue is that evals aren't always transparent, accessible, and verifiable. So, many of us have seen these charts uh on these model publisher uh notes when they release a new model, but what's the problem with that? You know, we don't actually know how these benchmarks were set up. There's a lot of configurations you could use for the models themselves, and also how the benchmark is orchestrated and facilitated.
And we don't always know what's actually being tested here. Um I'll give you one real anecdote, which is that we had published a benchmark um with one of these AI labs, and another competing AI lab came to us and said, "Hey, we don't like the results um of, you know, this particular benchmark you published. Let's run it on our own." And so, they ran it, and then they published it with much higher, much better results.
And the difference was that they were optimizing it for their model. So, they had used compaction that they had provided through their API, and we didn't for all of the models we ran. So, the results you're seeing don't always reflect the actual state of things, and that's a problem. Number three, big circle represents all the world and its knowledge, small circles represent AI researchers and technical professionals.
They're like something like 30,000 AI researchers, at least that's what Google AI search told me, and I think there's like 30 million software engineers, data scientists, technical folks. We expect AI to help most of humanity, but then a very small percentage of people are creating all these evals. And if something's not not being evaluated, not being benchmarked, we cannot hill climb on it, we cannot know how good we are at those things.
And what will this lead to? More of these cognitive edges or jaggedness with these models as we're already seeing now, and that's only going to get more exacerbated over time as we see superhuman intelligence in some areas and just like very mediocre performance in other areas. That's not equitable AI that we want that will benefit all of humanity. Um this is a kind of fun anecdote, but not many AI researchers are also wastewater treatment plant engineers.
I give you this example because this is an actual benchmark built by one of our users. He's He lives in Turkey. He's been a wastewater plant engineer for 20 years. I don't know what that entails, but it's clearly a very important job cuz he built this benchmark cuz he cares and he recounted a story where, you know, he's been doing this for 20 years. There have been severe incidents in his country where there was some incident, people didn't follow the safety protocols, and people ended up dying as a result of that.
So, he built this benchmark to evaluate how AI could help him in his job and help avoid these incidents in the future. So, this is like a proprietary novel data set that he's created from his own experience. Doesn't live anywhere else on the web, doesn't live in any of the AI lab focus area because that's not something that's economically productive for them at the moment. So, it's very important um why we think um we should work on open source contributions to the eval space.
All right, um, enough of that. Talking about the solutions ahead, we're working on a couple solutions, but it's tough. So, I'll very quickly cover the first two at the top, and then I'll hand it over to Michael to kind of deep dive into the remaining two products at the bottom. Um, so at the top left, we have hackathons. We have a platform, um, that lets anybody host a hackathon, and I'll talk about why I think it's relevant to the problem of evals.
Second, on the right-hand side, we have agent exams. Um, we wanted to democratize the process of taking evals. We heard a lot about open claw and consumer agents being a big thing. But, the problem is that most people don't care about evaluating their own open claw agents, which is kind of crazy. Um, number three, at the bottom left, we have game arena. That's where we have this evergreen benchmark, where models are playing PvP games against each other.
So, it's, um, Elo score-type rating, and it's forever, um, hill climbable and unsaturated, because they're just fighting against each other, and there must be one winner and one loser. At the bottom right, we have benchmark. So, um, it's the product I run, which is basically a platform that enables anybody to build, run, and share evals, um, to the open community. Um, so very quickly, hackathons. Why I think it's important.
Hackathons are a great way to channel channel people's energy and expertise to solving a problem. I think we've seen that, you know, with the right energy, investment, and time, we can do a lot with very little. Um, I think a great example is the world galvanized over the past three years to make GenAI happen, and this is like a very small form of how we want to help facilitate that process, um, towards solving the right problems in this space.
Um, it's important that we put guardrails around the problems that we're trying to solve, so that people don't go crazy, but also give them enough space, so that they can flourish, and their creativity can show. And the results of everything will be open source, uh, for the benefit of everybody and not just a small group of people. Um but there are some challenges. Um and actually before I dive into the challenges, so on the screenshot on the right is a hackathon that we're actually running right now with the Google DeepMind AGI team.
So, Google DeepMind a couple weeks ago published a paper on how we can measure the cognitive faculties um of AGI. And so, we started this hackathon to focus on five particular faculties of the 10. And we want people to build benchmarks in those areas, and we want to give everybody the chance to contribute to AI research and not just a few people. And also knowing that everyone has something unique to contribute that these AI labs couldn't do so themselves.
Uh but running a hackathon platform isn't all that easy as well. Um I think these are fairly self-explanatory. Maybe I'll talk about um the second and the third one in particular. Um the second one being that we need to provide them the right tools in order for them to do their best work. Um it might sound trivial, but things like, "Okay, if you have a thousand participants, um everyone's operating globally online. How can we give them the tools like hosting their own data sets to have access to AI models?" That's something I never thought about before, but a lot of people, you know, coming from uh poorer background, they might not have money to pay for these five API keys to access all these state-of-the-art models.
And how can we then let them share their work in a way that's understandable through write-ups um so others can see, perceive, understand, learn, and build on the work that they've been doing. And then the third point is that as much as, you know, AI agents and everything you hear about this conference um are very good at a lot of things, not very good at like judging innovation and creativity. So, a lot of the work still requires human experts, and even alignment within experts is difficult and not trivial.
Um and that's something that we have to facilitate as part of this platform, too. Um the second thing that we're working on right now is what we're calling standardized agent exams. I originally called it like SATs, standardized agent tests, uh, but you know, there was a trademark issue, so I had to change the name. Um, yeah, it's a true story. Um, so how it works is that you just paste a one-line prompt your agent, and essentially it takes an exam, and we return a score for you on a leaderboard that you can compare its performance against.
Um, this was a very experimental MVP that we just launched last week. Um, and I think this is important because if you look at AIEvals that people are doing, it's kind of two, um, ends of the spectrum. You have like research labs and enterprises, you know, using gray trust, using all these state-of-the-art technology to set up to measure their agents and their models. And then on the other end, end, you have consumer agents.
People who are building open claws, and then, you know, filing 1,100 security advisories this morning, as we heard. Um, but most of them aren't actually testing their agents before they're sending them out to the real world. Um, which I think is a huge problem, as we've seen, and will become even more important. Um, so like one conversation that came up this week was, how can we maybe do more safety-focused exams? So you can do a quick baseline of your agent before you send it out into the world to run your inbox, to run your Amazon accounts, and to do stuff for you.
Um, so I talked about the first point. I think the second and third one are quite interesting. Um, on the second one is that when something's accessible, we want to make sure that it's also, um, challenging enough. And so we have this spectrum of if we make something too difficult, people can take the exam, but no one finishes it cuz it runs for too long, it's too difficult. But if you make it too easy, then it doesn't give you the right signal for what you want to measure.
And then finally, the chart at the bottom shows maybe there is a market for agent consumers. Um, we only launched a week ago, and we have, you know, hundreds, like 500 plus agents already evaluated on our exam, um, without us even really promoting it very much. Um, so I think that's um an interesting insight that we've gleaned from this experience. On the right hand side on the screenshots that you see there, um so we posted in Mopbook and then we started seeing these weird kind of spin-off posts about people sharing agent sharing their exam results and even like an SAE prep course um that came up on Mopbook.
Um so, you know, that's always interesting to see. Um with that, I'll hand over to Michael. >> Yeah, thanks, Nick. Um yeah, so I really selfishly wanted to talk about Game Arena and benchmarks with y'all, AI engineering, um and mostly kind of give an overview of how it works, some of the really cool things we've seen in it, but mostly want to talk about the challenges that we've been having with them. And so, please come find me at the DeepMind booth after this to talk about them.
I would love like some more insight on these kind of things, too. So, uh for Game Arena, it's a So, benchmarking platform. One of the problems with benchmarks is they very quickly get saturated. Uh we see this with community benchmarks, we see it with AI and like researching benchmarks, as well. So, Game Arena is an approach to help us work against the saturation by just having PvP and so you can never have saturation because you'll always have one model able to compete against others.
So, saturation might just be for a while a model is the best. Um really quick engineering slide here of like how is all this set up and how does it work? So, when we're trying to figure out games to put into Game Arena, we want to analyze like separate capabilities of AI models and so we try to pick good varied games. So far, the ones we've invested a lot in are Werewolf to like uh play around with what's best at deception, uh poker for like the randomization, also some of the deception and how good uh like Grok loves to go all in on poker.
Uh Less models are a little bit less crazy, a little bit more conservative. Most interesting, some of the newer generation of models are worse at poker because they are more risk-averse. And so, you just see these personalities uh start to emerge over time. And then, chess because anytime you're analyzing ML things, you have to be analyzing chess. And so, uh for the quick uh just overview of like how all this works, we design and iterate on a game, figure out a good game we want to do, make sure models can actually play it, and then spend a lot of time on iterating prompts to try to make sure that our prompts are fair.
And uh this is all like open source and like uh very viewable for like what's happening. So, if you want to check it out, I put the GitHub link there. Really all of the things that we've talked about in this presentation so far are live on Kaggle. A lot of them are open source, so please come look at it, play around, give us feedback. We love it. Um so, after that uh we work on building this harness. Mostly we've done open spiel games so far, which is like an RL um framework.
Uh and so, like we test like are these models better playing at random? Sometimes they are. Usually they are, but like sometimes not that much better. Are there any other interesting properties that emerge? Uh finally, we end up running the simulations. Uh we use LLM model proxy. This is actually available on Colab if anybody uses that to just like talk in a consistent way to all of the models that we want to run these games against.
Uh and so, it runs on top of the Kaggle simulation platform, which was like initially an RL platform that we had for Kaggle before LLMs became a huge things. Uh we schedule game runs. Uh it uses Bradley-Terry pairwise to try to like not have too many games we have to run. I'll talk about that in a second. Um and then finally, we publish the results. We have all these LLM conversations. We stick them in a data set that's available on Kaggle.
People can check it out and um like learn things from that. Uh we put it on the benchmarks to show the Elo scores and then we also have a game visualizer for all these things, so you can go and, you know, see Grok all in on poker hands. Uh that's a little demo to the left of that. Um so, now as promised, some of the really big like challenges that we've had with this. And so, love to hear all y'all's ideas and like different ways to approach this.
Um you can imagine it gets very expensive very quick. I'm sure you all have seen um you know, your Claude 4.6 bills or something along those lines. So, you can imagine that for poker in order to get statistical significance, we had to run about 400,000 poker hands. Um and there's many turns inside of each of those hands. You can imagine what those bills start to look like. And so, the Bradley-Terry pairing um is like part of this, but any way that we can get statistical significance without having to run millions of games uh is great.
And like we're always trying to think of new ways to like be able to be sure that the models are best at the things that we're claiming they're best at, but um you know, running as quickly as possible. Uh it gets a little boring to just watch like LLMs play against each other all the time. Like for some of the initial games, it's pretty fun. Um but you know, as we're trying to build this out, um it might get a little bit repetitive.
And so, trying to figure out ways to engage Kaggle's community to be able to participate in this process. Um and so, you know, even things like could we have a hackathon that somebody provides a prompt for, uh for example. And then like, you know, we would have like prompts given by our community as part of the competition to like play these games and see who prompts like the model the best and like climbs on a leaderboard.
And then comparison over time is difficult. Uh old models disappear, new models come along. Sometimes when you're talking to a model endpoint, if you're not talking to yourself, they're not exactly honest on what model is happening in the background. Uh so, that's always a little bit of a difficulty. But yeah, go check it out. Pretty fun. Um so, yeah, moving on to benchmarks with my limited time here. So, um what is this not?
Is this not like a production evaluation platform? There's plenty of people talking about that over at more if you want to go and run this for your production code. Very cool things. This is much more about community involvement. Anyone can like build and run and share evals in uh hopefully open and verifiable way. Uh for time's sake, I'll just talk about this really second, but we basically it looks very similar to the production evaluations platform where you write some assertions.
So, like this thing worked or you know, walks uh wet or it's a dry. Like you can check, okay, does this contain a towel? We also do LLM judging similar to the production platforms. Uh these all get grouped together in a task. They then get evaluated against a collection of models that the users want to run against and then all these tasks get aggregated together in a benchmark such as the wastewater treatment one that Nick was talking about previously.
Uh so Paige who just presented in this room before this actually made this a nice little task for us. It was parsing an SVG uh from XKCD and say can you recreate this SVG? And so you can kind of see the code for this is over on the left. Um one of the models is a little bit outdated so Sonic 4 created this uh nice reproduction below it. Um and then uh Paige created a number of assertions so like, you know, can it generate an SVG at all?
Does it have the correct text? And some other checks and then um an easy way to compare these things side by side. Uh so yeah, this is some of the challenges with this uh that like inspiration incentivization are hard. If you want to like it's not hard for production evaluations platform cuz you're you know, shipping your thing to consumers. They care about like does this model work or not? But to just inspire people in the community to create benchmarks that like other people find interesting.
Uh we've had good luck with hackathons um and like uh you know, uh Kaggle has a points and like metal system and things like that. So we have some things baked in the platform to help inspiration incentivization, but it just takes a lot of work to write a good evaluation. Uh for agentic benchmark execute like oh yeah, um so when we started this people were really interested about analyzing just models. As we've moved more onto what are agents doing, it gets really hard to figure out like what we're actually testing against.
Uh so I pulled out this little thing from Morph LLM uh paper or like blog post that they published on March 16th that basically called out that you know, against SWE-bench Pro the six frontier models are within a couple of percentage points of each other. Definitely like go check out this blog post. I haven't specifically verified it, but it does seem likely to me. The thing that really matters a lot for coding performance is what harness is it running inside of with like a 22% difference depending on the harness.
And so, that can get, you know, really tricky of like are you testing the harness? Are you checking the model? Things like actually ambiguity under test. Um and then again, fast release and deprecation cycles and models, it gets a little bit tricky um to figure out what we're or like to be able to do comparisons over time. Uh so, yeah. I think that's about time, but as I mentioned, we'll all be at the D-Mine booth um for the in-between times or you can email either me or Nick here, but thank you so much.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.