Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 2:02
2.5x that video's typical replay level
are re-bumbling around the work itself. And the important question here becomes a lot less about what is your title and more what part of the system can you own? Now, I like this taxonomy quite a lot.
Said at 1:56
The graph counts replays. It does not show where viewers stopped watching.
Words
980
Runtime
6:06
Speaking pace
161wpm
Reading time
4min
161 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> Awesome. Well, hey everyone. My name is Aparna, one of the founders of Arize. We work with some amazing teams to help them build evals. Um, and we have an incredible lineup of talks for you all today at the evals track. Um, it's happening in room 2005 and there's going to be amazing speakers from Term Bench and Uber and Snorkel kind of all happening after this. Um, but today I'm here to talk to you about the
81 words, the words spoken in the first 30 seconds at 161 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 61 |
| Average words per sentence | 16.1 |
| Longest sentence | 43 words |
| Questions asked | 2 |
| Sentences containing a number | 6 |
Most used terms
Filler phrases
30 in total: actually 13 · um 5 · kind of 4 · I mean 2 · like 2 · right? 2 · literally 1 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> Awesome. Well, hey everyone. My name is Aparna, one of the founders of Arize. We work with some amazing teams to help them build evals. Um, and we have an incredible lineup of talks for you all today at the evals track. Um, it's happening in room 2005 and there's going to be amazing speakers from Term Bench and Uber and Snorkel kind of all happening after this. Um, but today I'm here to talk to you about the future of evals.
Evals have gone from the new skill that every PM and every AI engineer has to learn to the thing that every serious AI team is betting on. We've been really fortunate to get to work with some of the best AI teams in the world. So, we get a front row seat into not just what's happening when they're building their actual agents and before they actually ship, but actually the evals that teams are running on their live production agent via their traces.
Little bit of some stats for you guys. We run over 100 million evals every month. The average team runs about 12 different eval jobs with the top teams running over 3,800 different evaluators. And offline evals, online evals, they each have their own place, but today what I'm actually going to talk to you about is the teams that are running evals on their traces. This is actually what's helping teams figure out what's working, catch their failures, and that's the type of data you need to fuel your continual learning loops.
And the industry kind of agrees. I mean, all the CPOs of Anthropic, OpenAI, all you know, GDB, you have Garry Tan saying, "Evals are everything you need." And the whole industry kind of agrees. So, we added evals, they catch all the failures, right? Here's the problem. When we were building all of these first-gen evals, the thing that we were actually evaluating has changed underneath us. In 2023, it was about just answering a prompt.
In 2024, we started to see all the frontier models. They've added tool calls, they've added reasoning, they've added deep research. Now, what we have is teams running loops on real-world data with sub-agents kicked off on long-horizon tasks. Every one of these was actually a massive jump in complexity, and we didn't just make the problem harder, we actually got a fundamentally different type of problem. What that meant is that as these systems got more complex, so did the way that they actually fail.
We're really lucky cuz we have our own agent that we've built, Alex, that lives in our UI, and we get our kind of get to feel this pain ourselves. Every time the frontier labs added new functionality, we added it to our agent. And now Alex can has much longer memory. It has the ability to create dynamic UIs. It can go search across an enormous volume of traces. But, we also realized that it would forget context. It wouldn't know when something was done.
Um sometimes it would just get stuck in these loops. And the key thing here is that the classical LLM as a judge evals, that probably many of you have written in this room, just weren't for us to be able to catch all the types of failures that we were experiencing. I mean, it's just fundamentally different, right? You have a deterministic flow, and now what we have is literally every time a user interacted with Alex, it would create a new UI.
That's a fundamentally different trajectory. So, this led to our really big revelation. What if the best way to an evaluate an agent was actually with an agent. Doesn't mean that all of the ways that we did evals, with deterministic evals, with LLM as a judge, classic evals, doesn't matter anymore, but it just means that we have a different type of tool to solve a different type of problem. Agent as a judge is about adaptive dynamic analysis.
LLM as a judge just gives you a fixed rubric with these fixed scores. It's what everyone's doing, but when your agent's doing completely different trajectories every time a user puts in data, it just means that you need a fundamentally different type of eval. My take is that most teams today are doing the first two, but the future of evals is actually having all three. And today I'm actually excited to share we've released agent as a judge to help our teams on their eval journey.
We've released signal. Signal's actually a long-running agent that can read traces sent in, discover patterns of issues. Um, it can figure out types of problems that a classical LLM as a judge eval just would never be able to do with these deterministic rubrics. It's helped us figure out very subtle failures that you wouldn't even think of doing, such as something going on in a loop for multiple times, it was calling the same tool for repeatedly long time, the trajectory was inefficient.
And actually what this does is because it has all that analysis, it can go put up a PR and put up a fix. So, if you want to learn more, come to our come to our booth. We're right by the OpenAI booth. We'll give you a demo, we'll show you a bit more about it. We're also, like I said, taking over the evals track, so come to room 2005. We're going to be talking a lot about the future of evals and what they look like. And if you just want to hang out with our team, we're throwing a viewing party for the USA World Cup game tonight, so check out the Luma and register to come join us.
Awesome. Thank you all so much. >> [music] [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.