Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 2:02
2.5x that video's typical replay level
are re-bumbling around the work itself. And the important question here becomes a lot less about what is your title and more what part of the system can you own? Now, I like this taxonomy quite a lot.
Said at 1:56
The graph counts replays. It does not show where viewers stopped watching.
Words
3,240
Runtime
19:48
Speaking pace
164wpm
Reading time
14min
164 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] >> This is a clinical note an AI wrote from a real consultation. Take a few seconds and read it. It reads like a routine headache. A new headache, likely tension type, take some paracetamol, come back if it doesn't settle. Looks completely fine, doesn't it? Here's what's missing. In the room, she also mentioned her jaw aches when she chews. A new headache, over 50 with jaw pain on chewing, that's giant cell arthritis. And untreated it can take her
82 words, the words spoken in the first 30 seconds at 164 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 250 |
| Average words per sentence | 13.0 |
| Longest sentence | 49 words |
| Questions asked | 9 |
| Sentences containing a number | 6 |
Most used terms
Filler phrases
28 in total: actually 8 · like 7 · sort of 4 · kind of 3 · um 3 · uh 2 · you know 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] >> This is a clinical note an AI wrote from a real consultation. Take a few seconds and read it. It reads like a routine headache. A new headache, likely tension type, take some paracetamol, come back if it doesn't settle. Looks completely fine, doesn't it? Here's what's missing. In the room, she also mentioned her jaw aches when she chews. A new headache, over 50 with jaw pain on chewing, that's giant cell arthritis.
And untreated it can take her sight within days. It's a same day start steroids now emergency. And that one line, it never made it into the note. On the page it's paracetamol headache. And nothing in the note is technically wrong. It's the dangerous part is what isn't there. And so that's what I'm going to talk about today. The dangerous failures are often the ones that actually look completely fine. Firstly, who am I?
I'm Seb, medical doctor by background, and now I'm Composo, where we build AI evaluation systems for high-stakes domains. So, that one was a subtle kind of error, but sometimes it's not subtle at all. A man in his 20s sees his GP for a sore throat, tonsillitis. The AI writes that up. It gets him chest pain, suspected angina, diabetes medications he's never taken, and an address for a hospital that doesn't exist. And I I really like the LLM for this one.
I think it's it's a good attempt at a hospital name. Um and weeks later he's invited to diabetic eye screening for diabetes he doesn't have. That's genuinely a real case that happened recently. Obviously, these kind of crazy ones someone notices, but it's those quiet ones that sit in the record uncalled that are the most challenging and can actually do a lot more damage. And they're not rare at all. In the largest real-world study of these notes, about 1 in 20 carried an error that was serious enough that it could cause significant harm to the patient. 1 in 20, that's not theoretical in testing, that's in production on real patients.
And that's only the serious the ones. If you widen that lens to all errors, nearly 1 in 5 had an important omission and more than 1 in 10 had a hallucination. And AI is being deployed at scale across healthcare fast. Ambient Scribes are one of the leading cases, already in about a third of US practices and climbing. Physician AI use doubled last year and none of this is tracked. So, for most of these systems, there's no adverse event reporting at all.
The errors never show up as incidents, they just sit in the record. So, errors this common that are going unseen is is quite hard for me to believe that it's not already affecting patients. It's not that we checked and it's fine, it's that we're flying blind. And this isn't just a healthcare problem, it's every high-stakes use of AI. Healthcare shows it more viscerally because here being confidently wrong can be life and death.
But, everything I show you can map straight back onto other domains as well. So, here's what I want to do. I'm going to show you what exactly is going wrong, why it's going wrong, why the systems we built to catch it don't work, and a suggestion at how maybe we can start to fix that. So, first, what's going wrong and why? So, LLMs are getting good, obviously. They don't make stupid mistakes anymore most of the time. So, it's not about dumb errors.
Everything here came out of three of the best production Ambient scribes on the market. Ones that we all know. We generated a load of notes across them last week. And this is exactly what's going on right now. This is every failure we found. Each dot is an error colored by type. Left to right, how much it matters. Bottom to top, whether a strong automated check catches it. And that split is the point. A handful up top get caught.
But almost everything sits below the line. The ones I care about most are these on the bottom right. The high stakes and missed ones. Let me show you what a couple of those looks like. So, a woman comes in with a headache. Doctor asks, "Did it come on suddenly or build up gradually?" She says she doesn't know. It just happened. The note records that as abrupt sudden onset. And sudden onset is a red flag. You can see why it just happened could maybe be interpreted and inferred as abrupt onset.
But, that's a feature that points to a bleed on the brain. She never said it. The model decided it. And now that one word drives the whole workup. Here's another. Doctor suggests running some tests. Patient says, "Can we just try try antibiotics instead?" They agree, hold off on the tests, treat, and see how it goes. Note records the opposite. Arrange tests today. It kept the plan that they talked out of, not the one they chose.
Every line in the note reads fine because it's not really a hallucination at all. It's not wrong. It was there in the original, but it's just not what they ended up deciding. So, why are these happening? There's, you know, in ambient scribes, there's first transcription and then generation. A lot of it does happen on the transcription layer. It can be words misheard for their sound-alikes. So, Humalog heard Humulin. Two insulins on completely different timelines, so swapping them could crash a blood sugar.
Hyperthyroidism becomes hypothyroidism, the opposite condition. Or a drop to no on uh no evidence of cancer that becomes evidence of cancer. So, these these are really hard problems, and they are common. Not the ones I'm going to focus on, because most of what goes wrong is actually even with a perfect transcript. It's the model reading the words correctly and still doing one of three things. Either it adds something that was never said, it changes something that was, or it omits something that should be there.
Now, the blatant version of each of these is is really easy to catch. The hard part in all three is the same. It's telling whether that thing that was added or changed or dropped actually matters. It's detecting that slight over inference versus the dangerous fabrication. The harmless rephrase versus the meaningful edit. A dropped line of small talk versus a dropped allergy. So, the ones that matter slip through along with all of the ones that don't.
That call which different matters is taste, effectively. Not aesthetic taste, but essentially judgment. It's It's whether in this context a missed allergy might kill someone or is not important. And I think there's there's three properties that really matter about this. It's tacit, so your domain experts have it, but they can't fully write it down. It's contextual, so the same detail is critical in one note, noise in the next.
And it's moving. The model changes, guidelines change, two good doctors disagree, different hospitals have different definitions. So, there's no fixed target to write down. And so, the model knows the facts, ultimately. they're extraordinarily capable but what they lack is a sense of what matters here for this specific example and that's why even brilliant models make these mistakes. So one natural move you're never going to make that generator perfect and generator is cheap generation is cheap so stop fixing it at the source.
Let it write put a checker after it pass only what clears the bar. And that checker should be the easier job the generator has to get everything right and it pay attention to lots of varying instructions. Whereas the checker only has to find the one thing that's wrong and just focus on that task. You can also give it more time more tokens the exact failure modes to hunt for. Evaluation should be easier than generation.
It's the asymmetry of verification verifies law that's why AI is raced ahead anyway you can cheaply check the answer maths and code. And doing this is exactly what the best teams do they put a lot of energy into evaluation. It starts with the gold standard which is expert humans reviewing notes which obviously works offline but you can't put a human on every note in production. So they automate it. They build a serious system and some of the best versions of this that I've seen are you take the transcript and the note and context put in front of the judge a detailed rubric for faithfulness with worked pass and fail examples.
The rubric maybe auto optimized with GPA or something like that maybe you have some deterministic NLP to sort of count up medical concepts that are differing between the two. That's a powerful system and yet I pulled all of those errors earlier out of ambient scribes in an afternoon. So if the evaluation is this good how are these errors still getting through? So I built this system and ran those same notes through it.
And it scored most of them fine. It flagged a handful of them and signed off the rest. >> [clears throat] >> But one in five of those clean passes still had some sort of serious error buried in it. And often that was an omission. The things that should have been there and actually quietly weren't. And that's the best version of a judge I've seen in a lot of teams and it waved them through. Why did it do that? It's not stupid.
It's a frontier model, serious engineering behind it, more than clever enough to read the whole encounter and catch every obvious error. And it's not blind, either. And and that's part of the trap. If you take a note that says start amoxicillin, when the real decision was actually to wait and see, it's faithful to the words, amoxicillin did come up, but it's a lie about the intent. A good judge might catch that, might.
But whether it flags that versus the other doesn't have other things that it could comment on depends on it knowing what decision matters most. And so it's not blind, it just can't tell what counts, essentially. So the note passes confidently and you put a judge like that in front of your system, you've not added a safety net, you've added a second silent failure that just nods along with the first. And here's the root of it.
So in math or code, the verifier comes with free. A unit test, a compiler. But for is this note safe and complete, there's no unit test. You have to build the verifier yourself and verification is only easier than generation for the easy bit, i.e. spot the difference between transcription note. But that's not the hard bit. The hard bit is knowing of all those differences you've seen, which matter. And that's harder than writing that plausibly good general note in the first place.
Because that standard of good was never written down anywhere that the judge can read it. A rubric that you pre-specify is only the taste you could write down. The taste that matters is the part that you couldn't. And so here's here's a bit more detail on what what matters looks like. Two patients, both with blood in their urine, both notes dropped the same kind of line where they'd been on holiday. One had been to France, the other to Lake Malawi.
Same English emission, same shape, same mistake. Well, not really, because blood in the urine obviously warranted away and you're going to have to investigate it, but the France trip is irrelevant. The Lake Malawi trip is the diagnosis. Fresh water in sub-Saharan Africa means schistosomiasis until proven otherwise and it completely changes what the management plan is. So that same dropped line in one note is pure noise, in the other it's the answer.
And which one it is, you simply just can't write all of that down in advance. So if you can't write it down, you can't write taste down, how do you get that into your evaluator and your whole application system? Well, we've answered a version of this before. RLHF exists because you can't write the reward function for good. You learn it from examples by showing it. The only question is where you keep what you've learned.
And there's three places. You can either specify it up front, you can stuff the prompt, write the perfect rubric. We just watched that fail essentially. You can bake into the weights, fine-tuning or continual learning, but for a standard that's still moving and a score that has to be explainable, the weights I think are the wrong place to keep that. They go stale, they can't tell you why and you can't change them without a retrain.
So there's the third option, which I'll show you, which is you essentially just keep the taste as the examples themselves. Past judgments, expert corrections, references, and for each output, you retrieve the ones that bear on it into the judges context, add one and it's live on the next call. You can point at exactly what moved the score. For this problem, it's both better and also cheaper to do. So, that's the way to do that is one repeating loop, three steps.
Discover the failure modes from real outputs, capture how your experts judge them, calibrate every output against that, and when the standard moves, the loop moves with it. So, in more detail, discover. You don't write that rubric in a vacuum. You have to put the system in production and look at the real outputs. Cluster what goes wrong and the failure modes surface on their own. You name them. This is your failure mode ontology.
Discover from your data, not guess on a whiteboard. And you can't shortcut it. The ways that a real system goes wrong are effectively unbounded and synthetic test cases only cover the failures you already imagined. The ones that hurt you are often the ones that you didn't. And you'll only find those in real outputs. So, this ontology is your map, what to capture judgment on, what to retrieve against, including the failures that you never thought to check for.
After that, it's capture and then calibrate. So, those discovered modes, they're not a checklist that the judge runs, but they organize everything. What you What you ask your experts about, how you index the cases that you'll retrieve, and capturing is a simple part. You put real outputs in front of your experts. Clinicians spend a focused few hours leaving comments. A session doesn't have to be a month-long labeling project to start with.
And you collect their judgment. Not just a score, but the reasoning and corrections. And over time, you build up that record of how your experts actually judge. You then calibrate. That's the the the generic part of this you can write down once easily. For example, be faithful or don't drop anything important. But what you can't write down is what counts as a serious miss for this specific note. That's contextual. And it shifts from note to note.
So, what we recommend is you assemble that on the fly. For each output, your judging agent pulls in everything that bears on this one case. It's memory of the most similar outputs that it's judged before and how they scored, the expert corrections that apply, the reference documents and guidelines. Just context engineering per output. And crucially not just one pre-specified rubric in a vacuum, and not a model that you have to retrain every week, but a full sort of case-specific standard assembled for this output.
And it's a loop as well. Every output you judge, every correction, sharpens the next. And when a brand new failure mode appears, Discovery surface it, and it flows straight back in. And so, to make that a little bit more concrete, that headache that I opened with, the one that was really a possible blindness emergency, here's the kinds of things that you would want to pull in for that note. The nearest cases that your experts have judged, not this exact patient, but the same shape, maybe a red flag filed as routine.
Uh the corrections that apply, like a new headache over 50, um suggests something that you need to check red flags on, and some criteria and guidelines, and you pull all of that in. It hasn't memorized this case. It's a capable model. And handed the right context to reason from, held against that, the dropped red flag stands out. It was never actually hard to catch. It just didn't know what mattered. And so, if you take that same data set of generated notes from the start and pass it through these three judging systems, the first, a strong off-the-shelf judge with a rubric frontier model, um it's better than a coin flip, but it misses most of what matters.
The second, that sort of serious system that we talked about before, rubric, deeper, maybe some detonate stick checks, better again, but still missing quite a lot of what counts. The third, the judge running this loop, discovered failure modes, calibrated for output against what experts judged, is performing a lot better on this specific data set. Same notes. The only thing that changes is what the judge was shown. And the difference here, it's not more compute or a better prompt, it's that the first two fight taste and lose.
They guess the criteria, they freeze one standard, and they go stale. This repeating evolving loop does the opposite. It discovers the modes, fits the standard to each mode, and keeps learning. So, you might not write chemical notes, but if you ship anything where being confidently wrong has a cost, the contract review that misses the clauses that change the deal, the support agent that promises a refund you don't offer, the same thing is true for all of those.
It's watched, if at all, by a judge with no taste for what matters in your domain. So, three things. Discover your failure modes from real outputs, don't guess them. Capture your experts' judgment on them, the standard they can't write down. Calibrate every output against the cases that they've already judged, not a static rubric, not a retrained model. Then keep that loop running. And if you take one thing away, easiest place to start is your experts leaving free-form comments on real outputs.
That's the real That's the raw material for everything else. Your judge can verify anything that you write down in advance, but the standard of good never could be. And so, stop trying to write it all down in advance, and just start capturing it case by case, and evolving it. That's why evaluation can't be a thing you build once and freeze. The standard it checks against doesn't exist on paper. It has to be discovered from real outputs captured from the people who hold it and kept alive as it moves.
Evaluation isn't something you have, it's something that you do continuously over time. Thank you. >> [applause]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.