Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 11:58
4.6x that video's typical replay level
issues. Uh I also invented OS certification. I just close the tracker whenever I want, so I have my life back. So, does this work? Yes, sort of. >> [laughter] >> Which leads me to act three, slow the down. Everything's broken.
Said at 11:52
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
The graph counts replays. It does not show where viewers stopped watching.
Words
3,102
Runtime
16:45
Speaking pace
185wpm
Reading time
13min
185 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[music] Hello everyone. Uh welcome to uh don't be data poor. Uh my name is Anoj. I uh I lead AI at Interior. Um just a bit about Antior. We are a clinicianled AI company um built for health plans uh backed by SEOA and NEA. Um and what we do is we run AI transformations for health plans. Um as part of which we build um agents for uh several high stakes uh healthcare administrative workflows in production uh things like pri authorization, payment integrity, uh heat measures etc. Um it's it's
93 words, the words spoken in the first 30 seconds at 185 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 120 |
| Average words per sentence | 25.9 |
| Longest sentence | 213 words |
| Questions asked | 9 |
| Sentences containing a number | 6 |
Most used terms
Filler phrases
282 in total: uh 167 · um 38 · like 37 · actually 13 · kind of 9 · sort of 9 · right? 7 · I mean 1 · basically 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
[music] Hello everyone. Uh welcome to uh don't be data poor. Uh my name is Anoj. I uh I lead AI at Interior. Um just a bit about Antior. We are a clinicianled AI company um built for health plans uh backed by SEOA and NEA. Um and what we do is we run AI transformations for health plans. Um as part of which we build um agents for uh several high stakes uh healthcare administrative workflows in production uh things like pri authorization, payment integrity, uh heat measures etc.
Um it's it's okay if you're not familiar with any of these workflows uh because a lot of the work that we do can actually be summarized uh in in the same way. It's uh policy guided decision-m over highly unstructured data. And the unstructured data looks something like this, right? You have uh uh you have these scanned fax bundles containing medical records full of patient information. Um a not so fun fact uh is that I think around 70% of medical communication still happens via fax.
Um and fortunately or unfortunately this is the data that we end up working with the most. It is a very uh rich uh and information dense uh data that we see here. Um the data distribution here is uh it comes from a very long tale of rare uh cases with very nuance scenarios. Uh it models an entire clinical trajectory for a patient. Uh and every single person's journey is very different. Uh it also presents itself in uh varied formats.
So you have like uh things like bad handwriting, tables, checkboxes, um key value pairs, images, um a lot of tough data to deal with. But I I personally think it's a very fascinating source of data that we see here. like it's it's it's like sort of like an observation through a very fuzzy lens over an entire person's lifespan. Uh it's really unique and I'm sure you must have heard this like enough times today already but uh in healthcare uh the the baselines for accuracy are just uh exceptionally high.
Uh 95% is not good enough. Um at interior this is why we invest very deeply in data sets and evals. uh and these unstructured medical records are are a staple source of uh data for these evals and we we work with this kind of data in almost every workflow that we try to automate but the problem is we can't really keep this data u it's phrase it uh we can't even derive information from it and most of our contracts prohibit us from uh from doing anything like that um even things like uh redacting it anonymizing it and keeping derivative copies like there's a strict no completely off the table.
So nothing really survives in any sort of data set that we want to persist over a period of time. [snorts] So so what this talk is about is like what do you do when the data set you most need is also the data you're least allowed to keep and the and the bet that we the answer that we put our bets on is that we can kind of synthetically generate this data ourselves. There's been a lot of focus on synthetic data recently. uh you've like frontier labs uh uh striving to generate synthetic data for continued pre-training for RL for computer use for agents uh so it's it's it's a hot topic and it's it's a hot topic on our minds as well uh and the moment you say generate like the first thing that comes to mind is okay can we can we try to use an LLM to generate synthetic data and I think you can I personally believe LLMs are a fantastic tool to generate synthetic data and several teams have already demonstrated uh this already there's been some papers uh in the healthare space outside the healthcare space people have successfully used uh LLMs to generate synthetic data for for different purposes there are some known challenges in trying to use these LLMs uh uh to create data especially we're trying to oneshot the whole process it's really hard to generate diverse realistic looking synthetic records and this is even more of a problem uh when you're trying to doing when you try to do this at scale so u often times these medical records uh are over 300 pages long and it's like imagining if you wouldn't ask an LLM to write a novel for you in one shot, right?
So, it's the same reason why you wouldn't use an LLM to just one shot in synthetic record for you. Um, and LM seem to suffer from this very strange mode collapse problem when it comes to generating like diverse uh data, creative data. And I think there's two main reasons for it. The first one is uh like Aish mentioned in his talk earlier, there's very little exposure to this data source in the pre-training data corpus.
Um, and today's objectives for pre-training and post- training are are largely uh they're only they're not incentivized for creativity or diversity really. They're incentivized to be helpful assistance. So with these challenges in mind, I'll walk you through like one of our approaches in how we uh manage to build a pipeline to generate synthetic data. Um earlier I mentioned our our forward tasks look something like this, right?
You have workflows and tasks that start with some unstructured data and a policy um and you execute your policy against that data. You follow this reasoning trace through it u and you arrive at some sort of an outcome which is your label. So this is our forward task. Uh and the idea we had was to try and reverse this process. uh can we actually sort by sampling a random label um figuring out a reasoning trace for that label and then trying to generate data backwards from that.
Uh the idea here being that if you can actually uh sample these two things uh with enough diversity u we will have we will be able to generate data that's conditioned on diverse set of inputs allowing us to kind of circumvent the diversity problem a little bit. Uh so just a quick aside on policies. We've talked about policies a bit but uh let me just clarify what these really mean. Right? So this is an example policy we have for for a CPAP device for patients.
Uh this particular one is for a medical necessity review workflow. Uh and it it sort of outlines all these diverse set of conditions uh that a patient might have um in which a CPAP device should be approved or or rejected. Um so and this policy as well as many other policies you can think of these as uh essentially decision trees that outline all these sorts of conditions um um that that dictate how some outcomes are met.
And at Antier actually we we spend a lot of time and energy in trying to model these policies explicitly as decision trees. We work with symbolic representation similar to decision trees and it helps us achieve a better accuracy and consistency score when executing them in LM based workflow. And and the reason why I'm bringing this up is that uh by having this sort of symbolic representation of a policy, you actually have a way to kind of deterministically sample different reasoning traces for a given outcome.
So back to our idea of like reversing the process, right? uh this this sampling of reasoning traces from the policies uh what helps us get that diverse conditioning input to then generate medical records from uh and the key idea here is that the distribution here uh that we sample from is is a much more uniform uh and effective prior distribution than what you'd normally get from an LLM and one added benefit of sampling this way is that in theory you're able to test uh for far more scenarios than you would likely get from production data sources.
So what I mean by that is like say you get a sample of 200 cases from your customer uh um and and and you try to like have an eval that measures performance against that and you get a 95% score uh it doesn't really tell you about uh what you what your performance would be in those rare edge cases that are not in that data set. There'll always be rare edge cases uh that outside that distribution just because of the fact that our data is so uh highly variant.
So for those for those familiar with Cynthia like uh they follow a similar pattern uh of sampling scenarios from a symbolic causal state representation there's a few other folks in the space who are working with these symbolic representations to uh to generate diversity and synthetic data generation. So let me walk you through the rest of the pipeline. All right. So uh once we have this diverse set of samples as our conditioning input, what we did was we built an LLM based pipeline that uh follows uh a course to find pattern to progressively uh build up a medical record layer by layer.
So here we first start with creating some patient invariants like the biological sex, the birth date, the blood group. uh we use that along with our reasoning trace uh with NLM again to produce an ordered list of uh events and provider encounters that a patient might have had and we call this the patient journey. So this is a high level uh you can think of as a high level uh overview of what a patient might have gone through in their lifespan um captured by a list of events on a high in natural language and in the real world it is actually only during these uh encounters provider encounters that documentation is really generated at least for the data that we get uh most of our data source data is generated during these provider encounters.
So we model exactly that in our pipeline. uh uh we first generate a document plan for each encounter and then based on that and the preceding history of the of the patient. We uh we fan out into generating the actual documents uh um to hydrate them with actual synthetic information. Uh and this course to fine layering uh is actually what allows us to keep uh the different prompt payloads in the pipeline very token efficient from both input and output perspective while also enabling uh we this also helps us enable to scale across uh longer patient journeys.
So you can scale this uh pipeline. You can have a much longer patient journey uh and you can just fan out and generate documents that way without uh overloading the context windows of your LMS. Finally, we have the sort of refinement loop in the end uh that we that uses a set of eval to provide feedback uh to improve specific parts of the generated documents. Uh for example, one of the eval is an LLM based check for consistency um between all documents.
So this makes sure that there's no contradictions or uh u inaccuracies or conflicting information between two documents that are generated. And this is important because we we have a parallel fan art process uh that is used to generate these documents independently. And because we started with the labels for this particular uh pipeline run, uh what we actually also have is uh an ability to kind of use those labels uh run uh and and compare those against the generated uh uh medical records to see if uh the task that we originally used actually matches is the data is in concordance with the task inputs and outputs.
So we can do this sort of roundtrip check to ensure that our data is actually in sync and by default get correct labels by construction. So um in theory uh this is a really nice property to have like you can basically skip the ground truththing expensive ground truthing process you need uh for for data for for your data. Uh one thing to clarify here is that uh so far all the generation has been happening just in plain text and markdown text.
Um it is possible to go from that to a rendered PDF. uh but we don't really see much value in in doing that uh because uh we have state-of-the-art PDF parsers today. They're available to everyone and they just allow you to convert any sort of complex PDF into a nice markdown representation. So all of the synthetic generation uh um and evaluation happens in the text domain. So this is just an example of like a pipeline that we created from scratch and it's uh it's very easy to build.
It's largely fully LM based. Um but but who came up with this, right? like who who am I to know anything about what a good um medical record looks like? U so how do we know if this is any good? And I think this has been mentioned a few times today already but like you really don't like uh no AI engineer would ever would like you want your domain experts to be the ones telling you what's good, what's not good. Um and which is why we believe that uh it is of great value to empower your domain experts to own your whole data pipeline.
And specifically we uh we do this in two ways right uh we enable our clinicians to kind of interject at each point u in the generation process with a human in the loop mechanism. So at any point a clinician can steer the generation process to make a medical record in the way they want it. Uh we often see our clinicians use this uh to to first look at cases that happen in production, get some interesting ideas uh and then use that uh use those ideas along with this uh the steering in this pipeline to make cases that look similar to what we might see in production or they have seen in production.
And this is what makes the data generated from this really useful, right? Like you can actually model your uh your failure cases beforehand or even after they after you see them in production. And secondly, I think most importantly, we let our clinicians also own the whole logic of the pipeline. Um, we do this by modeling the whole pipeline as a skills-based workflow running on a generic uh agent harness that we built internally.
So, every every uh kind of section here you see uh all the way from the patient journey to the document generation to the document enrichment to the evals all of these things are skills uh that run on our agent harness. As an example, if a clinician wanted to say maybe add support for a new document type, let's say for a new customer, they wanted their intake forms uh to look a certain way, uh they could easily just make a new skill file for it, uh attach it to the pipeline, um and and and voila, there wouldn't be any engineering changes required.
So, it's completely clinician own from that perspective. And just an aside, generally, I feel like skills are really an amazing interface between AI engineers and domain experts, especially in vertical AI. uh we see this being um we we see this being modeled in several of our other workflows both for internal use cases and in production as well. So some results from this right so even though we only really use uh synthetic data for evaluation at the moment there's already a lot of merits that we get from it.
Um roughly 90% of our data sets are already made of synthetic data. Uh this helps us uh maintain a very high uh production accuracy score um for across many customer deployments. Uh the pipelines that that I just showed you uh already we are able to achieve a a very high fidelity on this generated data. Uh in a blind review clinicians were not a were only able to distinguish uh synthetic from real about 60% of the time.
So room for improvement but uh uh but but it's but it's it's close and I'm I'm quite quite it's quite promising avenue for us to invest more more here and and the the the the fact that is the most interesting to me and uh what I really what I'm really excited about is that all of these data sets well most of our data sets today then are created just in time for these customer deployments right you can you when you have the ability to like create data from scratch so quickly uh you can kind of uh you don't need to depend on on on waiting for data from a customer you kind of just model all your edge cases, simulate them and test your workflows before you go live with the production class go live in production.
So some takeaways if you're looking to build your own synthetic data pipeline in healthcare or even another domain um try reversing your inference workflow. Diversity should always be sampled from a from an appropriate distribution for your use case. uh try to emulate the process in which uh uh the data was actually generated. So like I showed you we were trying to sort of like we were using LLMs. We're trying to emulate how uh our medical records might actually be generated during patient encounters.
Uh so and I I would highly recommend you try doing that. Uh and the fourth most important thing I think is uh when you're when you're making a data pipeline like this, it's really important to give your domain experts the keys because uh these are the people who know uh about your data and and and they will help you uh drive towards a recursive self-improvement and not the engineers. Cool. So you don't need a PHI problem for this.
Uh anywhere uh the data you need is ephemeral, sensitive or even expensive to label, you can think about uh generating data yourself. Uh, and hopefully you won't be data bored. Thank you everyone.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.