Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Words
5,084
Runtime
26:06
Speaking pace
195wpm
Reading time
21min
195 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
[Music] Hey everyone, Eric here. Today we're going to learn about Baze rule and some probabilistic thinking and stuff. Okay, so here's a thought question. Imagine I told you that I was going to interview someone uh for this class today. Uh what would the probability that that person would be a man? Okay, if I say, "Oh, yeah, and by the way, the person I'm going to interview is 7 feet tall." Now, what would you estimate the probability that that person would be a man as? Okay, so you probably estimated in the second case that
98 words, the words spoken in the first 30 seconds at 195 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 404 |
| Average words per sentence | 12.6 |
| Longest sentence | 65 words |
| Questions asked | 64 |
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
[Music] Hey everyone, Eric here. Today we're going to learn about Baze rule and some probabilistic thinking and stuff. Okay, so here's a thought question. Imagine I told you that I was going to interview someone uh for this class today. Uh what would the probability that that person would be a man? Okay, if I say, "Oh, yeah, and by the way, the person I'm going to interview is 7 feet tall." Now, what would you estimate the probability that that person would be a man as?
Okay, so you probably estimated in the second case that it was much more likely the seven foot tall person was a man than a woman. Whereas in the first case, you probably said it's 50/50 chance, right? It could be one or the other. Uh, and this is what's called BA rule, right? What you're updating, you're updating your belief based on some new information. Your belief is in the population of the world, people that are my interview, half men, half women.
But in the list of people on my interview who are 7 feet tall, there's a lot more men than women. Okay. So, basically what you did, here's a graphic depiction of what you did. There's a bunch of people in the world, half of them men, half of them are women, but some of them are over 7 feet tall. And most of those people, most of the red people in this are men because that's just the way it works. Uh, that's Baze rule.
So I'm going to teach you a formalized way to do that so that when you get new information you can update your information. Okay. So this is the important thing. Sometimes we have a belief about the probability of something. That's cool. And then we get new information. You can and should update your belief, right? And we see people doing this all the time. If you go to a game, you're like, "Oh, our team's going to win.
Our team's not going to win." But then like you get like 10 minutes beforehand and your team's down by 70 points, you're like, "Yeah, we're going to go home." You know, because you've updated your belief about whether or not you're going to win and you're like, "Oh, but if the score is tied and there's 10 minutes, no one leaves, right? This is B rule." But you should update beliefs and this is how to do it. Base rule is a proper way to do this.
So you might say, "Oh, what's the probability the visitor is a male given this new information?" Well, what's the probability I have a disease given the test result from the disease? What's the probability someone's the right person to hire, given what their resume says? There's a lot of different places you could use this. But basically, anytime you have a probabilistic belief, this may happen, this may happen, or this may be the truth, this may be the truth, and you get new information, you should update that.
Okay? So, we have two different things. We have the test, which is the new information. It's not the truth, it's just the new information, right? Then we also have the truth which is the belief that we want to know about. Right? So the truth is whether it's a man or a woman coming to visit. The test is whether they're over seven feet tall. Right? That's the new information. And here's an important thing. Tests can be wrong in two different ways.
One, the test could say positive. So could say that the truth is what you think it is, which is um good, I guess. Or the test could say the the the thing that you don't think is true is not true, which is also okay. But it can be confusing if the test says yes, but the truth is no. Or the test says no, but the truth is yes. So let me give you example. You guys have heard of COVID probably. Most people are unhappy about this, but it does a lot for my class because people understand the concept of false positives and false negatives now, right?
So, if I actually have COVID and I get a test that says that I don't have COVID, then that's a false negative because the test is negative, but I actually have it. On the other hand, if I don't have CO, but I get a test back that says I do have COVID, that's a false positive. So, there's two different ways that this test can be wrong. It can tell me I have the disease that I don't have, or it can tell me I don't have a disease that I actually do have.
So, keep that in mind. Tests can be wrong and they can be wrong in two different ways. And that's what we need to to keep track of. Okay. So, let me give you a question here. Okay. Um and and these are real numbers um from a guy named Ginger for symptom-free women aed 40 to 50 participating in screening using mamography. And this is important because there's recommendations of when you should get these particular tests.
And at the time this this data came out, uh the recommendation was 40. Okay. The probability one of these women has breast cancer is 0.8%. If a woman has breast cancer, the probability is 90% that she will have a positive mammogram test. If a woman does not have breast cancer, the probability is 7% that she will still have a positive mammogram test. So that's a false positive. What's the probability the woman who has a positive test actually has breast cancer?
It's kind of a big deal because if people are getting screened, they're going to be getting positive tests sometimes and then they need to know what does this mean? What do I do about this? Because it's unpleasant situation to be in. And so you can use B rule which says the probability of A given the probability of B is the probability of B given the probability of A times the probability of B divided by the probability of A.
Or don't panic, you can do it a different way. Okay, change all these probabilities into frequencies. Okay, so there's B rule. Okay, you don't have to actually do it that way. Instead, do it this way. For symptom-free women aed 40 to 50 who participate in screening using mamography, the following information is available. Out of 1,000 women, eight will have breast cancer. Out of these eight, seven will have positive mammograms.
Of the remaining 992 women, 70 will have positive mammograms. What's the probability the woman with a positive test actually has breast cancer? And you can see that it's seven divided by 77. it becomes much much easier to do it if we convert it to frequencies which is exactly what you did with the original question of the 7 foot tall visitor. Okay, so I'm going to show you a formal way to do this. We're going to use trees and they're going to branch off to keep everything clear, but you see how much easier it is when we convert all this stuff to frequency instead of trying to use that big um formula.
Okay, so we have several steps in here. I'm going to write down the steps. Okay, first we're going to pick a number of people. Then um we're going to start by dividing those people between the two truths we're interested in. Has a disease, doesn't have the disease, is a man, is a woman, is a good hire, is a bad hire, whatever. We're going to divide it between the truth that we're interested in. Okay? Uh based on our current belief, then we're going to connect um with lines and label everything.
Why do we want to connect everything with lines and label it? Well, we want to make sure that it's readable and understandable to ourselves and to others so we can go back and double check and make sure we put everything in the right place, put the right numbers in the right place, and that we can share this information with others because you're not always just doing this for class or doing this for yourself. Sometimes you're doing it to show other people what the right way to proceed is.
Okay. Um then we apply the test to the two different groups because in general a test will give different probabilities of a positive or negative based on which group you're in. So if you have breast cancer, your probability of a positive test is different than if you don't. If you're a man, your probability of being seven feet tall is different than if you're a woman, etc. Okay. highlight the ones that um the ones with the test result you're thinking about because in general you have the test result that people are seven feet tall and you want to compare just those people and then you do some easy math which I just showed you.
So let's go through one. Okay, so 10,00 women we already picked the sample size and the truth is that 992 of them don't have cancer and that eight of them do have cancer. Okay, that's the truth. Everything's labeled. Everything's divided up in between those two groups. Okay. Then we apply the test for the women who don't have cancer. A positive test. Still 7% of them are going to get positive test. That's 70 people. And uh the rest of them are not going to have a positive test.
So that's uh 922 people. For the women who do have cancer, 90% chance. I'm just rounding here. 90% chance. That means seven of them will have a positive test and one of them will have a negative test. So if the question is what's the probability that somebody with a positive test has cancer, you want to say okay what how many people actually have a positive test and cancer divided by all the positive tests. So that's seven divided by 7 plus 70 which equals about 9%.
Okay or 0.09. Let's go through the six steps again. What are the six steps? First you pick a number of people. Then you start with the truth. the thing that you have a belief about and you divide the people according to that. This is the thing you actually want to know. You connect the lines and label them. Then you apply the test to the two groups also with labeled lines. Okay? Then step five, you highlight the ones that have the test result you're thinking about because this is the whole point of BA's rule.
Now that we have this additional information, we can throw some of the people out of consideration, right? And then number six, you just look and see how many that have the truth also have the test result you're interested in divided by all the people that have the test result. Okay, so let's do another one. So imagine you have to go take a lie detector test. Okay, now I made up these numbers because no one's quite sure, but imagine that the probability that a given person is lying is 4%.
So in other words, every sentence that comes out of your mouth is either true or false. Imagine that four percent of them are lies and 96 of them are true. If you tell a lie, then the lie detector test will be positive 99% of the time. So what does positive lie detector test means? It says that you're telling a lie. Okay? So I' I've seen a lot of cartoons. So I'm pretty sure that what happens is a red light comes on it goes.
That's a positive lie detector test because it's a lie detector, not a truth detector test. Uh if a person does not lie, then the lie detector test will be positive only 10% of the time. Um, which means 90% of the time a green light and a ding will come on. Okay. So, if a lie detector test indicates a person's lying, like I said, hey, did you kill that guy? And I say no. And it goes, what's the probability that I killed that guy?
All right. Let's go through the steps. First, let's say that you're going to ask me a thousand things. Okay. In 960 of those things, I'm going to tell the truth. And in 40 of those things, I'm going to lie to you. Out of those 960 things, 96 times there's going to be a that says I'm lying. And out of the 40 times I'm actually lying, 39 39 of those is going to go and tell me I'm lying. So out of all the times that it goes and says I'm lying, what is the probability that I'm actually lying when that happens?
And you can do the math here and it comes out to be about 30%. So, if you ask me if I killed a guy and I say no and the lie detector turns red and beeps, that means I probably didn't kill the guy, right? Anytime a lie detector test tells you that somebody is lying, that means they're probably telling the truth, which is why we don't use lie detectors in court because they're terribly unreliable. But here's an interesting question.
If the green light comes on and goes ding and says I'm telling the truth, what's the probability I'm telling the truth? Well, now I can look at the other branches of the tree where I have truth. And that becomes >> 864 divided by 864 + 1, which is a 99.88% chance that I'm telling the truth. So, these machines should be called truth detectors because they're actually pretty good at that, but they're terrible at detecting lies, which again, this is why we thought they were a great idea and then we stopped using them in court altogether because all you have to do is hire any statistician to come in and say, "This is stupid.
We can't tell. If it said that they were lying, that means they're probably telling the truth. And if it says they're telling the truth, that means they're really, really, really probably telling the truth. Plus, we also don't actually know what the positive rate is or how often people lie. So, um, they're very bad tests. Okay, let's do another one. Um, this one's personal to me. So, um, when I was in high school, we all graduated from high school.
And, um, we all sort of went and did our thing. I had a friend who wanted to join the army. And one of the things you have to do when you join the army is you have to get a bunch of tests, including at the time time they call it an AIDS test. I call it an HIV test. But so, he had to go get this test. And he got this test and the test came back and said that he had HIV or AIDS. Um, again, it was the 80s, so we didn't know what the virus was called.
Um he was very sad as you can imagine because in the 80s that's basically a death sentence. Um the disease had just just been invented actually I guess it was invented in the 60s but it just became popularized. It was highly stigmatized and 100% fatal at the time. So he was very very sad about this. Um, as it turns out, he went not he went to his own personal doctor and uh got an other test and that test said that he didn't have the disease.
And um so he proceeded with his life. It's been 30 or 40 years. He's shown no symptoms. So I'm assuming that the first test was wrong and the second test was correct. Um, but it dramatically changed his life because he didn't join the army because they would have given him a second test and he said, "There's no way I'm taking a third test. I'm just going to go with this first test because I don't want to get another uh positive result." So, here's an interesting question, right?
This test is actually very very accurate in the sense of specificity and um sensitivity, which means the probability of detecting if you have it and the probability of detecting that you don't have it if you don't have it. Specifically, these are real numbers. So, for HIV test, um, if you get a positive test result, there's two tests you actually have to get a positive result on the Western Blot and the ELISA test. The probability of a true negative is 99.99%.
So, in other words, if you don't have the disease, there's a 99.99% chance that they will tell you you don't have the disease. Um, on the other hand, the probability of a true positive because this is a pretty good test is 99.99, sorry, 99.9, right? So, if you have it, it's 99.9% likely to tell tell you you have it. So, about at the time about one in 10,000 people had it. I think that's probably a similar number today or maybe less.
But at the time, about one in 10,000 people had it. So, let's use Bay rule to see if you get a positive test back, what the probability is that you actually have it. So, we're going to pick a number of people. In this case, I'm picking 10 million. When you pick the number of people, pick whatever makes sense. In this case, because there's so many decimals, and I hate decimals, I'm picking 10 million so that there won't be any decimals.
If you're doing something where you're doing eggs, pick some multiple of 12. I don't know. But in this case, we're picking 10 million. But any number you pick, you get the same answer. You just have a bunch of decimals. Okay? So, if we do random testing on 10 million people, and this is important, this random testing, right? If you have symptoms, that's a whole different thing. Similar with mammograms. You find a lump, you go in, that's a whole different thing.
But they're not recommending that only people who have a family history of cancer or anything like that get tested. They're recommending everyone get tested after a certain age. So that's random testing. Same thing for the army. You're going to join the army. They're not saying, "Hey, you know, what's your sex history like? Do you use intervenous drugs?" They're saying, "No, we're just going to test everybody that joins the army." Okay?
So 10 million people try to join the army. 1,000 of those people actually have HIV. One in 10,000. Um, and 9,999,000 of them do not. Out of those 10,000 people uh that have it, because the test is 99.9% accurate, 999 are going to have a positive test result and one of them is going to have a non-positive test result. So, that would be a false negative. It's negative. says you don't have the disease, but that's false.
So, it's a false negative. Out of the 9,999,000 that actually don't have the disease, only 0.00 I'm sorry, 0.01% of them will have a um positive test. But because there's 10 million of them, that also comes out to be 999 people will have a positive test result that don't have the disease. So this is a false positive. It's positive, but it's false. And then 999,980 blah blah blah will have a negative test result and life will be fine.
So what's the probability that somebody with a positive test result has HIV? Turns out to be 50%. Because there's an equal number of people. So here's how to look at it. there's only a one in 10,000 chance that the test gives you a false positive, but there's also only a one in 10,000 chance that you have the disease. So, there's an equal chance that you actually have the disease as to giving a false positive. So, obviously, if you get a false if you get a positive, then you're like, well, it's a 50-50 chance.
Either I have the disease, which is very rare, or I have a false positive, which is exactly the same amount of raress. And in like I said in my friend's case, it's been many years and he seems to be fine. Um, but this is a really really big deal that people don't understand this and don't know how to update it. So there's this guy named Ger Genarinszer who uh has done some studies where he sent his uh his his graduate assistants out to see what HIV counselors actually do and what they will tell people.
So this poor guy had to go out and get I think he got 20. Yeah, it was 20 HIV tests and it took um almost a year to do this. Why did it take a year for him to get 20 HIV tests, you think? Well, because it's very important that it's random testing. So, he has to tell the people that he has no high-risisk behaviors. One of the high-risisk behaviors is intervenous drug use. And if he went out and got 20 tests in one day, his whole arm would be full of needle marks.
And they would say, "Well, hold on. your risk, right? Because we're updating based on BA's rule. This is some new information, right? Um, if you're an MS MSM, men who have sex with men, then that makes it higher. If you're intervenous drug user, that makes it higher. If you're a hemophiliac, that makes it higher. So, there's a lot of things that can make it higher. So, the graduate student had to go on and say, "No, I'm a um a straight male non-intervenous drug user in a long-term monogous relationship." Right?
So that that's tells you what the probability of actually having the disease is anyway. So this poor guy had to go get these 20 tests. Um and the people didn't do well. This is in 1998. So he sent all of the testing centers. He's German as you might imagine someone nameer would be. Um so he sent all the people all this stuff, published these papers and said, "Hey, your HIV counselors aren't very good at this. So here's the truth of the matter." Anyway, he redid it and got a new graduate student that went out in 2015 to go to the HIV counseling centers that he had already done this to years before, you know, 15 years before and sent them emails and and and letters.
Well, it was letters because probably not email in 199 1998 um just to see if they got any better at it. So, here's the paper. You don't have to read all the paper, but I do think that that's a picture of the actual guy who um went out and got the the testing. Anyway, so here's some here's some clips from the paper. What occurred was um out of the 32 counselors, 12 were physicians, 17 social workers, and two social education workers, and one nurse.
Okay, so 12 of these 32 people actually had medical degrees. They were medical doctors. There was a nurse, a couple of social workers. Okay. And um one of the questions this guy asked him when he was getting his test is um if I test positive, how likely is it that I have the virus? 29 counselors provided incorrect information out of 32. Notice that 12 of these guys have medical degrees. 18 of these said that it would be about uh 100%.
Um but basically all these people told this guy that if he got a positive test result that uh he was going to die. Well, who's gonna have HIV? Because I guess we don't necessarily die of it anymore. There's a little table here that I like to show you that has um a little table that has some of the actual responses that people have. So, there's a column here that says numerically correct. These are all the people that gave them the correct answers.
And numerically incorrect, which you can see it's all there. Um, as I said, 100%. As I said, 99.9%. Uh, yes, as I said, the test is 99.9% certain after the window period. As I said, 99.9%. As I said, absolutely certain 100% so to speak. Now, notice the as I said keeps coming up. So, he also asked them what the sensitivity and spec specificity of the test was. In other words, what's the probability that the test will tell me that I have HIV if I have HIV, which is 99.9%.
Um, but then he asked them, what's the probability that if I have a positive test, I have HIV, which is a completely different question. All these people told him I already answered that question. Right? This is a really important thing that I want you to understand. The probability of A given B and the probability of B given A are very very different numbers. Right? If I say, hey, we're going to have a guest who's a man, what's the probability he's over seven feet tall?
That's super tiny. But if I say I have a guest who's over seven feet tall, what's the probability he's a man? That's really, really big, right? It's the same thing here. And usually I want to know what the probability of the truth is given a test. Not the probability of a test given a known truth. Okay. Uh so use B rule apply it. Use these sorts of trees. Oh, here's another one. I was searching for information on this article.
I went to Google and guess what Google says? It says if you get a positive test result, there's a 99.99% chance that you have the disease. So hey, thanks Google AI for telling people this. This is a really big deal, right? So here here's an example that came from Gird's book, right? This woman got a positive test. And so again, this is in the 80s. And so she was sad, I guess, is an understatement. She gave away her children and went to go live in this halfway house where she could just go to die with other people who had this disease.
Um, the 80s weren't pretty. Anyway, um, as it turned out, it was a false positive test. But in the halfway house, she met this guy and that guy did actually have HIV and they got together because what's the point if they both have it? They're not going to give it to each other unless you don't actually have it. So, she gave up her children and then she accidentally killed herself because of this positive test. So, it's kind of a big deal that people should understand how this works.
So please make sure you understand that the probability of A given B and the probability of B given A are very different things or could be very different things. Okay, let's summarize. You hold a belief about the probability that something's going to happen or something's true. Okay, then you get new information. You can and should update your belief about your probability that is going to do and we naturally do that.
But B rule is the right way to do it, right? So, if you want to say, "Hey, when should I leave to avoid parking?" You could apply B rule. You don't have to wait until your team's down by 70. You could say, "Hey, look, there's a 99% chance that we're going to lose if we're down by five at this time. Let's get out of here." Um, okay. Tests can be wrong in two ways. How are those ways? You can get a false positive or a false negative.
They should actually be named backwards. It should be a positive false and a negative false because it's a positive result that's false or negative result that's false. Right? So positive result says you have something, but if it's false that means you don't actually have it. So the test says you have it, you don't actually have it. It's a little confusing. Think of them backwards. Okay? The probability of a disease given a positive test is radically different than the probability of a positive test given a disease.
And this doesn't apply just to diseases and stuff, right? Say we want to let somebody into to college, for example. We say, "Hey, you got to take an SAT." Well, that's a test. And then we say, "Is this person going to be a good student?" We don't know, but we have some probability of that. And then we get this SAT score and we add that SAT score to that. And we integrate and update that data. Right? So, anytime you're getting new data about something, right?
You hire a consultant to your business to say, "Should we do this?" And they say, "Yes, that doesn't mean you should do it. That just means you should update your belief about whether it's a good idea or a bad idea. Okay? So this doesn't apply just to diseases. Also, you should hold beliefs as probabilities because a lot of times you don't really know what's going to happen. So if you hold beliefs as probabilities, then you can update those probabilities easier instead of saying, "Oh yeah, definitely it's going to be a guy that he's going to interview or definitely it's going to be a woman he's going to say it's 50/50 and I'm going to update it.
He could have a really tall woman coming in. I don't know." Okay. So hold beliefs of probabilities. Update your beliefs based on new information. Use B rule in the form of frequencies. That's all I have. You can practice some more later. Thank you. Bye-bye. [Music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.
| Sentences containing a number |
| 87 |
Most used terms
Filler phrases
134 in total: um 42 · actually 27 · right? 26 · uh 18 · like 11 · basically 4 · you know 3 · kind of 2 · sort of 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.