Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

iTeach · @iteach1056
Words
18,804
Runtime
2:29:08
Speaking pace
126wpm
Reading time
78min
126 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Hello and welcome back to ST 118 statistics and applications session 9. In this session I would like to discuss with you the subject topic of the case when time uh is involved in regression. Right? We all understand that time is money and time drives everybody's uh lives and all businesses. And so it turns out that it also drives a lot
63 words, the words spoken in the first 30 seconds at 126 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 1,004 |
| Average words per sentence | 18.7 |
| Longest sentence | 307 words |
| Questions asked | 232 |
| Sentences containing a number | 150 |
Most used terms
Filler phrases
835 in total: uh 451 · right? 172 · like 95 · you know 42 · um 24 · kind of 21 · sort of 17 · basically 9 · I mean 2 · actually 2.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Hello and welcome back to ST 118 statistics and applications session 9. In this session I would like to discuss with you the subject topic of the case when time uh is involved in regression. Right? We all understand that time is money and time drives everybody's uh lives and all businesses. And so it turns out that it also drives a lot of the data measurements. And since we are always uh measuring data, capturing data and put it into our use, right? uh in terms of finding out the coefficients through linear regression then we should take a look at it.
It turns out time does have some special uh influence on the way we deal with it mathematically or not just that it it is physically driving all the measurements and driving the growth and driving the the withering of businesses and uh data. But it also um needs to be treated with respect. All right? Especially when we are putting time as a variable into our linear regression models. Okay. So next we want to talk about um the models that include time.
All right. Uh applications These might be examples and to see how they play a part. And uh well we would then talk about the auto correlation as a property that is uh causing issues and therefore that we have to pay attention to when we put in time as one of our explanatory variables. Okay. And then of course we want to learn about Durban Watson test or autocorrelation of errors in the linear regression model. Okay. Uh and of course uh we can always look at a an example a simple example.
Okay. And of course all these calculations are all available in open libraries. That means you don't have to pay a fee or software purchase. uh in Python in RAR all of high quality computations high quality in the sense that they are optimized they are fast they are also reliable that is the numbers that come out you don't have to suspect that they might be having some bugs inside and that you know don't take it for granted uh that it's always bug free right especially in the earlier days of the library development so in my days when I was learning stuff uh sometimes times we get a an answer that we ourselves suspect didn't look right.
Turns out there was a bug, right? Because the programmer didn't quite understand the mathematics behind or it wasn't tested fully or it is complicated, right? The math is is complicated. So while 80% of the calculations are done, there is always this 20% part that is always rather tough to get it right or something like that. So um but then uh you know we we do have very high quality uh functions and libraries available uh these days and so to do any of these they are straightforward.
So this example just like all all all simple numerical examples in my videos they are meant for sort of a slow motion frame by frame right uh analysis to see how the numbers move from one place to another so that we at least for once right in our lives get to see how the numbers are being computed and thereby developing a better inner field right of how to interpret the results when they are very quickly uh calculated and then we say okay with this number what do I do right and we then know what to do okay so let's get started um with the idea of uh when time is involved yeah so in our example we uh of of linear regression we've been talking about uh wanting to investigate explain or predict right so uh using the regression model once we collect the data we pipe it into the functions to get the coefficients then we do we look at the the test results right model is good variables are performing and they are significant uh adjusted R square is good and so then we rely on it to do our investigation so we could use the model to um identify which factor is driving that response variable.
We could use the model to explain right the strength right oh this is uh the the the primary reason right we of course the model doesn't explain the reason we are not doing the why here in our linear regression but through our contextual interpretation we can uh attach a reason and say this is of primary influence uh of our y variable this is the secondary. This is the third uh you know factor that that is driving our our response variable.
So when time is involved in regression right we will put in the time variable. So for example perhaps we might want to say that uh health right the health of a person can be explained by sum b 0 plus b1 of the person's weight and income thinking that you know we all know Weight of a person affects the person's health a lot. Right? If obesity is one such a problem, right? So if you're overweight, then it affects your health.
Uh we might also say that well if the person is unwell and you have high income then you can afford to correct your weight and thereby you know uh recover your health or you know improve your health. So we we we look at this as a health index. The higher the better. Weight right and then income. So we are trying to establish this relationship. And uh this might be okay right this might be okay because uh we might have a demographic right in the in the city in the population that uh has varying income levels.
We're not looking at the ultra rich. you're looking at, you know, perhaps the media media middle quartile, right? Uh so there's a good v variety of income levels, there's a good variety of weights, right? And we study this. So that that's probably okay. However, there may be a um particular segment of people, right, where the income and the weight might be correlated. So if you're looking at small wrestlers, so the heavier you are, it seems to be your seems to be that you're paid more, right?
Uh because you tend to be able to win whatever sumo fights they they are doing. So they they really add on to the weight, right? And if you have high income, you can add on to your weight and and so on. So of course, does it make it less healthy for that sumo wrestler? Maybe they since young have been training their bodies to take it. So their health is not too much affected. I don't know. Right? So it what I'm trying to say is that if you are looking at a particular segment then maybe we can't tell so much anymore.
They might be uh related. Then in that case we have multi-olinearity. We learned about that right? So all that stuff all the five assumptions of linear cor uh linear regression still apply whenever we perform linear regression. So far so good. But then here comes the new kid on the block. Right? The moment we add another variable but this time around it is time. Uh right now what happens? Because when time is added, we might think that well as time progresses just just we just you know take the samples and time is a variable just like any other right.
The the the important thing to realize is that time is not really just like any other variables. The first thing that is already very different, right, is that time variables they are always ordered or time five is always larger than time four. Our time here doesn't have to be a calendar date. It is almost always uh kind of coded into integer variables. So with fixed interval measurements into time 1, time 2, time three.
So time five is always after time four and before time six. But when we do statistics, we do averages, we do sum of squares, right? We do all the errors, right? If you recall, that's why knowing the formulas and the mathematics behind is always a good thing. Even though we may not have to be called the details but you must have heard of our discussion about sum of squares and all these right. So the thing is sum of squares summing finding average all these don't care about the order but time data cares about the order and when we put in time here it's the order is being discarded or ignored when the order does matter.
Correct. For example, to find the average income, we just average it up. But the when time is involved, you want to uh it becomes not reasonable to average your future income with your uh part-time income when you were a student, right? And then say your average income in your life is uh $4,000 per month. It's very maybe it's quite a meaningless number as far as anyone is concerned right the average income in your life.
Uh so because it doesn't help that particular individual to plan or do anything right and it's like just jumbling up of the future and the past together and then coming up with a number. So when time is involved we need to be careful with a few things right we need to in uh interpret it carefully we need to uh look at the meaning does it make sense or not right so uh typically time will be added as uh two kinds of uh two kinds of variable so you can add time as a quantitative variable right so which is what I was saying uh time one time two time three that's called a time series uh way I'm not talking about intervals right interval time intervals can is just an interval so that's fine time interval is not an issue but absolute time stamps right so that's that's going to cause the problem time uh time intervals the duration as we call is not so comparable.
For example, you take uh 10 minutes to run 2.4 km, right? I take 15 minutes to run 2.4 km. That's fine. You are better than me, right? You're you're you take lesser time than my time interval. So that's fine. You take lesser duration than my duration. That's that's that's comparable, right? So that's totally no problem. And then I train and train and train. Perhaps I can run I don't know 9 minutes 50 seconds right so it's it's uh you know in a way independent of age independent of the race independent of uh you know you so long as you but maybe highly dependent on training and uh training hours and uh the instructor or whoever right but whatever it is I'm not talking about time duration I'm talking about the time uh time stamp so for time stamp.
We can put in the time series, right? We can put in the quantitative time stamps where you minus the two timestamps, we will get a time duration. But typically we put time series. So for time series then it's always unit interval in between and we have to uh exercise the discipline to always time uh uh capture the data every minute, every hour, every day, you know. So daily sales for example is a timestamped uh quantity daily sales.
So every day we have a time uh every day we have a we capture the total sales for that day stop prices. So uh goes by take t take t take right so if you are at the trading floor you have a trading terminal it gets updated every second then the time stamp is uh having a unit interval of seconds. So that will be quantitatively. You can also add in time variables as a qualitative label. So you can treat it that way. So uh such as the time when it is one, time when it is two and all these are 01 variable.
So you turn on turn off that time and because you're capturing multiple occurrences of people of certain weight and height and health right uh at time 1 time 2 time three time four. So you want to uh and suppose the time is measured by years then it makes sense to say that okay let's analyze the second year after the participants enter into our program. So what happens what about the third year? So we switch off uh time two and turn on time three and so on.
It means you're treating each time stamp as a population. So, so as a population of health, weight, income, and that's it. No. So, so, so in some sense, the time disappears. We're just looking at the time three population, the time two population. So, after we turn on that time two and turn off all other time stamps, then we're only left with time two. That means the second year the participants are involved in our program of study.
What was their health weight and income numbers? Then we do regression. Right? So then it basically means if we have 10 time stamps that means we study the the same group of participants for 10 years then we have 10 regression models one for each year in that sense. So uh essentially that wouldn't cause a problem because the time would have uh sort of disappeared right into our proper handling of the time series. So that part is not a problem or less of a problem.
So that's not what we are going to uh focus on. Not so much a problem uh problem in regression. I'm just saying okay but this will be where all the issues come in. Huh? So uh what are the issues? Well, the issues could be that uh that uh we have uh what we call confounding, right? Confounding when you add in time, right? Then suddenly you think about it or maybe after doing a sort of a naive run of the regression model the the numbers come out and you think about the numbers the co the coefficients you you find find it very odd very strange or the p values suddenly make the whole model uh very bad not significant right or something like that and then you think about it is because probably time is also driving the income and rate so time is driving income and wage and then driving health.
For example, when you are younger, you tend to that means all else being equal, right? That means you don't pay attention to your weight. We don't care about the income. Our biological body tends to be weaker. We we catch a cold more easily. We suffer from fever more frequently and more lengthy. we recover slower than when you are a you know strong young adult right so your immunity rises this is all biological it's not about weight and income okay uh you could be having you could be obese you could be having high weight low weight but your immunity compared to your younger self is always going to be higher yeah simply because you are having more body mass than your younger Well, even if you're skinny, right, your your skinny young adult body size is still going to be stronger, right?
Heavier than your skinny young uh teenage or child age body. So, the immunity and all that stuff as time passes your immunity and all that. So your health on average increases already your income and if you capture the the correct segment or rather the wrong segment then that is the time when your income grows right very much with your with the passage of time right like example if you are age 16 17 you're some somewhere in the high school you're doing well in class so you want to earn some part-time uh income.
That's fine, right? And you cannot you cannot almost by the arrangement and the fact that you're a full-time student, you cannot be working full-time legally, you probably cannot be hired as well. So somehow by human rules, right, you cannot uh be earning a high income. So you can earn part-time income which to you might be a lot but if you look back uh now assuming you are now 20 plus right so you look back ah that's not an income right but at that time 14 15 16 17 you earn some money that's like wow right a big welcome uh side income so as you as time passes you pass 20 21 22 right and in university even if you do site income uh part-time tuition and whatnot walking the dogs, mowing the lawns and all that.
You can earn more, right? Because uh people know you are of higher age, you are more intelligent, you are more adult, right? You understand contextual uh reactions even not taught and not told. For example, if you walk the dog, at the end of walking the dog, the owners will not teach you or brief you, but you somehow know that you should, you know, bring back the dog, tie it up properly, maybe give it a bowl of food and water, and then you lock up the place so that uh the dog will not be starved or what not before the owner comes back later on or something.
If you're giving tion right and the parents might both be working and so nobody is at home uh you know the customer that is the student the young kid would then open the door and of course the parents will expect you to also look after the house in some sense because while you're giving tution some stranger wants to come into the house they should have that peace of mind that oh you are adult enough to know that hey even though this is not your own house, uh you shouldn't let strangers in something like that, right?
So, so these are like common sense thing that okay, they would then pay you more than they pay a 12 year old kid to help to give tuition, a 13 year old kid, a genius, right? But to help to give tution, but not enough life experience, right? But you have the life, so you're worth more. So whatever the thing is, I'm trying to say as time passes, your income increases. Yeah, your health improves even if your income didn't increase, even if your weight didn't change.
Right? As time passes, your weight increases because a few things. You choose the wrong time stamp. You choose the wrong C uh uh study profile, the population profile, right? Wrong time stamp because the time you choose, right? It's one year by one year basis. So over 10 years. So it's a long experiment with good funding. Fine. But it's a long experiment when the variable that you also includes the weight here changes or uh shows significant changes.
Right? So over a decade for a teenage it will he or she will grow right into a young adult and that time period the weight could increase right maybe not as drastically not the gradient is not as drastic as when a baby grows to a young child but teenage to grow into a young adult there can still be weight increase whether you are increased because you have more fat become obese because you eat more you can afford more more snacks and more unhealthy food to eat or you can afford to go to gym and have sports and you grow muscles so you become heavier right in a good way you you become heavy because of muscles that's fine but whatever it is the case time when as time increases weight increases or as time increases you don't pay attention to your weight right your weight decreases But then because it's a long time study your weight drops but then it is the it is still uh that means your 24 year old body right ex extremely skinny is still going to be heavier than your 14 year old body because you are the skinny DNA type right no matter how much you eat you are still looking skinny right fine that's good you might think it's good but no matter what I'm saying your weight at 24 years So will still be uh heavier than your 14 year old weight right so that's why you it is called confounding and that means we know contextually uh the way I narrated to you you would agree I don't have to show you data to prove right so uh if I choose long enough time then and I I contextually the variables that I've included in the regression they also experience changes growth with time.
Then we have confounding effects. Confounding effects simply lead to the breaking of the independence. Fair enough. Yeah. So weight no longer become independent of time. That income doesn't become independent of time. As you as you grow, right, your 14 year old self cannot earn income. 60 year old you earn a little bit part-time. 18 year old, 20 year old, you earn more part-time. 22 years old, you graduate and find a job, you earn more. 24 year old, you get promoted, you earn even more.
So over this decade, you experience as time increases, an increase in your income problem, right? So multi- colinearity comes in. So confounding is what we talk about the contextual influencing our as time increases weight increases income increases and for that matter even health increases right but the mathematical level confounding leads to multiolinearity right we talked about multiolinearity in the previous video multiolinearity leads to what leads to you still will get your B 0 B1 B2.
So you still get a model but this model starts to get fuzzy or you don't know whether this is B1 B2 they can be trusted or not. Are they under estimating the health? Are they overestimating? So that's problem and when that comes in then we cannot use this model very much. Right? Hopefully adjust R square shows a bad low number at least we get a signal and worse is that adjust R square also shows reasonable number then we have a big issue because multiolinearity wouldn't be readily or consistently uh uh discovered by adjust R square.
So because adjust square was calculated based on assuming there's no multiolinearity right so that's why when we break those assumptions all the indicators that we talk about they don't you know blink and show red light warning warning you know you have t trespass upon the assumptions uh is no longer reliable no such thing because it's a silent uh kind of uh cancer in the model itself so it the model doesn't know that that uh you your data had already uh broken the assumption.
So that's the thing and uh when we put in time then we also put in something that that uh don't quite occur if we don't put in time which is auto correlation of the data of the Y's. as the response variable. Uh so the thing is that uh many variables tend to change with respect to time uh in a way that is gradual and consistent. gradual and consistent. Uh they don't they don't uh sort of have wow changes in a short time.
Okay. Of course given long enough time that's why the the time period the the range of the time period and the intervals how frequently are we sampling make a lot of uh difference whether we suffer from the ill effects or not. Okay. Uh in the in the previous case here uh we were saying the confounding right? If you are sampling quick enough within a short time interval like um university year one students only okay university year one students the income are all you know there about right so uh and we're talking about year one one year time period so there's not much of a time so maybe we measure over 12 months right and so your income wouldn't be expected to vary by a lot right and and so it will be more like oh I do tuition for 3 months then quit already because the parents uh didn't want me to continue uh I do another tuition for uh 6 months because the child graduated that's why and then I move on to second year right so when the time range is short enough and the intervals are properly chosen because we choose 12 months not uh not daily right daily every day we go and measure the health weight then you'll be a lot of choice, right?
Because daytoday it would not change at all more or less, right? We would all agree to that. We don't have to really prove it, right? So dayto-day your weight probably won't change by too much. So it's probably random plugation because if you if you uh weigh yourself before you uh go to toilet clear your bowels and urine and everything and then you weigh after that there's going to be like uh maybe 3 400 grams difference right so uh all these will be random in nature so yeah so so that's not going to be very uh helpful to to your experimental goal but then you wouldn't have confounding autocorrelation.
So the the thing is that if we measure the time the range is long enough and these variables they do change rather drastically given the range right and then you sample and then you sample and the samples right if it's if it's too frequent then from day to day your weight won't change right your weight won't change too much or at all from month to month it will change it may change right depending on whether you are increasing your weight maybe consciously or subconsciously or reducing your weight right so the thing is when you have autocorrelation first of all what is autocorrelation autocorrelation means the the health is dependent on the previous measurement of health so what's your health this month.
Okay? So in other words, your health this month, health this month strongly depends on health last month. Can I say that? Of course, it may be also dependent on the kind of uh uh participant profile you pick, right? What is your population? Yeah. Um for baby health, right? So, newborns who just grow without consciously taking care of their own health. That means they won't actively uh take precaution to maintain health to sustain health to improve health right then very much it is it is uh kind of uh based on their own biological strength their body bodily strength right and also due to randomness so uh a baby's health right next month is of course highly dependent on the previous month's health.
A healthy baby next month can become slightly less healthy, right? Uh a little bit healthier, but then there is so much so-cal health you can go, you can't be going towards health, infinity hell, right? Health as a measurement say 100% is the best health. Let's say we all normal people, we are at 50. So the best you can go is 100. So there's still a lot of room. So assuming the maximum is not a problem, that means hitting a maximum is not a problem.
So the baby next month could be 51 compared to this month 50. But another baby, right, another baby that is less healthy. So this month is 40. Next month better better would it be 50. We within one month, right? That's why the time drives a lot of the rate of growth, rate of improvement, rate of deterioration that we all know, right? That we all know. So that's what except we don't know what's the rate. So that's why we study this.
But what the the the the autocorrelation is is that the for health for example and for many measurements that are natural not like health like uh weight and height and all these and and for that matter society kind of measurements that means uh it is human it is in a way the participants are all human so artificial you can purposely don't go for income you can purposely don't take care of your health but collectively it won't be the case that the whole population of Singapore, the whole population of a city won't want to take care of their health, right?
It won't happen that way. So that's why at a at a social level, society level, uh the measurements of this society, this population tends to be also depends on your previous value. Okay? So the health of this person is highly dependent on last month's health, right? So that's what we call the order one uh autocorrelation you depend on one cycle before one cycle before. Yeah. If turns out that the health of this person depends because it's no longer just baby but adult.
So adults sometimes drink right sometimes smoke sometimes uh eat nutrient supplements to improve their health. Sometimes perform diagnosis very frequently some of the health nuts right. um then you will find that hey maybe the dependency is not just order one so the health of this person depends on last month's health and last last month health right so then there will be order two autocorrelation so is there autocorrelation that is the question to be studied that means we need a detector the detector is certainly the Derbin Watson test right but that's not to test solely for health Din Watson test is to test for this entire model whether the remnants that is the residue right contain autocorrelation with with each other when the residues contain autocorrelation then this current model is not good enough even if adjust R square shows somewhat respectable number it's not good enough because there's a big chunk of residules that are not being explained.
That means your basically your so-called predicted Y is very bad from the observed truth the observed data. We'll come to that in a while. But just trying to see what is auto correlation, right? But um first thing first is not only health that has this uh autocorrelating of order one but also uh let's look at GDP right uh versus time as time passes GDP increases or not As time passes, GDP increases or GDP decreases. Hopefully not, right?
As time passes, GDP increases. Now, as it increase, can it double? Is it possible that Singapore's GDP goes from 2025 some uh X number 400 billion, 500 billion uh to double it in 2026? Is it possible? Nothing is impossible. Yeah. But does it make sense right? Is is it is it even in our discussion can can government say we want to plan to double our GDP in one year's time you can say that for over a decade right oh we want to do that so we should focus on AI development for example you can do that over 10 years time but not over one year's time we all would tend to agree to that why not why not over one year's time it's simply because time drives everything correct or not?
Yeah. So, we all know that time will not drive the GB GDP to such a drastic rise or for that matter drastic fall. Will Singapore's GDP become zero next year? Not impossible, right? But that would mean some sort of a externality like COVID 19, right? So, was it 2020? Uh it it there was a big dive in the GDP 2021 big dive in the GDP, right? But also not going to zero also right it's that's but let's say it's so low that we just we can basically consider zero fine that's possible but that's externality which is a different case.
So putting those uh rare but impactful externality events aside, you mean you know those things that would disrupt the the smooth uh progression of the linear regression model, then GDP basically grows in a way that depends on its previous value. If you think of GDP as a person's weight, if you have a high GDP uh at the start of the year, then you tend to also right because you have achieved that high GDP, you tend to also already have a population and the equipment and the land and the whatever assets to earn not only that but more, right? more than your that year's GDP.
So if assuming suppose that year is uh $500 billion and because you have gotten that $500 billion already in 2024. So it shows you your country already has the people, the talent, the equipment, the land or the knowledge to use the land, the various stuff to be able to earn that much. then you tend to be able to earn maybe 10% more but that's very high 2% more 2% of 500 billion right that's uh 1% will be 5 billion so 2% 10 billion 10 billion more which is very reasonable right it will be most you know within what most economist opinion to to say well very plausible Right?
And so on so forth. But if your starting base is only 50 billion, 50 billion, right? Then we say next year we're going to have a increase in GDP. We'll do this this this right? We already have the asset, talent, people, equipment, land to earn 50 billion. So we project that we will earn how much? So if you say we project to earn 2% people will be quite agreeable right maybe that's a plan only but then you say okay I earn 2%.
So good you want to say you earn uh 10%. I want to earn 5 billion more so we become 55 billion 10% right 10% people will start to raise alarm 10% are you kidding right so because that that's that's going to like make your people work uh 24 hours a day work 24 without sleep right uh or uh have a lot more overtime or you will need 50% more production equipment which you can't afford to Why? Because you only have 50 billion of GDP to to begin with and so on so forth, right?
So it becomes rather unbelievable to say you want to grow by 10%. To say you want to grow by 5 billion, right? If you have a base of 500 billion for this year, plausible. Why plausible? Because it's 2% of 500 billion. From 50 billion, you will say you want to earn another 5 billion. It's also 5 billion. But then we say not plausible not so. Why do we say not so? Because that is 10% of 50 billion right? So we don't agree.
But if we say 2% then we agree. So you see the way we describe GDP is we have a growth of 2%. We have a growth of 1%. Oh unfortunately there's a shrink of GDP of 0.5%. That kind of meaning. So we always use percent of the previous year even in our discussion even in news even in uh reports plans whatever right we don't talk about the absolute we we do talk about but you understand what I'm saying is that the absolute amount matters less when you plan when you analyze when you you know explore the feasibility and so on than the absolute bad than the percentage.
So that's why the percentage is based on previous GDP and the way we say it that way already hints at us that the GDP is strongly related to time and because we use percentage that you know is another another hint it's not to prove that it is having autocorrelation but it is a hint that you may have right a very strong autocorrelation okay for that matter we can also Look at uh inflation rates with time. Artificially created uh indices like inflation rate tends to defy gravity, defy physics.
Uh so for example stock prices changes with time and if it's short enough time for example stock prices of a startup a bank blue chip company which is very stable and all that right but short enough time within a month the daily changes for example right because within a month we all know that the companies policy culture CEO management bonus scheme whatever treatment of employees customer base whatever right could hardly have changed much fine Day-to-day there's always some dynamics, right?
But over a month you won't expect drastic change at all, right? Not Yeah. But over 10 years, yes. Over one year there will be some. Over five years there will be more. Over 10 years a small bank can become a large bank, right? A multinational can collapse into a small bank, right? Uh so big things can happen given enough time. So that stock prices stock prices when you have short enough intervals right like let's say daily stock price of DBS back over one month whichever money you pick over one month consecutive daily prices it will more or less be just uh random fluctuations and it can okay uh not just in a very rare instance but uh sometimes money flow in You don't know when but they do come in sporadically and then because of changes in um monetary policy, government policies, uh incentives and whatnot, right?
Or money being driven away elsewhere, right? Because for example, um US interest rates reduces. So there's no point putting money there in US. So they move elsewhere, maybe move to Asia, partly coming to Singapore. And so uh and then suddenly the stock price goes up. So it can happen that the stock price can go up and down, right? And can spike and then fall uh kind of not so according to real life physics, real life rules because of uh the ease of movement of money.
But if there's a cap, if there's an artificial cap, say the stocks exchange says, of course, this doesn't happen. But if the stock exchange caps it at, oh, your price movement cannot go more than 5% any day. Then even if a lot of money comes in, they can only max out the 5% stock price for that day, right? So, so if you watch daytoday changes, it'll be 5% 5% 5% 5%. So it becomes like a very gradual increase that is managed right because the money actually would have made the stock price shoot up 50% in one day but couldn't because of artificial uh rules that cause it not to do so.
So artificial rules tend not to not to uh make the current value depend on the previous value or it could or it could right but naturally occurring uh measures will tend to do that right so most of the time that is like inflation now inflation rate is like uh for example consumer price index and all they are artificially created, right? Because it's an average over the prices or the price increases. Um, but then look at the price increases.
Those those prices the the the the ingredients that participate in that uh averaging formula. Those prices come from what? Come from food, the clothing, housings, transportation. These are driven by the actual usage right uh the society kind of acceptable price changes those prices cannot double overnight or else some all hell will break loose right so those prices cannot drop to zero overnight or else chaos will happen too so it is very much uh monitored and uh uh uh controlled in that sense that uh you the businesses want to charge more for profit but they cannot just artificially double the price overnight they can't right they want to attract business and give promotion they also wouldn't all of a sudden zero out their prices for all their products I mean you don't hear that that means is collapse of the business right people will start snatching up the products so because of all these mutually uh enforced forces opposing forces So the prices tend to be stabilizing in a dynamic way.
You want to go up, sorry, I will complain, right? You want to go down, uh, uh, consumers want the prices to go down, businesses will say, well no, I'm not going to give up my profits. Right? So, so there'll be kind of a type of war and at that right tension, the prices will start to creep slowly up or down. Yeah. given uh the the environment the business environment the uh overall business climate. So inflation rate tends to increase slowly with time tends to decrease slowly with time and it will be dependent on the previous inflation rate.
We will be surprised if we see that in uh let's say they always announce inflation rates right from uh the Fed discussion. So inflation rate 2% 2.4%. And the next meeting they say inflation rate now has tamed down to 2.1%. We can consider you know uh reducing the high interest rates or whatever it is. So they 2.4% down to 2.1 is also quite drastic already but then it's still acceptable. 2.4% drops to 4%. Think about that, right?
It would take a lot of price falls, right? Like all prices fall rather significantly all at the same time. Then it will lead to such a drastic fall. But then imagine what would the society become right in that uh in that uh time period that led to the observed 4% uh consumer price index. It will be chaotic, right? So it will not be having drastic drastic relatively speaking right but no matter what inflation rate depends on its previous uh value uh especially over the right time intervates and many of these rates even though they are financial rates they are artificially crafted formulas still very much uh versus time very much that the previous value drives the next value.
So we have sales versus time. So for companies especially mature mature established companies that sell to public, right? like consumer goods, fastmoving consumer goods, shampoos. So, PNG, Unila, right? They sell shampoos, they sell laundry, detergents and all this. Now, can they double the price? Uh there'll be a lot of complaints. There'll be a lot of uh noise. So much so that government will step in. Why are you doing this?
Right? So, so they cannot just unilaterally just do that unless there's a change in the environment, right? No. So uh sales for this kind of company's products they they are already established right they have a steady customer base their products are used by a lot of consumers right and so that they're sold all over the world so it's very like stable it's like a big tree right now yeah you can you double the prices can you double the volume of sales will people suddenly overnight wash their hands hair three times a day rather than one time a day.
Change their practice, change their habits. It's tough, right? It's unimaginable that people change their habit overnight. Can people change their habit over long periods of time or over a decade? Sure. Right. Singaporeans are continuously being being educated, right, through campaigns, through advertisements, through slogans, through reminders, in our bills everywhere to watch our use of water. You hear it once, you hear it twice.
Don't you know just government saying that, right? You don't pay attention. You hear three times 30 times over a year in various places, right? 300 times over 10 years and gradually you say, "Well, perhaps I should play my part." Okay, you know, you're not being forced to, but you're just being educated to gradually appreciate that Singapore is small. Every water droplet is very precious, right? We cannot manufacture water uh at least not cheaply and it takes a lot more water to make one drop of water.
So it doesn't make sense. So we better uh be very prudent and use only that much water. Okay. So finally we get it right. So it's like over time yes the water usage per person and therefore the whole country uh decreases right. So, so that would be a sort of a manage campaign to drive the usage of water down. So, that is basically right how how uh the dependency can happen right gradually when you measure long enough time.
We cannot be overnight having our use of water. Maybe we did do that, right? So when we we when we were young we shower and then just let the water run because it's fun and we don't care about water saving. But when we are adult maybe we end up using less water than we were a kid even though we have a bigger body size right to bathe to uh whatever it is. Uh and how did that happen? Or because we are more conscious. So we can be educated to half our reduce right our water usage.
But can people just do that in one day in one month? Yeah. So that's what I'm trying to say that the dependency is there the sales whether you talk about dollars or volume cannot have cannot double for the company they want to double it but that's not possible for the people they want to save on buying the laundry detergent and shampoo right and shower gel and the toothpaste and all that. I want to have it right if I can but half right I cannot do it overnight but with practice with printing reminders all sticking all over my house my toilets then maybe I'm reminded again and again and I would uh urge myself discipline myself a little bit every day and finally over a period of maybe 3 years 5 years I could right finally reduce by 30%.
Okay, maybe not half yet, but uh significant reduction is observed, but that's not overnight. So again, the the range of your sampling of time interval, the time interval itself and the range they dictate the magnitude, the strength of the autocorrelation. So but all these inherent dependency on previous sales it's already there. If I'm a big water user right I have the habit of showering five times every day then when I manage to finally reduce my use of water probably it will be two times every day.
Yeah. If I am a person who showers only once every other day, once every two days when I finally reduce my water usage, it may be that oh okay then in that case it's a good reason for me to shower once a week. And that is the kind of relative reduction that we are talking about right now. So, so the dependency on the previous value is still there although the magnitude of the uh the absolute quantity is all uh different from people to people.
Um others could be unemployment rate, productivity and so on so forth. Right? So basically a lot a lot if you can already see the pattern here right. So uh they are all finally although some numbers like interest rates uh consumer price index all these they might be formula of combining other variables into one like taking the mean taking the median whatever it is those ingredients that feed into this formula they themselves they are generated by ultimately people right or if you're talking about say uh studying you know pets behavior so it's about dogs and cats and animals So if you're talking about that then ultimately it's also about those animals behavior which is going to be biological organic and hence dependent on the previous self.
So uh for a particular species of dog the maximum size it grows to whether you're measuring the overall length from head to tail or weight right will be dependent on its previous weight. If it's a small kind of dog small breed right then it will still be growing. It will still be growing just like human but maximum size is this. If it is huge kind of dog then it will still be growing but it will be growing at a faster rate while young and then the overall size and weight will be larger right but it will not become infinity it will not become something out of control right yeah so so again a lot of these organic biological and naturally occurring for example the number of trees in the forest right again you cannot double all of a sudden.
It will grow and grow and grow, right? But it it won't suddenly double in density. So all these things they are pretty much within our expectation of daily life's behavior. So in that case then the number over time right? So time is a variable. Now remember that's what we started talking about this. So if time is a variable inside your explanatory variable then you better be careful that that uh this auto correlation my current value depends on my previous value that my exist okay so if that exists then that means there is embedded tangent uh tangent gradient there because if I depend on my previous value that means there's growth rate because we are we are dependent on time if there's growth rate Then that means that means there's embed embedded growth rate within that within the response variable.
So the embedded growth rate might be accelerated that means there's acceleration in the growth rate that means d² y / dx² instead of dy dx being a constant it may not be a constant. So if dydx the gradient gradient m3 here is not a constant then we are in trouble. So that means we are trying to use a linear model to to uh mimic or reproduce a nonlinear data set when time is involved. Right? So when time is involved basically uh there can be new problems to take care of.
Okay. So today we want to look at what is this very time related problem called auto correlation. Okay let's look at autocorrelation. Auto means with itself. with itself with myself. What do you mean by with myself? With Y. We're talking about Y, right? So the Y has correlation. You know, you recall the correlation, right? So X and Y, are they correlated? Linearly correlated, right? So uh is Y correlated with Y? That's always a one.
If time is not involved, am I correlated with myself? Using the Y sample, you correlate with Y sample itself is a one. Perfect correlation, perfect positive correlation. If I change, I change. If I change a lot, I change a lot. That kind of meaning. But when time is involved, we're talking about am I changing a lot of over my previous self and there's now a sense of previous because we're talking about t minus one. So if we talk talk about uh 51 as auto correlation of order one order one lag and lag means shift oh my GDP this year depends on shift one order one so my GDP this year depends on last year's which is often times what we hear what we read in the news reports right uh quarter a year to year there is a 1% increase in this year to year of course with pre previous year not of the next year right which hasn't come not of two years ago definitely not right so it's always year to year when we report numbers like that and the 1% always is more or less there8% 1.2% 2% yeah negative.1% that 1% okay so it's always hovering around that percentage then we get a sense that it is likely to be a autocorrelating sequence of numbers because remember if I am growing 1% this year next year and following year let's say they are all 1% of the previous year if I start off with hundred billion dollars GDP then my 1% growth is 101 billion but then the subsequent year is 1% of 101 so it's compounding now it's more than 102 already and of course you can say okay after rounding it's 102 but then it doesn't go like that going 10 years into it you're going to get more than 110 billion you're getting much more than 110 million so percentage growth is a little bit like compounding effect already right So that's why we will strongly uh when you hear things like that uh you should now suspect that there is a autocorrelation or there can be we need to check hopefully we don't have but unlikely.
All right uh company has been increasing it sales quarter to quartarter about 2% every quarter. Okay. Right. So that is another triggering point. So when we regress if you didn't put in time perhaps it's okay but if you put in time if you're studying whether the grow in number of employees affects company sales there's no time right so fine then it's okay but when you plug in time as well number of employees and time try to explain the growth in our company sales and yet you'll be hearing news like quarter to quarter grow uh year to year quarter to quarter growth right 2% then we better check for auto correl correlation.
So here the idea is that the total autocorrelation correlation is uh of course something to do with covariance. Okay. And uh the denominator is going to be the sum of squares because the divide by n or n minus one they cancel out. And we are only looking at the y values. Okay. The y values throughout the times one to end. Yeah. Uh each y value of the y bar then square. So it's square of deviation. This is quite quite understandable, right?
This is like the numerator of the variance of y. That's all. So it's like ss y essentially. But the numerator is where it the magic happens, right? So the i is going to starting from two. We would like it to be one, right? We would like it to be one. But because we're going to shift the y mean it means we're going to make a copy of y and shift it a little bit so that we can say at that time t = to 2 y1 explains y2 at year 2 GDP of year 1 multiply by 1.01 right or explains right year 2 accounts for GDP of year two right now.
Yeah. So we got to shift by a little bit. So the autocorrelation is going to be y of this year away from the mean no square because we're going to multiply by y of last year t minus one away from the mean. Okay. And because we're going to have last year I has to start at two. The index i. Do we use i or t? T. So the t starts from time two. So 2 - one right 2 - 1 is one. So then this would make sense and and the one deviation is multiplied by the two deviation.
Then of course the two deviation is multiplied by the three deviation and so on so we add up this and the one here right the one here is the order one. And if we were to uh have other orders then we would not minus one we would minus uh two right but then if we minus two means we going to shift two cells down we better have a lot more data. The more lag we are trying to study the more data we would have to skip right because we need more data to form the lag then we can see what is the effect.
It's like saying that uh to study the lag two order two lag autocorrelation we need year one and year three data skipping year two par see see how the the two years ago the GDP influences this year's GDP and that can happen why because government policies they may say that to encourage trade I give you rebate but you cannot claim rebate immediately for this year you have to claim 50% 50%. Or maybe you claim 60%, remaining is 40%.
Okay, no matter what the the effect only dies off two years later. Now it's artificially created, right? Because of the government policy. Uh and the government policy is trying to sort of make sure you don't sort of trade and then run away. You have some uh uh carrot to continue trading. And the second year you make more trades and then you have more claims in the third years. So you continue the trade uh more like an encouragement because government doesn't have to give you rebate right but I give you rebate but I hope that you won't run away with taking the rebate and then go right I I would also like you to stay but I don't want to have a rule that say you cannot go away right I mean if you want you go right but you just forgive give away give up your rebate which is not available to you after all right the government didn't have to give you the rebate so assume that kind of policy is crafted this way then we should expect that uh because Singapore has a lot of trading companies trading country right so the trading companies will trade and then claim rebates for example this rebate I don't know if if it exists but assuming if it exists if it exists then the companies will claim right and Singapore is stable most companies will be very happy to continue staying and then they can claim some more of the lef over 40% right while doing more trades.
So why not? So they will uh claim the rebates the 40% only in the second year after the trade was start right. And so the lingering effect will be created because of this policy which is artificial. Artificial because people created this rule, government created this rule, not a natural uh phenomena in that sense, right? because this this can also be taken away like after implementation this doesn't seem to help the GDP at all okay take away don't lose money right something like that so so if this policy were implemented then we will see a uh GDP two years ago can influence this year's GDP because of the rebates right and and the autocorrelation we have to take into account the lag to effect the order to autocorrelation.
So then we will do a five of two calculation that basically is the idea. Okay. Now uh so so uh this is this is just a definition and it can look a little bit uh intimidating because we got to think of the lag right. So let's start with a bunch of collected uh GDP GDP data but we're not going to use the actual GDP just something very simplistic. Let's say we have a 6 5 76 8 97. Uh essentially these would be because of the time being 1 2 3.
You can think of these as quarters or years. So we have seven time units. Typically when we come down to this uh numerical scenario we don't have to worry about the details because whether it's measured by quarter, day, month, year doesn't matter anymore. We just concern ourselves with the numerical number crunch. So what I'm going to do is to calculate the auto correlation of this bunch of y values by first uh we need the y of t minus one and the y bar y bar is everywhere right?
So first we calculate y bar. y bar is uh and our y t of minus one is going to be a copy pasted and shifted value. Always remember that we want to say use the GDP of year one to measure or or uh uh calculate for the GDP of year two. Right? So, so we want to align the six with five and the five with seven. And after a while, you just basically will see another way to do this, which is to copy then shift down. Right? This way of thinking is easier.
Uh let me just finish writing first. 8 9 and you see that the seven is being skipped out. Right? So, so mathematically what we want is that you know this is the shift right? So this is the shift but programmatically when you use R or Python or for that matter even Excel you can do it uh you should do the whole vector copying copy paste then shift down by one cell that's a lot easier to do programmatically not so but the thinking is to do this we want to use six to autocorrelate with five or to try to explain five see if there is strong or weak, positive or negative auto correlation.
Um notice that there is uh numerator there's no squaring. So it can be minus * minus like coariance it can be uh minus * plus resulting in a negative autocorrelation. also possible. Also possible positive autocorrelation, right? The last tier is one. Next tier is one uh 1% more. Next year is also 1% more. So it's positively growing. Okay. So this could be GDP. This could be uh population of a young city. This could be whatever it is some kind of compounding even negative autocorrelation GDP this year is 100 billion next year is 99 billion you know subsequent year is another 1% lesser another 1% lesser okay so it's like a falling GDP right by percentage not by 1 billion every year but by 1% every year so it's more going to be more than 1 billion several years down, right?
But the first year is just exactly 1 billion. That's the difference between just subtracting 1 billion as a constant and multiplying by.99. Okay? So then what? Then we want to calculate uh two things. We need this and this, right? So we need yt y at time t minus y bar. Okay? And we need y at time t - 1 minus y bar. So because y tus one doesn't exist non undefined, right? Because we are talking about time zero which is undefined.
So we're going to skip this. This doesn't exist and we will just be uh doing uh yt5 minus 6.8571. So we will get a negative uh uh oh sorry uh this will get a negative 1.8571 and because we're minusing y bar we can just simply do it as well. uh because the denominator includes time one which is defined this one is defined so we should also do that and we get0 now I won't want to keep on writing all the fourdigit numbers this is not a lottery comparison here uh because the rest will all just be 7 minus 6.8571 85 71 right okay then 6 minus so we can just do that but I just want you to be clear of what I'm doing here right and also to check that hey the first two you get the same digits as me the rest you will see some sort of a 1428 8571 and all that stuff and I just give you the last two so that you can compare and make sure you are on the right track as my calcul calcation 29 and the last one 0.1429.
All right. Yeah. So because the the formulas are the same so you can just repeat and base. I I don't need to show you the rest. Uh because if you're right on these four numbers you should be right the for all the numbers there. And for yt - 1 - y bar is you use this column and the first now now I'm talking about this first cell which is blank. So this first cell is blank. Second cell is to use 6 minus 6.8571 resulting in8 uh 0.8571 right uh 1.8571 8571 again all the way down right to 1.14 29 and 2.14.
Okay. All right. So now the sum product sum product of uh these two columns. Okay. Which two columns? Um these two columns we want to perform a sum product. A sum product is simply that with the first cell times the first cell gives the first cell times the first cell plus the second cell times the second cell plus the second last time the second last plus the last and times the last cell right which is the sum product of y t minus y bar y tus one minus y bar.
So that gives us uh 2.9796. Did you get that? Okay. Even if you use calculator, you probably can do it. Uh but you can do this in Excel, in R, in Python very easily using num. Okay. Then next we have sum of the squares of uh this one everything. Okay. because the definition is over the entire t = to uh 1. Now some software namely in uh in Python's stats models.tsa TSA time series analysis uh there's a parameter for you to say okay I don't want the all everything I don't want to match the the skipped sample so if my if my summation is over from 2 to n I want this to be also from two to n the adjustment that means we it will we will ignore this one you have a means to tell the function to ignore this one in other words you you can try to standardize the number of samples involved in the auto correlation Although it is also acceptable right uh officially to to uh include everything in because it's easier to calculate everything uh and we can do that right so for small l that means uh low orders order one order two order three that means you skip one two or two or three relative to the sample size if your sample size is large then order one two three small orders include include don't include you're not changing the denominator by too much.
So then typically the autocorrelation if it's positive if it's negative the conclusion remains whether you choose to skip or not skip the first few uh so-called error squares. All right. But when sample size is smaller and the order is large like order we are not calculating order one but if we calculate order four then we are like almost like dropping off more than half the data right then it can be very serious in terms of the conclusion when there is no autocorrelation you can conclude there is right simply because you're picking only the last two or the last three then then we need to be very careful so uh small sample size is always not a good environment context to begin with.
As we all know in all statistical calculations, uh touching reality only small number of times not good. But here we're just illustrating the formula. Uh so here this formula we're trying to uh of course do the sum of squares yt minus y bar uh itself two times. It's like it's like I copy uh yt minus y bar times itself. That's why it's auto. I I square it because it's two copies of that. So my number is 10.8571. Okay.
And finally I get the autocorrelation fee of order one is uh 2.9796 over 10.8571 and the ratio would give us 2744. Okay. So there is some slight positive auto relation going on and whether it's serious or not depends on our own assessment. So even though everything is based on data formula is always well defined. You and I will calculate the same data. We'll always get back the same autocorrelation. But is it serious or not?
Is this positive autocorrelation bad enough? right is up to our assessment depends on what these numbers mean. All right, if it is progressive test scores, my own test scores, right? And I'm seeing autocorrelation positive. A good something like uh as time passes my scores are improving upon the previous scores. Maybe this is GMAT score, uh TOEFL score, you know, then then uh a high positive correlation, autocorrelation could mean I have a high learning rate, something a term that I borrow from uh artificial intelligence for my actual intelligence now, right?
That I'm picking up. I'm improving uh really significantly every time I do another test. This practice makes me even better, even better, even better. So that's good. uh but for human perhaps this is already very good uh that I'm improving there is a positive autocorrelation right and if it is negative to neutral that means I'm getting more and more frustrated why am I repeating right why do I have to practice practice doesn't help so the more I do the more frustrated I become because I don't believe in it and so I have a negative autocorrelation so I I think you know you can you can put it in your own context you can have a feel of that right have a almost an emotional feel of what autocorrelation means uh because we are all we are trying to autocorrelate our performance with ourselves our previous self right in the previous time frame time measurement okay so that is the calculation involved okay next let's look at a test that would allow us to uh gain a understanding of whether the model which model this model you know example model whether our model um has issues involving the residues having auto correlation with its previous error >> because now we have a sense of what is previous.
When time variable is added in, it is meaningful to then say the error incurred at time one, right? What's the error? The error is the model, right? The model uh producing a predicted y at time one when we also have a measured y at time one. So then there's a discrepancy. We learned about the residual in the previous sessions in session seven, sessions eight. And the first error, the first residual predicted minus the y does it predict right or is it correlated with?
Right? Is it correlated with is the next residual correlated with the previous residual? Right? Or is the first time the time equals to one residual predictive of the time two residual which is itself predictive of time three residual and so on. If the whole residual series is autocorrelated whether positively or negatively that means the residues themselves are predictable based on past values. Correct? Because that's autocorrelation, right?
That's trying to say that present value depends on the past. Suppose it's on we're just talking about order one autocorrelation and order one occurs a lot more uh naturally frequently than order two three uh order two five maybe right like I say the government policy order three could also be because the rebate is uh permitted over 3 years order 6 7 8 9 10 harder and harder to come by right nothing is impossible but uh it's just rarer for us to say that oh you know what that phenomena must be order 12 autocorrelated data series or what it's it's just uh harder to imagine.
So while we don't have a maximum for order uh naturally occurring data most of the time will be order one is sufficient to tame the problem due to the inclusion of time variable most of the time. All right. If you want go on to order one and order two and maybe even order three, right? But then don't try to think of I need a very perfect model so I will try to do one to 20, right? That that may not be necessary nor even good for your modeling.
Okay, coming back to the detection of errors. Okay, so remember this is strictly about errors Watson test, right? So let me highlight uh that this is for errors the residuals and we have a good and well definfined uh formula or definition for residuals right the predicted minus the observed values of y nothing to do with x here right so autocorrelation is a concept that can apply to anything and in my illustration just now I calculated for the value of y which is this.
You can of course calculate for values of weight and income. But our concern in the context of linear regression right when we apply autocorrelation will be of the response variable because if response variable itself is uh strongly explained by right it's highly connected to its previous value then there may be a case for autocorrelation right right so we calculate Autocorrelation most of the time for two things one is the y variable the other one is the residual residual again I say the third time now is different from the y value because the residual is y hat maybe it's good to put that in here is y hat of t minus yt t not y bar.
So it's a vertical difference between the star and the dot or the dot is on the predicted line. The star is the data point like this. In other words, since the yhat is involved, yhat is this entire formula, right? This entire model, not even here. This one is the data, right? So the y hat is the right hand side evaluated given the same data points as weight income and time. So the autocorrelation of errors in other words the derby wasn't test is done when you have done the model whereas autocorrelation of y I don't even need to run the model you give me the y I can start calculating.
So that's the main difference. Don't mix them up. Right? Okay. So, Dvin Watson test. Let's look at uh the test statistic. How do we how do we calculate the Durban Watson test statistic? It's always like rock climbing kind of a cliff. And I want to see if I can save this side of the space. Okay. All right. So, what is the test statistic? Um the Durban Watson Derbin Watson Derbin Watson test has the hypothesis of H0 that the population auto correlation of order one, right?
Uh is zero. And of course, row one is not. So uh no auto correlation order one has autocorrelation order one. Now we need to of course quantify that uh oops I'm covering my my own writing. We need to quantify that it is of order one clearly. Yeah. And we will only talk about order one. Like I said order one occurs much more frequently and almost like readily and we need to worry more about order one effects because we hear that we read about that. just think of the GDP numbers.
Uh and uh next thing of course when we put down order one is that uh if there's no autocorrelation of order one can there be auto correlation of order two uh possible right there are some phenomenon I uh it's not so frequent so I cannot be listing them uh readily but there may be a sort of skip one year then happen that kind of uh dependency Uh so if that is the case then uh perhaps due to policy or otherwise you skip one year and then it affects the performance.
So if that's the case then you have like zero or very little order one autocorrelation but existing right some level or even strong order two auto correlation in practice. However, seldom, right? I I would say very uh again, I don't want to use the word rare because if it does happen, then it will happen in that niche area uh very frequently, right? Uh but not easily or commonly experienced by most people. That that's what I'm saying.
Furthermore, if order two happens, then order one also tends to be in influence. That means order one will influence usually stronger uh more and then order two will uh come into influence but usually the strength is lesser. It's a little bit like echo effect. Oh, when you make a sound in the mountain ranges, the first echo comes back strongly and first, right? because it get bounced off the nearest cliff and then comes back.
Then after that it the the voice your voice will also go on to hit other cliffs further away. So it will take longer time and because of dissipation it comes back weaker. So it will come back weaker and later just like echoes and typically that is also the case in most variables we measure in real life. Okay. So uh for now let's focus on order one analysis. If we can arrest this beast, then we should be on to a good start, right?
So, uh we now want to see if the test statistic from Derbin Watson is going to reject the null hypothesis or not. If so, right, just like the global model F test, if we reject, then model is significant. If we reject here then there is autocorrelation which is bad news. Okay. So uh what do we do with the test statistic? So the test statistic is with respect to the names Din and Watson DW right uh similar to the auto correlation.
It's a bit like quote and I quote sample sample calculation right. So uh we do this right 2 to n e to the t or e subt minus e t minus one squared. Now let's look at this one first. This is residual or error at time t at time t minus the previous error. Now there is a uh a very very important presence of structure here. It's like the previous in in performing regression previous makes no sense, right? Because previous value for weight is not the same necessarily the same as previous value for height individually.
But previous residual and uh and the current residual they make sense here because we're talking about time. So uh there is strong exploitation of the sequencing of time and just as it should be that means we are respecting that this variable time has sequence has order. How do we mathematically respect that? that means include that orderliness into our calculation thereby having different results if the order were to be changed and that is exactly this uh t and t minus one.
So the order is not so much the the priority of summing the error differences but rather the lag no the the the lag in the sense of uh lag by one time unit or lag by two time units then it will be t minus two. Okay. So then again this is a bit like autocorrelation. We will do this as a ratio over the entire residual squared as the so-called total variance without dividing by n or n minus one taking the average that is without taking the average we just do a sum of squares of the errors.
Right? So sum of squares of the errors. This trick we have been using a lot of times whether in testing for uh equal variance assumption in the F test. We also do that uh we we just use the sum of squares right. Then we we do that in ANOVA. We do that in uh of course the global model uh F test. So we we do sum of squares. We're very used to sum of squares, right? But the numerator here is where the structure is being exploited.
Now the previousness on the sense of previousness is being exploited. And once we calculate this test statistic then we will be able to uh obtain a number right. So this number right this number is between 0 to four and I purposely stretch it out. Okay. 0 to four. And it depends on various things amongst them. Your significance level, sample size, right? Just like t distribution. But in addition, uh it also depends on your number of uh predictors and that is that is here right.
So it's like x1, x2, x3. So then k equals to three and there are three predictors, right? Then we got to look up the the right uh critical value for Durban Watson to test uh to to obtain the boundaries. Uh we will use the critical value method to compare with the test statistic here. Right? So because for derby roson if you look at the critical value table they will have two values and that's not because it is uh two tail two tail you will get four critical values right something like that u more because first of all for derin was test statistic closer to zero implies um strong positive with correlation or autocorrelation, right?
And nearer to four and you cannot get a value larger than four. So nearer to four, you get strong negative auto correlation. In the in the middle is the value two and two has a rather special row in that it means no auto correlation of course order one. All right. And then you have of course uh various shades right going outwards. So, uh, green light here, right? So, then you have somewhat of a orange, you know, orange intensity.
It's from two to three. There's no fixed number because you're trying to obtain the correct uh critical values. And when you go on to here right it will be negative autocorrelation. And when you go on to here you will have uh positive auto correlation. Okay. Yeah. So just look up the table for Din Watson. You get two numbers for one tail test, four numbers for two tail test. The two numbers are such that uh if your test statistic fall between the low and the high number then it's inconclusive cannot tell.
So you probably have to redo with more sample size or wait for more time periods to gather more uh samples and then you can try again. Okay. And so if that uh if you're doing two-tail tests, you will get four four um four uh uh critical values. So if the Dubson test statistic, you know, falls below or falls above the high uh high of the four, the highest number of the four numbers or falls below the lowest number of the low number, then uh you will want to conclude that there is there is autocorrelation.
Okay, when you're testing for yes or no autocorrelation, we don't care whether it's strong positive or strong negative autocorrelation because if that is autocorrelation, our our regression model is in danger of its reliability, right? We are not very sure we can use it anymore. Remember the Wson test is about the model. Okay, let's write that down. Okay, this test is for the model to test uh so it's used for uh checking auto correlation of model errors.
So if you don't have a model then you don't need to run the Watson test. Yeah. Because this E of T you can only get when you get a model right you calculate the residue. Okay. So, so think of the Watson test as a sort of a like just like COVID 19 test kit. A test kit to let us know if we are suffering from such a problem in the residue. Okay. In the residue. Now, why would the residue imply anything about the the uh problem about the model.
If we have a model and remember this model here, it can also be polomial. It can also be logarithmic random variables. Remember in our session 8 we talked about that. So it can include nonlinear variables. what would have been our nonlinear variables get transformed right in our session eight into linear variables. So what that means is if our variables they are able to account for a lot of the fluctuations in y then what remains should be noise right should be independent even when time is involved and we kind of like don't give any special care to time then what remains ought to be just noise but dby was says But let I'm a sensor that cares about time because you did put in time.
Of course, we don't run the Wson if we don't have time. Okay, the time variable in our explanatory variable sets. So when is involved then was says but I'm going to treat your time here, right? And then check if the subsequent residuals they are predictable. They are like 2% of the previous value. Remember the GDP example right so if errors subsequently they are 2% of the previous error right later on later errors they are later residues they are 2% of the previous residue which is 2% of the previous previous residue and so on and therefore Wson is going to be you know either near zero or near four right that that means there is auto correlation in the errors what does that mean that means the errors will be larger and larger, right?
Remember the compounding growth kind of situation, right? Now, think about it. As time passes, the errors are larger and larger. That is the conclusion. So, as time passes, errors become larger and larger in a disproportionate sense because compounding is this is nonlinear, right? It's like like an exponential growth. So as time passes, another way to put it is as time passes, residual grows exponentially. Which means what?
Which means your model is uh totally poor in explaining health when time becomes larger and larger because you're making more and more errors. See that even that's despite your your variable here could have been income squared or weight weight to the power of three you know or square root of income or something like that nonlinear right despite that as time increases your predicted health and the current observed data it's growing exponentially or so that is revealing uh the fact that your model is not so healthy, right?
It's unable to cope well when time increases. That's why you see even if health itself is nonlinear, right? We could put in other transformed variable to account for the nonlinearity and thereby reducing the residue the error the the differences between predicted and the observed data including when time is involved if our nonlinear variables are good right they are able to recreate the nonlinear shape observed in health in the y but if not if we didn't transform form the variables let's say right and health is nonlinear with uh weight income time then the residual will have to take in the rest of the unaccounted for errors right now because data and the prediction if your prediction is up to data error is squeezed out to be zero if your data and the prediction prediction is very very huge distance away from the from the data maybe at first when When time is small, it is very close.
So area is small. But when time is large, it has grown exponentially. The gap has grown exponentially. Then the gap is precisely the residue, the error term here. That is why the Wen test will go for the error term as the unaccounted for gap. Right? Does it grow? You know, does it does it depend on his previous cell? And that enough that is itself enough to describe a lot of uh nonlinear behavior whether it's exponential you know parabolic or whatever it is right uh if you depend on your previous self doesn't matter to what degree very strongly like exponentially very weakly like square root kind of dependency or whatn not there's non uh that there's dependence on the previous self there's some sort of a recursive description future values are some modified slight slightly modified version of the previous value then it will show up as rejecting the null hypothesis right so then we will then uh doesn't mean that it is invalidating the model perhaps we can still use the model because the samples are so rare we just use for the lower time periods we don't make the time a large uh quantity right for positive correlation uh autocorrelation and otherwise we may try to slice the time and whatn not do other things so we can do some reactive measures depending on what the outcome is right so that's why the Watson has a very uh general ability to help us detect uh time oriented errors which is auto correlation So hopefully now you understand why we look at the residules.
Residu is like CSI detective digging out the rubbish the the unwanted residual right from the so-called criminal right the potential accused you must have left certain rubbish right so I look at the rubbish I can take out uh a lot of uh extract a lot of evidence right whether the food you you ate the the shop you bought the food from the DNA, the saliva, the whatever it is, right? So we can extract from and that is what the binos uh totally to to number crunch and and uh microanalyze the the residuals and exploiting the structure to detect uh dependency about this previousness.
So other tests don't quite care about T and T minus one and also cannot care about uh I and I minus one because at in other regression models that don't involve time the the I sample and the I minus one sample they are supposed to be independent first of all right so uh doesn't matter okay let's uh take a look at an example perhaps to help bring out the so-cal power uh the power the the effectiveness of David Wson test right in uh helping us to detect autocorrelation and uh then we see how we should react to it.
Okay. So perhaps we can use back our example here over here. All right. So our example starts with uh this y = to y = yt right yt 6 5 6 8 9 7 and we were saying that this is because we are using time equals = to 1 2 3 4 5 6. Okay. Now suppose we just do the usual linear regression in this case simple linear regression. Right? So our model is going to be uh our intention is going to be uh y is yt is equal to b 0 + b1 * t right so using t using t to explain yt so if yt is GDP then we say as time passes time explains GDP, right?
So as time passes, Singapore GDP grows. So time is a explanatory variable. Although it sounds strange, that's what coming that's what is coming out from linear regression, right? Uh so time is a explanatory variable of GDP in the in our response variable here. If we do that, then what happens? If you just do your number crunching, you will get um we get B 0 is uh 5.1429 and B1 is 0.42 42 right. So um the model is completely soft.
We get this. How do we know if it is suffering from the time illness and that is auto correlation. So what we do is we will calculate for uh y of e our predicted y. Okay. So our predicted y is yhat of t. predicted y is going to be 5.5714 by just plugging in one into the model here. Right? So we have our b 0 b1 so you can calculate yourself. I just give you two numbers. Second number is 6.0 zero and then I will spare you the rest all the way to the last number 8.14 uh okay maybe last one 8.1429 [Music] rounding to four decimal places but if you have more decimal places please keep them so we have these all right then we have uh the errors so the errors will be E of T.
E of T, right? E of T is going to be Y hat at time T minus Y of T, which is this minus this, right? So these two give E of T. So like these two give E of T. Okay. So, uh, yhat minus yt yt is the soal correct number. We don't doubt the data, right? So, my e of t is0 4286. Second value, easy 1.0. Third value, you keep going. All right. All the way to the last one. 1.1429. Yes. Did you get the same? Okay. So that's the signed error difference.
And we need the signed error difference first because the second term here is E of E of uh T minus E of T minus one. Oh, okay. Just watch carefully. Right. So we are talking about this term E of T minus the E of T minus one. the one. Okay. So the first term we we are going to talk about t0 which is non-existent, right? So this doesn't exist. Cross out. Second term is 1 minus0.4286. So we get 1.4 286. Third term right.
So it's uh I get -1.5714 because it is uh0.5714 minus one here right so you get this and so on so again I will spare you the all the numbers here 2.4286 4286. If you get your ET correct and the first two terms of your ET minus ET minus one, then it should be safe for you to get the rest of the numbers correct. All right. Uh let's do a quick check. So sum of squares sum of squares uh we get uh sum of squares. Okay. So you can do sum of squares very easily in uh Excel sum sq right are even easier python straightforward right so you can do that the number I get 5 then the sum of uh squares of this so let's do that this one.
Okay. Goes to this guy over here. And we have uh sum of uh sum of squares of e t minus e t minus one. Right? If you tabulate this column, then it's very easy. Just do the repeat sum of squares. So 15.244 9. Okay. So this time around this right sum of squares square each term and then add and this gets uh recorded here. And Durban Watson is none other than just uh dividing 15.2449 by so Durban Watson right is 15 over 5.743.
So we get uh 2.6679 66 depending on what is the alpha sample size we have uh we have six here because we threw away one we have uh oh no we included that so we have uh seven here so depending on your choice of alpha right because our k is one we t as the only time variable, one explanatory variable. So depending on your choice of alpha, we will end up getting four values, right? So DLD, DLD high uh and then another DLD on the left side and the right side.
So this is leaning towards the right side meaning there's a bit of indication of negative autocorrelation which is of course uh not in agreement with our pos slightly positive autocorrelation. So perhaps in the end it's inconclusive or it doesn't it doesn't it is not strong enough. the the autocorrelation whether negative or positive is not strong enough to make the Watson reject. So that we would say there's no autocorrelation right.
So perhaps it doesn't matter that means indeed right indeed uh well this is not GDP data so it shouldn't be indeed but suppose this is GDP data then if we fail to reject then it is uh okay an acceptable model at least for this set of GDP data for Singapore for this time period. Okay. uh but as time passes Singapore may change its economy the structure and activity and all that then GDP may be more more uh susceptible to fluctuations then it may not be dependent so much on its previous set right so there's some dependency but not so strongly so in this case it seems to be not so strongly uh positive not so strongly negative so there's some fluctuation so maybe after all it doesn't matter But that's not what I want to show you, at least not completely.
What I want to show you is what if we use an auto reggressive model. So we use an auto reggressive model. What is an auto reggressive model? So let's say we use auto reggressive model. of order one. This means once you see the pattern you will be able to understand this rather deep word right auto reggressive model what is that right but actually it's very uh sleek and idea is very simple that is I'm going to explain why with y itself after of right after all the Wson test says reject let's say this is a number that is so positive that that it rejects right so we reject and say there is autocorrelation and we noted that it's slanting on the positive side uh positive test statistic uh more than two that is and so uh there's some sort of we don't know whether it's strong positive strong negative and how strong we're going Say yt is going to be explained by itself, right?
So some y intercept and its previous cell. Correct or not? Yeah, it's some sort of a 1% of its previous value. Correct? Except we don't know it's 1%. Maybe it's 1.133%. Maybe it's 0.994%. We don't know. So we're going to let the data tell us. But the whole thinking don't you think is exactly trying to find y intercept and the gradient. So we just use linear regression or yt using explanatory variable a shifted y values and that's it.
And because these are numbers in real time these are numbers happening historically from y of t it is possible for us to do it. That means we're not twisting time. We're not trying to use future value, future GDP to back project to present. We're not doing that. Right? So this is completely feasible. We are not uh flipping time upside down. We are not doing something that is uh against physics. So we can and we will do that.
And so in what happens next is that uh we are going to do exactly this. So first thing first is we will use y of t minus one to regress with y of t or just to put in place explanatory response explanatory response like this. So y of t is very simple. We just copy 7 6 8 9 7. Yeah. And y of t minus one is you see we would like to use the previous value to explain the next value. So we like to say GDP of year 1 explains GDP of year two not the other way around.
So don't shift wrongly. Don't shift upwards. end up you're going to use I don't know seven to explain five then you're talking about using year four to explain year three GDP then we have a problem right so shift it down in a sense that we want to say use six you see by now we lose time already we don't have the t involved in our model here we don't have a time and we don't have a time then then we are not implanting another potential source of time related problem.
You see, so no doubt yt and yt minus one they themselves the values themselves are affected by time but time is not in our arena in not in our model. So we are not introducing the time related issues for linear regression. So anyway using previous value to explain five six explains five five explains seven right. So after a while you get to see that it's again shifted like the way we calculated our auto correlation. Okay.
So, uh importantly it is this way, right? You don't shift the other way around. Then we will be in trouble. We are like saying using next year's GDP to explain this year's GDP, right? Then you don't need a model that's just talking about history, right? Okay. Anyway, what do we do? Very simple. We just pipe into any linear regressing uh device which by now you should be very very familiar and uh we would get we get b 0 equals to uh 51077 [Music] that's what I get and b1 the gradient which is the growth percentage remember 0.2769 um in this case it will be bad news as GDP right it's like saying that next year's GDP is 27.69% 69% of previous year's GDP as was collapsing country.
So uh we hope we are not in that situation but uh end up I think if you plug in GDP yeah your your B1 the gradient will be then interpreted as growth rate like like compound growth rate if like 1.01 would mean 1% growth every year. So uh yeah so that would be how we can we can expect the numbers for now it's just uh uh you know regressing using some randomly generated integers but we see that we can then calculate all these again right so we will calculate uh e of t same thing just to be fair comparison and uh e of t minus one.
Yeah. So, E of T we will be creating uh I need to shift the values because I in order to get E of T I need to get the predicted value and I think it's good idea to show you a few of those predicted values. So, first predicted value is uh yhat of t Right. So, which means we plug in the six multiplied by B1 and then add the offset. So that whatever x value we use to derive our b 0 b1 we would now use back the same x value has to be so as to get the predicted value for the same x uh for the same explanatory variable value.
So if you do that then the first number you get should be 6 7692 just to give you another number 6.49 4 923. All right. And then you keep going all the way down to the last number. I get 7.6 exact. Okay. Now we can calculate the e of t because it is y hat of t minus yt. So this we can calculate uh we get 1.7692. Okay, another number to go with minus0.577. All right, and then let's jump all the way down to the last number which is exactly 0.6.
And we also need E of T minus E of T minus one. Right? So that's why we cannot square it right here because we need e of t minus e of t - 1 not e of t^ 2 minus e of t - 1 square. So so we rather keep the signed e of t. So which is to say um in order to use a order one right AR1 AR1 or auto auto reggressive order one model we lose one time unit right because we need the the sort of like we wait for one year of GDP so that we can start saying okay next year's GDP is going to be 1% additional to this year's right so for this year if we are starting this year starting our records this year we cannot predict for this year so we have to skip one year so so our prediction also will skip one year because we don't have a x to go with so we will skip that and because of that our error will also skip one year and now for error to happen we have to skip another error which means now we will skip two years because only after uh uh or rather only at the end of the third year we will have this uh second error term here which is05077 to minus of 1.769.
So when we do this subtraction we get and now we lose two terms in this column. Right? So we lose two terms. Let's just note that. All right. So that is the structural gap introduced by our formulas which exploits the orderliness of the time series data. So this is time series data still because there is time involved but this time involvement doesn't plague us because there is no time variable directly participating in the explanatory variable right so it's okay nevertheless is to justify for the meaning of ET minus ET minus what does it make sense to say the previous value answer is yes it still makes sense right but just to give you another confirmation I will list the next error 5538 and all the way down to the last value which is 2.2769.
All right. Okay. So just like before we will obtain the sum of e t² right uh again we'll we'll square each number and then add them all up right so for this I get 9.1692 and the summation of e t minus e of t minus one squared I get 18.166 so our derin Watson value will be the sum of e minus e - 1^ 2 so it's 18.166 66 over sum of e squ. So it's 9.1692. Okay. All right. So now what happens? I get 1.98 9813. Okay. Now uh very very honestly I didn't massage or tweak the numbers such that ultimately I get such a nice Watson number like almost exactly on two which is to indicate that no matter what's your alpha you're not going to be able to reject right so there's no autocorrelation there's no autocorrelation this one depending on the alpha you choose and the critical values Right?
If the critical value is maybe 2.2 2.3 2.4 you know it's just above. Right? So then you will end up rejecting. So or may not reject. So this part not so clear but there's a slant towards the four side which means non uh not totally non autocorrelated. That means some autocorrelation but here it is unlikely to cause any rejection. Thereby we conclude that there is no auto correlation. Right? So in this case uh this implies likely to be some autocorrelation.
All right. And this means no auto correlation order one of course right and this is great because uh we uh we are now able to then very very uh sleep well because we we know we're not going to be plagued by potential accusation that hey you know what uh GDP P data you're not supposed to use time variable right in your linear regression models so that uh autocorrelation may come in did you check for it? Yes, I did check number one.
I didn't use time variable even though my data has some sort of a timebased sequence the yt and the yt minus one. And number two, if order two uh uh if if order one autocorrelation exists despite using this, right? despite using the order one I still have order one autocorrelation then it would show up and end up it didn't show up right so I am totally sleeping like a baby I'm not worried at all right so this is a very good model right to account for the nonlinear growth in GDP as time passes right so the GDP grows and it grows at constant stant rate like 1% or more or less constant rate plus noise but that would have been totally accounted for by the C1 value here and of course uh now that I said this I should change this B 0 to C 0 and B1 to C1 so that it makes sense okay so that's our AR model AR model auto reggressive model of order one R1.
Okay. So advantage is whenever your Durban Watson shows that your original uh model where's our original model right this model our original model has uh or is suffering from some sort of auto reggressive behavior and in practical purposes auto reggressive behavior typically means order one auto reggressive behavior. Then uh the Derby Watson test statistic will be our lighthouse blinking blinking blinking right telling us potential problem.
And when we try this perhaps because of that right because of this indication by the Wson test we use an auto reggressive model. That sounds super professional, right? But the whole idea is very simple. Right hand side use back the response variable. But the previous time stamp has to be previous, right? And you can also say plus c2 y of t minus 2 if you say that there is order two influence, right? But you're going to test for that.
And uh when you do this this becomes auto reggressive model. Okay. So it sounds tough but it's very simple idea. And when this happens, we have the benefit of testing it again. And we see that it no longer is uh you know suffering from autocorrelation in the residual in the residue which is which is stronger, right? which is the reason the the conclusion from this residual analysis is more generally applicable to uh to the model right to say that the model is not having unaccounted for uh excessive growth in the residuals.
So that means the model is more useful to us, more reliable, the integrity is higher. Then for us to analyze the y of t which is to just calculate the autocorrelation. The autocorrelation is just the y of t just the standalone y data values right and that is enough for us to calculate. So we can ar arrive at a conclusion of uh positive or negative autocorrelation but that's just like saying that we obtain a sample we take the sample mean if the sample mean is large means the population mean is large if the sample mean is small we say the population mean we're not doing hypothesis test not uh a very allrounded test right to uh make conclusions about the population and number two we haven't even attempted to account for or to regress a model for y of t right so we have no regression model but over here was test requires us to first make a model okay you can make a simple model difficult model right then was number crunch the residuals like dig out from dust bin all the trash and see if the trash still has remaining food, right?
They the food is still solid, right? And they are not expired and you throw them away, right? So that is going to cause a problem and the minor will then reject the the null hypothesis and make us think again, right? So we think again when you reject the null hypothesis under the bin Wson test you can then consider a uh auto reggressive model typically of order one and that means use the response variables previous self to explain the response variables current self current value in this sense.
Okay. So with that, I hope you have uh expanded your statistical skills to be able to deal with more complicated data, especially when time is part of the analysis and time is almost always part of everything's analysis cuz uh time is money, right? And business is always involving money and time also affects our life, affects nature, affects everything that we measure. So armed with this skill then hopefully you can deal with these data a lot better.
Thank you very much for your attention in this video. See you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.