Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Tyler Reny · @tylerreny9233
Words
9,699
Runtime
59:39
Speaking pace
163wpm
Reading time
40min
163 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
hello everyone so this week we're going to work on learning a little bit about plotting I'm going to predominantly teach you how to use ggplot in a longer version of this script which I teach to my spe-489 class which is just on like a more intensive programming class some of you in this class have taken it I teach compare base R and ggplot much more frequently than I do here for this one I'm just going to teach you
82 words, the words spoken in the first 30 seconds at 163 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 8 |
| Average words per sentence | 1212.4 |
| Longest sentence | 2,552 words |
| Questions asked | 0 |
| Sentences containing a number | 8 |
Most used terms
Filler phrases
120 in total: like 38 · uh 21 · um 15 · kind of 14 · actually 13 · basically 10 · sort of 6 · you know 3.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
hello everyone so this week we're going to work on learning a little bit about plotting I'm going to predominantly teach you how to use ggplot in a longer version of this script which I teach to my spe-489 class which is just on like a more intensive programming class some of you in this class have taken it I teach compare base R and ggplot much more frequently than I do here for this one I'm just going to teach you ggplot because you might use a little bit of it for your final project so I want to introduce you to the syntax give you a sense of how it works and then you can explore more from there and decide if you want to keep using ggplot or if you keep on with r whether you want to use base plot instead base art plot so go ahead and open the file we'll work through this together the three main packages we'll use today are tidy verse The Gap binder data set and our color Brewer if you don't have this you can install it I've mentioned this to uh I might have mentioned in this class I don't remember but this is a tool kit that allows you to draw from pre-specified color palettes that are optimized for plotting so I'll show you what this looks like when we get to to Colors later um some of this tutorial was adapted from this art Graphics class at uh at Harvard I'm not even sure if this still exists but it did at one point so you can check it out if you want or search for um search for this class and see if you can find it there's there's tons of resources on this stuff online so um I don't don't hesitate to seek out more information more practice if you need it okay what is ggplot ggplot is I think it's the most widely used package in R so huge user base lots of resources out there for you um it's based on this kind of this idea called grammar of Graphics there's a whole book here that I linked to if you're interested in learning more it's a it's a a an approach to programming graphical displays of data um and it's it's built around this idea of an additive layers that build up a plot starting from Basics and add on each each tweak or each additional thing you want to add you add on just piece by piece it's extremely flexible it's a a theme it has a theme system so you can easy easily and consistently alter any individuals of the plot using very little code and as I mentioned many users so lots of help okay so as I said you independently specify the building blocks of graphics and combine them together to have whatever graphical display you want and so this usually includes things like data in the case of ggplot your data either has to be in a data frame or table you cannot plot otherwise this is different from base R base R you can feed two vectors in and plot them against each other in ggplot you can't you then specify an aesthetic mapping what data do you actually want to put on the plot so it doesn't know from this data frame you provide what you want your x-axis and y-axis to B so that's what you specify in your aesthetic mapping and then geometric objects so what sort of plot do you want you have the variables you want from the from the Aesthetics but you don't know if you want to turn those into lines or points or whatever then do you want to transform the data somehow do you want to alter any scales visuals colors all that stuff so each of these gets added on as an individual building block here the basic syntax and structure here is that you start with your ggplot function and within that you specify your data so the data first comma and then your Aesthetics AES parentheses and you see the the within those parentheses you specify any X variables your y variables your grouping variables if you need them coloring variables if you need them I'll walk through what that means and then notice the plus sign so this is the building block so this is your base this is like your blank canvas it's not going to plot anything then you add on your geometries so that's going to be your geometric object what sort of plot do you want so your GM underscore points is going to be points your G of underscore line is going to be lines and you can either pull this out of the original function here so just if you don't specify anything it'll assume you're using this data set and this aesthetic grouping or you can change it here and specify different things and then you can add on any sort of modifications to your plot so this could be color this could be x-axis y-axis ticks labels names anything you want and then a bunch of theme options where you can specify and tweak little pieces of the theme of the plot so I just walked you through each of these so GDP let's start let's play around with it and see if you can pick up the syntax it's pretty straightforward so let's start with this ggplot code so I mentioned this like sets up a a blank canvas so to speak and so here I'm using the gabinder data set just to remind you this is country year data it has life expectancy population GDP per capita and here I'm specifying the Aesthetics I want my x-axis to be logged GB per capita and I want my y-axis to be life expectancy so when we run that you can see here that this plot pops up and this is my blade canvas there's nothing on here yet but it specifies the range of the y-axis and the x-axis based on the data inside of the plot that I that I provided so a little note here when you feed any data to ggplot as I just said above the data needs to be in a data frame this is not true of Base R plots so here's let's say we have two random draws from standard normal distributions if you use base plot which is just plot your X and your y you can plot these you can see here the scatter plot if you try this with ggplot you'll get this error data must be in a data frame so you have to make sure that the variables are together in a data frame they're two columns here I can just go ahead let me get rid of that I can go ahead and put the data in the data frame and then plot it and you see we have identical plot in ggplot versus base r so pretty straightforward you can also if you don't want to feed it in so ggplot data frame comma Aesthetics you can also pipe it in and then you don't have to specify the data set here so it reads in as the first argument and then you just have the Aesthetics here so that's going to be identical so let's have your first exercise so you're going to use ggplot to create a canvas of a plot from The Gap minor data where X is log GDP per capita and Y is life expectancy go ahead and run the code and look at the empty plot so I'm going to go ahead because I'm going to do this I'm going to save this as completed so you have both all right so go ahead and pause your screen give this a shot or pause this video give this a shot see if you can write the code yourself see if it works and then go ahead and unpause and we'll do it together so I'm going to go ahead and say Gap minder I'm going to pipe it in ggplot my aesthetic mappings X is going to be logged GDP per capita y equals life expectancy all right so there's our empty plot so we we already did this Above This is not this is not new this is not anything breakthrough that's our empty empty pattern so next thing you want now you have to decide what's actually going on with your plot are going to go on your plot rather so if you go back to the original example what do you want to plot here points right we want to have a scatter plot and so we're going to add on the plus sign here so take this code from above add the plus sign and then add G on point and we'll draw the information needed to add points to the plot all right so there's our scatter plot now one thing you'll note is because we have year country year data there's going to be weird like patterns and correlations here because these countries are generally going to follow their own individual Trends and so if we were actually going to plot this and look at it because we were curious about the relationship we would want to subset this to a specific year first filter here equals equals I think there's a 2007 in here let's give that a shot and then get rid of that so here we take the Gap minder we filter it and then we pipe it directly into the ggplot and just note that you have to change from the pipe once you're in the GD plot you have to switch to a plus so it's the same kind of logic of building blocks but here we're piping all the way to the plot and then we're adding on the different pieces okay so there's your it looks similar but there's your true relationship for one country your observation for each each country rather than having a bunch of them okay exercise use the Geon point to create a plot with points where the x-axis is locked population and y-axis is log GDP per capita for data just from 2007 you can see you can copy and paste what I did here well you can't can't copy and paste it because you're just watching the video but you can type this and go from there so go ahead and pause and give it a shot and then come back and we'll do it together okay so let's go ahead and pipe it all right and then we're going to ggplot so x equals log pop y equals log GDP per capita and then we want G on point all right there you go you see basically no no relationship there if you ever want to just check and see what the correlation is roughly if you can't eye it we'll do this again a little bit later in this video but you can throw a linear regression on there and see if it's just flat which it basically is so more or less non-correlated let's just go ahead and check the correlation coefficient because I'm curious so cat binder now let's lock it log Gap minder pop log Gap finder GDP per capita and for each of these I want this to be here 2007. all right and this is going to throw errors because we have oh we don't have any missing data so there you go correlation is basically zero negative 0.04 um that's not going to be statistically significant you could also do this a linear model where Y is logged actually we're going to do X is log pop and Y is log GDP per cap so let's just go ahead and flip this around I'm just going to write this out right here because it's going to be easier okay we're regressing GDP all right we're going to log that on I don't worry about the syntax I haven't taught you any of this yet or regression syntax we'll go through this later pop data equals Gap finder cap minder here 7 comma all right there's our negative 0.04 we're going to wrap that in summary just look at yeah that P values 0.59 so straight up flat line um again don't don't worry about what I'm doing here I was just curious what that what that coefficient would be and it's obviously visually not statistically significant okay so what if you add the wrong geometry let's look at what happens when you add G online to apply this you just have points it's going to go ahead and try and connect them right so that makes no sense sometimes you do want a line though but it depends what you're doing or what what you're trying to plot so oftentimes you're going to have to rely on using D plier first and then pipe that into a ggplot so let's say we want to plot mean logged population size let's calculate that with d pliers feed it into GD plot so here take the Gap minder I'm going to add by year here okay Group by year summarize mean pop mean log pop so here's just taking each year in calculating the mean logged population and now I'm piping it into ggplot here I'm putting here x-axis mean pop on the y-axis and put the line through it all right so there's your basic there's your basic line plot okay so exercise here go ahead and use D plier to calculate mean GDP per capita by year and plot it over time go ahead and pause this should be pretty quick now give it a shot and then come back and we'll do it together foreign ER we're going to group by here summarize just call it y mean GDP per cap let's make sure that works great pipe it in now our x-axis is here our y-axis is y x equals here y equals y Geon line and there we go all right I'm also fond of dot plots with confidence intervals this is something I use a lot for my own research so let's create a Dot Plot that has shows the mean GDP per capita in 2007 with 95 confidence intervals across each continent so the Ci's here are going to account for the uncertainty given the number of countries and the variation of that variable with incontinence itself this is a little bit more complex so don't worry if you don't follow it at all all I want you to take away is how you actually put these confidence intervals and points on a plot but what I'm going to do is use D plier first to calculate means and confidence intervals then I'm going to feed it to GD plot with a new genome called GM segment that I use all right so you start with your data filter Year 2007 then we're going to group by continent and within each each for each continent we're going to calculate the mean GDP per capita and then a lower and upper bound on that confidence interval which if you dig back into your stats books is going to be the mean minus 1.96 times the standard deviation of the variable divided by the square root of n and so that's taking into account both the variation of the variable itself and your sample size and that's going to be your lower confidence interval that's going to be your upper confidence interval so if we just run that really quickly and look at it oop I included that type didn't mean to do that let's do it all right we can see for each continent we have a mean and then a lower bound and an upper bound on that confidence interval so we're going to go ahead and pipe that now directly into ggplot so here we're going to I'm going to reorder it here and so this function our X variable is continent but I'm going to reorder continent to range in terms of largest to smallest value on mean GDP per capita you can do it without this but I'll just show you what this looks like y equals our mean GDP per capita we have our y Min and Y Max which I specify within the aesthetic call here that's going to be our lower bound and our upper bound and then we add the Geon point and then we add the geom error bar this is going to be our 95 confidence interval I'm going to specify width because I'll show you what that looks like it's weird in a second I'm going to use this chord flip code which is going to flip our coordinates around uh so it's going to take our x-axis and basically put it up here on the Y and it's taking a y to put it on the X and I'll show you why in a minute and then I'm going to specify variable or labels for the x-axis and the y-axis I'm going to get rid of the x-axis variable and for y I'm going to say mean GDP per capita so let me show you each piece of this let's first put in points all right so because I'm reordering it it's going to start small here and take upwards so it's ordered now if I add the GM error bars in you can see it adds these error bars now if I let the default width so if I don't specify this and run this you can see it's like X-Wing fighters from Star Wars I don't know why these are so fat like I would not want to publish a plot with that so I just specify the width here all right chord flip is going to flip us around so remember it was this was on the x-axis this is on the y-axis um and then I'm going to clobber this label here because we don't need this we can look at this and understand that these are continents without labeling it and I'm going to change this so it's cleaner so instead of mean underscore GDP it's going to say mean GDP per capita and there you have it so these are means with 95 confidence intervals okay so let's go ahead and see if you can understand this code copy it and paste it and make your own plot like this so here I'm going to have a mean a plot of mean television watching by income category and so for folks making less than 25 000 more than 25 000 or have not reported an income and we're going to add confidence intervals to these estimates so I'm going to take this from the GSS data which I don't remember if we worked with it in this class but this is a public opinion survey I'm going to first make some changes to the income variable I'm going to create this categorical income so we have three three categories I'm going to group by that and then I'm going to again what I did above calculate mean and then lower and upper confidence interval and save that as a project as a data frame and so folks making less than 25 000 are watching about three hours of TV a day folks making more than 25 000 or watching about 2.2 hours a day and folks who don't report their income or watching about 3.6 hours per day so your job here is going to be to turn these three points into a plot with error bars like this so go ahead and pause the video give it a shot and then come back here and we'll do it together okay so I'm essentially just going to take what I have above and modify it okay so our x-axis is going to be income category and then I'm actually I'm not going to recode this I actually don't like the ordering so we'll fix that in a minute but I I didn't ask you to do that so don't worry about that but I'll show you what that looks like here our y-axis is mean TV our upper and lower confidence bounds are the same because I called them the same G on point GM error bar so let's just give that a shot first all right that works let's flip it around with cord flip all right so now it's currently more than 25 missing and then less than 25 here's our mean hours we see pretty tight confidence intervals there both because this is your bigger samples and there's probably not as much variation and so mean TV hours and we're going to get rid of x so let's do that okay so last thing we're going to do and I haven't taught you this yet unless I've taught you factors but I am so confused at this point what I've taught everyone across classes so let me just show you how I would go about reordering these not in order of lowest to highest but instead we want to manually change it so that probably more than less than missing at the bottom so mutate I'm going to take this variable income cat I'm going to set it as a factor variable and this is going to force it into a specific order that I specify then I'm going to specify levels and so we want let's do missing first all right and then less than 25 and then more than 25 and I always forget the order here so let's let's give this a shot and see what it does all right more or less stressing there we go that's what I wanted if it was the other way around you can wrap this in reverse and then run it it's going to flip it so that it does missing then less than more but I'm happy with that ordering so we'll get rid of that all right so more than 25 000 are watching less than 2.5 hours less than 25 we're watching what three hours missing data over three and a half okay if you want to add uncertainty to lines rather than points you can use the geom ribbon instead of GM error bar so let's see what this looks like I'm going to go ahead and just show you with a an example of a regression model here so I'm using an ordinarily squares linear model and I'm regressing TV hours watched which is continuous variable on age which is a continuous variable and here we see that for each year increase in age we have about .02 hours of additional TV being watched this is statistically significant I don't have controls in here so who knows if this conditional correlation would hold with other controls and and who knows if like there's some sort of an issue with endogeneity which I suspect there is there's all kinds of things correlated with age that are also going to shape TV hours watch so this is clearly not a cause of relationship but just for for an example here I'm going to show you how to how to do this so here we're going to create this is going to be a a predicted value plot so you may have heard predicted probability plot here we're not doing a probability we're predicting the number of actual hours watched for each age in this data set and so we're going to set our range on our x-axis that we want to calculate y hats for our predicted TV watched based on this model output and here we see we have respondents that range from 18 to 89 in the data set I'm going to use this predict function don't worry about this I'm not expecting you at the even at the end of this class to be able to do predicted probability or predicted values but predict is going to take the model object it's going to take a data frame saying we want to predict a y hat or a predicted TV watched hours for every age that we've laid out here from 18 to 89 and we're going to have uncertainty around that and here if I go ahead and run that we're going to have on the fit is going to be our y hat our predicted value our standard error.fit is going to be our uncertainty our standard error around that prediction and then you can ignore these I'm going to go ahead and add on an X variable which is going to be our x-axis for our plot fit is going to be our y-axis and then we'll calculate lower and upper confidence intervals here using lower equals fit minus 1.996 times the standard error and upper is Fit Plus 1.96 times the standard error and if you go back up here this is the same thing right the standard deviation of the variable divided by the square root of n is the standard error so we're just doing the same thing here and then I'm selecting just the variables we need and I'm piping that directly into a plot where I put x equals x y equals fit y Min equals lower y Max equals upper and then I'm going to add this G of ribbon and so in the Align over that so let's just go ahead and look at that all right so here's our prediction for an 18 year old predicted to watch less than 2.5 hours of TV per day and as you go up in age 289 predicted to watch you know three point seven five three point eight somewhere around that hours of TV per day of course this this lots lots of things you can do differently with this model it might not be linear even let's just plot the raw data because I'm curious ggplot GSS cat all right X is going to be age Y is going to be TV hours point and I'm just going to put a gamp smoother through here to see if it appears to be linear all right so we have oh interesting all right hold on I'm gonna I'm going to go ahead and make these little transparent so it's easier to see this line all right so basically we have it's not quite linear we have this dip here around 40. and then it increased around 70 and it starts to go up from there and so this is you know my guess is you know teenagers watching TV and then when you have children and families you're watching less and then you gradually increase over time once kids move out of the house Etc if we're assuming traditional traditional uh kind of family life here so that's kind of interesting and you have all these crazy outliers right some people reporting that they watch 24 hours of TV a day which makes no sense um but anyways beside the point what I wanted to show you is how to go ahead and make this make this plot here whoops try it again all right so this geom ribbon is going to put this uh this these coffee or these error bars here around the line and I specify Alpha equals 0.4 because if I don't it's going to be completely opaque it's going to be hard to see and then I put the line on top of that and I change the axes here to be age and the probability of average TV hours watched which is not actually right it's estimated so our expected value here all right so that's your GM ribbon so here I'm going to do the same model but this time I'm going to dichotomize TV watching to take the value of one if the respondent watched over four hours of TV per day and zero otherwise I'm going to run a logistic regression and then estimate predicted probabilities of 95 confidence intervals so go ahead and plot this line with the confidence interval using Geon ribbon from this code above so run this this is our dummy variable over four hours or less than four hours here's our generalized linear model this is going to be a logistic regression and we see this coefficient here is positive statistic significant but it's a logistic regression so we can't actually interpret this so let's go ahead and do our predicted values or predicted probabilities in this case here's our hypothetical ages again 18 to 89. here's predict has data frame mutate upper and lower all right so here's your data that you're going to use to plot this sort of line so go ahead and pause give it a shot yourself and come back okay if you plot a yes ax equals x y equals fit y Min equals lower y Max equals upper geom ribbon I'm going to set the alpha to 0.3 actually online all right so here we're now no longer working with the linear model where you have this is because we're working with a glm a logistic regression here we're no longer doing uh forcing this into a linear form and our interpretation here on the y-axis is now Labs x equals h y equals what I had before which is probability so we're back in the probability world uh watch four hours TV or more all right so the probability an 18 year old watches four hours of TV Premier per day is about 10 the probability that someone's 89 watches this much TV is about 30 percent okay so you've seen G on point G online GM error bar and GM ribbon there's some more here that might come in handy here's some just toy data so I'm going to show you what these look like all right so here's our blank plot here uh the data is just 315 y values two four six and then we have a label associated all right so we've seen points GM text will replace our points with text from the data frame so if you have a company label so in which case here we have the first point is a second Point's B third point is C you can do a g on bar in this case you're going to feed in GM bar stat equals identity meaning it wants to take those values 2 4 and 6 and apply them across the points here giant tile line which we've seen area path polygon I've never used tile area path or polygon so you can I I don't know when those might come up but if you want to draw shapes in your plot you can of more use than those last few I just showed you are probably going to be a histogram so here is a this is just random draws mean zero standard deviation one here's Gonna Be Be Your histogram so you just take this data in your as is your Aesthetics here we're just specifying one variable because we just have one one variable is X and then GM histogram is going to be a histogram and I like to specify fill equals white color equals black which is going to be white inside black lines on the outside if you don't specify this is what it looks like and I think it's kind of ugly so I do that maybe like box plots GM box plot it's much better looking I think than the box plot and base are which is kind of ugly but it gives you a sense of distribution and then violin plot which is kind of useful it shows you the distribution for each variable so you kind of see mid Max and then how how the data spread here I'm just filtering the data by 2007 X is GDP per capita Y is continent and so essentially it's just grouping those and showing the distribution for each okay so what makes ggplot really powerful is that you can do all this kind of grouping by variables and that's going to allow you to plot a bunch of different things by groups on a single plot I use these every day this is something that's incredibly useful so practice it make sure you understand it and it will make your life easier too so grouping allows you to group data by some variable in your data set and then have these correspond to just a generic group or some distinguishing characteristics so it could be color fill line type shape depending on what how you want to customize or distinguish between groups so let's say you want to plot mean logs GDP per capita over time by continent this requires a little bit of work with DD D plier first but let's say you have Gap binder you're grouping by continent and year and here you're calculating the mean log GDP per capita across each continent by year and here you see this little data frame it's going to have continent year and then mean log GDP per capita and so if I spec if I plot this without specifying groups it's not going to make any sense because we basically want a line to correspond to the trends for each continent uh over time and so you put in group equals continent here and now it's going to split it out and just have group have lines correspond to each of the different continents so this is better but we don't know what line corresponds to what so instead of group equals we're going to take the same code here and say color equals content and what it's going to do is slap a little Legend on here and it's going to specify these by by color so you can look at the legend and you can see how they correspond um if you don't want to use color you can use LINE type which is going to allow you to do this in grayscale and you can compare to the legend if you want to do color and line type which is usually what I do so it works both in color and if you are in grayscale and like a journal article or something and here I had how you do it in in base R but I'm just going to get rid of that okay you can also if you don't want a group like this you can have what are called small multiples so we use facet wrap facet wrap is a function that's going to allow you to make little plots for each different group so let's look at this pattern of log to mean GDP per capita across each country over each year or by year across each country in the data set so here I'm just gonna for to make this easier for us to see I'm going to just look at the countries in the Americas and here you can see if I didn't specify the grouping variable it makes no sense and so I'm using facet wrap and then tilde country so I want to group it by country and so you can basically produce a little line plot for each country in your data which is much easier to now compare and say like oh interesting it goes up for Canada down for Haiti flat for Uruguay Honduras right um and so this is a really nice way to visually look at your data compare across units in this case countries and pick out any visual trends you can use color and shape to same kind of effect here so here I'm going to filter year to 2007 I'm plotting cutie paper cabinet and life expectancy and I'm going to group by continent so here's points where we see colors correspond to countries on each continent here's the same thing with shapes instead so you can see by default it sets your shapes color and shape which is a little easier to see right and then if you want you can also do like bar plots I don't like this as much but I filtered it to just four different countries and you can look at visual Trends right over time compared between these another useful feature especially for exploratory data analysis is stat smooth this is going to throw a smoother either a Lois which is a local regression line or gam smoother which is like a local regression so don't don't worry about that stats too much just basically it's going to show you visual trends of how your data is related over time so here's from 2007 x-axis log gdpr per capita y-axis life expectancy and then you put this smoother on here and you see this is roughly linear and increasing monotonically if you want to then you decide it's a linear model instead of stat smooth here with SC equals false I just shut off the standard error I set method equals LM and it will put a linear model on there for you and if you say this as an object you can actually extract that data too that's drawing this letter bottle so you could use that okay exercise using the Gap minder data set plot log GDP per capita on the x-axis life expectancy on the y-axis use facet wrap to make small multiples by continent and then add a best fit linear regression line to the data in each plot what's your takeaway here so go ahead and pause give that a shot and we'll come back together all right Gap minder I'm gonna oh we're not going to group by we're going to go straight into ggplot all right X is logged GDP per capita y equals life expectancy G on point and then we're going to facet wrap it by continent let's do that first and just make sure it works all right so we should technically because we have data over time I'm going to also filter this just to the last gear in the data set just to make more sense okay and then we want a smoother method equals LM all right so basically we see a similar slope over time that has log GDP per capita increases across each continent we see similar increases in life expectancy which is kind of interesting right you might have reason to believe that in some continent that the relationship would be more or less robust but it's not Okay so we've learned to uh we've learned to do all that let's say we want to make some more modifications so here we're going to play around with this function called Labs where we can specify our x-axis our y-axis our title our subtitle right so x equals log g to p per capita y equals life expectancy here's title and here's a subtitle to a sense of how that all works says syntax is easy let's say we want to change the colors so here's GDP per capita by year across continent um here I'm going to draw from this is from the our color Brewer package I can set my palette to dark two which is just one of their many palettes I'll show you those in a second and then labels X and Y variables right so we have a different color theme Here let me show you what this looks like I'm going to go ahead and do a new window our color Brewer all right color advice for maps all right so these are all the different palettes that are offered in our color bird which is used for like cartography and plotting and each of these corresponds to a different argument you can feed here for palette so that's just something to keep in mind if you ever want to play around with this colorbrewer2.org you can look at color schemes these have been optimized to be the ideal for your eye for like cognitive uh processing of different colors and there are other things you can do like you can specify that things are colorblind safe print friendly whatever you can have diverging color schemes you can have qualitative color schemes so things that are uniquely distinct you can specify the different number of uh of classes here between lowest and highest so I recommend playing around with this and checking it out you can specify your own colors so here I had scale color Brewer so this is drawing a palette from color Brewer here I have scale color manual meaning I'm going to specify the colors myself you can use either hex values here or you can use right out the color names if you do uh our colors you can see what these what these colors are there's tons of tutorials here that shows you the different different names for all the different colors so you can spend a lot of time uh looking at this stuff so here I have two hex values green blue dark blue and then I put labels on here all right so here are my uh blue dark blue green and here are my hex values so this orangish color and this purplish color in these scale color manuals you can also change the stuff about Legends here both the name of the legend and the names of the different categories in The Legend So in this case continents so here I'm going to but I'll show you that in a second let's look at different line types now and so scale color manual scale line type manual so if you do grouping by line types we can see this is the default ordering but I can change the values here to be different different sorts of ticks here and you can go online and check out those different values here's here's what I would say in terms of changing the legend here so you can set values here scale color manual name equals continence Capital C so that's going to be different from lowercase and this doesn't make any sense you're not going to do this but if you wanted to type out different names here to correspond to the contents you can do that as well you can see I just changed the name and then these labels all right there's dodging and stacking options so this is going to come most in handy when you're looking at uh bar plots so we're going to use the miles per gallon data set and if you want to Stack lines next to each other rather than have them on top of each other you can do GM bar position equals Dodge it doesn't make as much sense in this case because there's some categories here that we don't have different labels for but if you did ever want to specify dodging things you can set position equals Dodge in a function um sometimes we want to dodge points that correspond with group estimates so here we're going to go ahead and group by continent and then we're going to create a dummy variable as data before 1980 or after 1980. we're going to summarize the mean life expectancy before 1980 versus after 1980 and here normally your confidence intervals would be 1.96 but I'm making it bigger just to give you an example of overlapping confidence intervals we can see here what this data is going to look like and then I'm going to plot on our x-axis continent on our y-axis the mean our lower and upper confidence interval and I'm going to use color to specify whether it's before 1980 or after 1980.
I'm going to slap on some error bars and let's go ahead and I'm going to show you without the dodging first so I just put on error bars I put on points and then I flip the coordinates around all right so if you had points here that are grouped right in this case colored and we can see that the lines the confidence intervals lie right on top of each other and so you can't actually see how big that confidence interval is we guess it goes probably to about here but let's say we want to see these separately so here we can specify in the error bar and the point arguments a position equals position Dodge and then we specify the width and we can see it's going to nudge those apart and so they're not touching each other and it's easier to see the data this way you can still see that it's grouped by continent but you can better see how the data exists side by side all right custom axis scaling this is where we can change the range of the axes so here let's say for example I have data on Afghanistan 1950 to 2007 we have a date on the population size I divided it by a thousand just to make it more tractable we do the same thing for China we see they're both similar increasing Trends but we see these y-axis are drastically different right China is a much larger country so if we wanted to plot these on the same plot right or next to each other to get some sense of the sort of scale of the difference you can specify a scale y continuous and that's going to change this y-axis and here we're going to range between 10 000 and 1.35 million and if we put the Afghanistan data on this we can now see this was the range for the China data we can now see that this increases mini School compared to what what's happening in China and if we put them together using facet wrap it's going to automatically force that y-axis to account for both and you can see how just drastically different these population growths have been if you ever want to put stuff into a facet wrap like this you can allow this y-axis to float meaning it will be different for both of the two y plots so maybe you're just interested in the overall trend it's hard to see the trend here because it's so flat but we let's say given the different baselines let's say if they're both going to increase similarly here's our float y you can put these next to each other now and have the y-axis floating so we can see maybe something we could have missed in this plot is maybe there's like a period of drop and then increase but you can't really tell this is going to show you that in fact there was a period of drop uh in the just after the 1980s and then it increased again and so that allows you to see a little bit more data depending on what you're trying to show there are also all these plot styling functions or options you have where you can add different themes so let me show you what these look like so here's your standard theme that you've seen so far gray background you can have theme black white which is going to get rid of that gray this is what I use for most of my plots minimal is going to get rid of the whoops nope that's Gray minimal's going to get rid of the lines all together and the gray gray is what the default is line draw is going to be sharper lines here on the outside theme light is going to be lighter lines on the outside theme dark because it's going to be a dark gray which I don't like themed void basically nothing and you can play around and add pieces back in if you want to clobber everything and then put in just like the X and Y axis or something okay so there's your rough introduction use this as a tutorial come back to it when you want to do some plotting here's a final exercise for you I'm going to give you some real election data from 2016 at the county level you've columns represent Latino as well as vote going to Clinton and Trump at the county level your task is to make this data long pivot longer which I'll show you how to do because I don't think I've taught you that yet so you can plot vote return on the y-axis against percent Latino on the x-axis Group by candidate and then you can specify that the Clinton points are going to be blue and the Trump points are going to be red and you'll add a lowest smoother to this so if you go ahead and run this code oh don't do that do this select from there all the way down here just run it it's going to create a little table for you it's going to have state county percent Latino Clinton and Trump okay the first thing you have to do is pivot longer so this is a wide data set we need just one variable for the y-axis it's going to be Clinton vote and Trump vote in the same variable and so let's show you how to do that so one two three so we want to take four and five and stack them so pivot longer four to five right is now going to we can see the data frame before was 58 rows when you pivot longer it's going to stack it so it's going to double that 116. we now have a variable a new variable called name which is going to specify which of these y values it corresponds to so in Alameda County the Clinton vote is 79.3 the Trump vote is 14.9 and the percent Latino is 0.226 and so every county is now going to have two observations in this data set and this is going to allow you to plot with an x-axis being percent Latino excuse me y-axis being value a grouping variable called name and then go ahead and make the other changes so pause this video try and do this plot yourself and then we'll come back together okay I'm just going to copy here Clinton Point okay A little smoother all right so we want percent Latino on the x-axis Y is going to be value color is going to be name so let's just try that first and see all right percent Latino here's vote percent so 20 to 80 here 's the name these red points indicate the percent that Clinton got in that county the green points are percent that Trump got in that county we're going to put a lowest smoother on here right so we can see these General Trends if we don't want if we want to force this in linear terms we can use method equals LM which is going to show us General Trends here scale color manual we're going to change this red to blue and this green to red let's try that all right there we go Clinton Trump and let's just tweak this so it looks a little nicer so name equals let's say candidate and then labels equals just capitalize these all right that looks a little bit better I didn't ask you to do all this but this is what I would do our x-axis is going to be percent Latino y-axis oops I need a comma here y-axis is going to be vote pres vote 2016. there we go so I fix those and then let's go ahead and use theme black white all right so here's here's your plot here's your story as the percent Latino increases the county level vote for Trump drops and the county level vote for Clinton increases obviously this is not a clean linear increase as we saw there's kind of like a an increase and then a drop and then an increase and same thing there's like a decrease and then an increase and then a drop for Trump but this tells us our general story and we of course because we're using aggregate data can't really say whether Latinos in these counties are driving this vote Trend but we can say with some uh or we can at least guess that there's probably some role here since the points out here at the highest end are the strongest indicators of what how Latino voters are doing so the question is are Latino voters out here in these districts that are more than 80 Latino similar to the Latino voters and here in these counties that are less than 20 that are 20 or less Latino we don't really know they might they might be different but if we assume they are the same then we would assume that the the Latino vote is consistent and it is definitely less supportive of trump than it is of Clinton all right a few other things if you want to stick around and watch how you can tile plots so here I'm going to create two plots one for China One for Romania and uh I'm going to or I'm going to create data and then I'm going to turn them into plots so here I'm filtering data by China I'm grouping by year and I'm calculating mean life expectancy by year and I'm adding on a label here at the end just to specify that this is data corresponding to China and I'll show you why I'm doing that in a second I save that as an object called P1 I do the same thing for Romania identical P2 all right I'm going to go ahead and plot these so here's the plot for China here's the plot for Romania I'm going to save these as plots too so plot one plot two all right two ways to tile these together there's a package called there's lots of ways actually but here's two ways that I use so there's a package called grid extra and grid arrange will allow you to specify plot one plot two and I want two columns so I want them side by side so that's one way there's a package called Patchwork which is really nice and super easy and using that you just add a plus sign that's going to place them next to each other you can divide them and that's going to place them on top of each other so depending on what you're trying to do lots of ways to do this let's say you want to add a third plot all right you can place these side by side using the Plus or let's say you want to put two on top and the third one on the bottom so Patchwork allows you to do this really nicely in parentheses you add the two and then you divide it by the third and it's going to put them all on the same same tile here um you can add titles too so here's the same same plot as this but I'm going to add title the surprising truth about Mt cars subtitle caption Etc and you can just see the syntax title subtitle disclaimer if you want to save plots you can use this GG save and so let's say we have plot three here we want to save this P3 specify the width specify the height and then the file name so let's go ahead and I'm gonna save it onto my desktop and I'm going to call this test plot PNG there's different file types you can do you can do PNG you can do jpeg you can do Tiff journals will take different have different requirements for the type of plot you do but I can go ahead and save this that's going to write an object to my desktop desktop test plot but let's make that a little smaller so you can see it [Music] so you can see that the width and height are specifying this kind of long long plot if you want to change the size of it it's going to automatically update the text accordingly so it's going to get bigger if you make the plot smaller so let's say width 3 or width for height three let's look at that same plot all right you can see it shrink it and it also made this text bigger and so it's going to automatically try and make that all the appropriate size for you based on the size your applause you don't have to start playing around with text sizes on your on your axes all right and then I have I'm not going to go over this in this video but some basic mapping stuff you can look at if you want um all right I know that's a lot uh my intention is just to kind of expose you to this have you come back to this play around and hopefully you can use this as a as a tutorial moving forward I'll upload both the empty script without the exercises done and then the completed script with the exercises done and you can play around with them and hopefully start to get a sense of uh of how ggplot works okay if you have any questions please reach out happy to help bye
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.