Create Better Graphs | Avoid These Common Data Viz Mistakes

DataCamp · Beginner ·📊 Data Analytics & Business Intelligence ·2y ago

Key Takeaways

Data visualization principles and avoiding common data viz mistakes using effective graph forms and design techniques

Full Transcript

Today uh we're going to discuss how to avoid some common graphical mistakes. Here is the agenda for today's session. At this point, you're probably wondering why anybody invited me here. And you're probably thinking something like anyone who uses a chart that looks like that has nothing to offer me on any subject whatsoever. Well, that's exactly what I think of people who use charts that look like this. I'll tell you in a little while why I dislike this chart so much, but first let's look at the agenda in a more sensible font. Uh we'll discuss a number of ways uh that charts can not communicate well or even mislead deceive the viewer. Uh you can see them here. I will go through them one by one. Uh first let's talk about pie charts. Pie charts are very popular. They don't communicate very well. This pie chart has five wedges. I'd like you to order them in size order. Normally, I'd ask people to shout out or put their answers on the chat, but that would waste time since there's a time delay. So, I'm just going to ask you to order them privately for yourselves, and I'll tell you from experience what I think the answers are likely to be. A few of you probably have no trouble, maybe 10% of you have no trouble seeing the differences. Some of you are very frustrated by this task. You just say they all look the same. Here's another way of presenting the data. And even those of you who had no trouble with the pie chart, I'm sure have more confidence in your answers here. Here you can clearly see that A is larger than B is larger than C, etc. You might wonder why I have no scale on this figure. It's because I showed this to someone and they said, "Hey, that's no fair. You have a scale on your dot plot and you don't on the pie chart." Well, the scale helps you to estimate the values, but that's not what I asked you to do. And the scale is not necessary if we're just ordering them. A graph is not always needed. Uh for a data set as small as this, a table can be very useful. The table here shows you the exact values. It also shows you that the um wedges sum to 100%. That's something that the dot plot did not do very well. We can get around that by adding a column. This is not the same data, but you see the principle that I'm talking about here between the labels and the uh bars. We've added a column with the exact values and the sum. Uh you can do this for any bar or pie chart. And it has the advantage of being visual um with the chart and yet having the advantages of a table. Here's another pie chart. Again, I'm going to tell you, normally I'd say, "What can you tell me about this chart?" Uh, but because um that's difficult in the webinar, I'll tell you what I'm likely to hear from you. Uh the first thing people are going to say is some of the wedges are smaller and some of them are larger. When you look at it a little longer, you'll notice that these are labeled item one to item 10. And you'll notice that the odd items are the ones that are smaller. The even ones are the ones that are bigger. I doubt that anybody sees much more from this chart. Again, let's show the same data as a dot plot. Here you can see that it's a very very structured data set. The pattern of the even ones is identical to that of the odd ones. Uh they're just uh offset by a constant amount. Um each even one is exactly 0.05 05 larger than the corresponding odd point. So we have a very structured data set that jumps at you in the dot plot. Nobody can see that in a pie chart. So again, pie charts do not communicate very well. Two people that I really respect in the world of graphs are Bill Cleveland and Edward Tuy. Tuy in his first book, The Visual Display of Quantitative Information, says, "A table is nearly always better than a dumb pie chart." And he continues to say, "The only design worse than a pie chart is several of them." Cleveland says pie charts have severe perceptual problems compare compared with dot charts. They convey information far less reliably. But I'm going to show you a chart that I consider worse than a simple dumb pie chart. And that's a pseudo three-dimensional pie chart. In this chart, we have four wedges. We know they sum to 100%. And I'd like you to guess, well, to estimate the percentage of each of these four wedges. Again, based on experience, I'm going to say that many people will consider this blue wedge, the one labeled one, as 40%. I've often heard that the orange one is 25, 20, and 15. Let's look at the same data in a 2D pie chart. Here doesn't look the same at all. Here, people are more likely to judge it accurately. Um, you can see that the pseudo third dimension distorts it. The wedge, the side here emphasizes this. Um, the way it's done, the cyan, blue, cyan, whatever it is, appears smaller uh than the orange. Whereas in a 2D, that doesn't happen. Here's the same data in a bar chart. Totally unambiguous, completely clear. The bar chart communicates much better than the pie chart does. So, we've seen that dot plots communicate better. Bar charts communicate better. Um, I'll be showing you a lot of examples that I found in magazines, on the web, and places um which leave something to be desired. Uh here is a pie chart where we've completely lost the nice fact that they sum to 100°. The way they pull some of the wedges out loses that one nice feature of a pie chart. This chart is completely the wrong type of charts. Pie charts should be used for parts of a whole and that is the only reason pie chart should be used. This data is data over time. Time series do not belong in pie charts. Um the data is for today, last year and 97. Today is 51%. This green wedge certainly does not look like it's over half of the pie. Um the three percentages do not add to a 100. This is a total no. Um again, if you use a pie chart, it should only be used for parts of a whole. Okay. So much for pie charts. Let's go to unnecessary dimensions and we'll get to the figure we started with. Again, I'd like uh somebody to tell me um I'd normally ask somebody to tell me how high the A bar is. And I will tell you from experience that I'll hear 1 and a half, one and 2/3, 1 and 3/4. Um everybody will guess some number between one and a half and two. Um, here's the same data in 2D. Now, the A bar clearly looks as if it's two. At that point, I'll say, you've given me different answers. Which one do you believe? And people will say the 2D chart because it's clear. If we go back to this one, people don't know whether you read 3D bar charts from the front of the bar, the back of the bar. I think what's supposed to happen here, but it doesn't work for me, is I think we're supposed to visualize a plane tangential to the top of the bar and it should look like it goes through two. But in any case, most people misread this. Why is that? I want you to notice that the bar does not t touch the back wall of the chart. There is a gap here. Uh this was drawn in Excel. The gap is an option that you can change in Excel but most people don't know that and the default is not zero. So you get this distorted interpretation totally clear in two dimensions. Uh look at the comparison. This I did several years ago in PowerPoint. PowerPoint was written by some company that Excel bought it from them many years ago. I think it was in 2007 that Microsoft redid the software and the Excel engine was used for PowerPoint but they made it backwards compatible so that whereas PowerPoint had a gap depth of zero, Excel had a nonzero gap depth so that even these two programs from the same software ware came out differently. Well, as I was preparing for this webinar, I said, you know, I did this a long time ago. I wonder how PowerPoint is today. This was drawn this week. And look, now it matches Excel. So that it's no longer true that you read PowerPoint from the back of the bar. Now you read it exactly the same as you would for an Excel chart. Again, you can change the gap depth, but most people don't know that. Um, here is different software. And here you read it from the front of the bar. Now, I find it totally totally unacceptable that the way you read software should be dependent the way you read a graph should be dependent on the software that you use to draw it. If you have no other takeaways from today's talk, I hope you will all learn not to use pseudo 3D pie and bar charts. They're confusing. They distort the data. And there are much much better ways of presenting data. If you want to use a bar chart, a 2D one is fine, but don't put this pseudo third dimension in. Here are some graphs that are even worse. Um, in this cone, the data is the same as I've been using 2 4 6 and 8. Uh, let's take the eight bar. First of all, it doesn't reach eight, but eight 4 + 4 equals 8. The bottom four should look the same as the top four and it doesn't. Um, the bottom four looks much bigger than the top four. This one is uses cylinders instead of bars, but the same problems exist. Here's a figure that show a real figure that shows the problem that we've been discussing. This uh grid line is labeled 50. This cylinder is labeled 51.2, but it looks as if it's less than 50. So, real graphs um suffer from the problem we've just been discussing here. We uh the figures we saw before were pseudo 3D because they didn't really have three dimensions. This one does have three dimensions. We have the horizontal, vertical, and the depth. Um, I've called this variable S1, S2, S3, and S4. It suffers because we can't see some of the bars in the back. Um, every row and every column is a permutation of 2, 4, 6, and 8. And we can't see the cyan, too, because it's hidden by bigger bars. Also, the four bars with the arrow all are six, but they don't look the same. How would you show this data? Well, the way I would recommend would be to um fix one of the three variables. I'm going to fix the S1 2 3 and four and then plot the other two for that one. So, we have something like this. Here's S1, S2, S3, and S4. And we see the other two variables. Uh this way of presenting data has various names. It's um uh Tuy calls them small multiples. uh it's like trellis graphics in or lattice plots in R and it's a very effective way of showing more variables. Another problem chart are bubble charts. Here we have two charts that look very much the same but okay the top one shows world confectionary sales. It shows data for Mars Cadbury craft etc. And if you look at it carefully and study it you can see that the data is encoded in the diameter. uh this is almost 10,000 this one is about 14,000 and uh that makes the di if you look at it carefully you can see that the data is proportional to the diameters. If we go to the bottom one, this is Nobel prizes by country. Um, and you look at it carefully, you see that the data is encoded by the area, but people don't tell you how they're doing. So, some of them, some bubble charts you see it's the diameter that represents the data. Others it's the area. We visualize area. So, area makes more sense, but that's besides the point. you don't know how they're doing it. What's more, people are not very good at judging area. We judge position much much more accurately. So in CA encoding data uh by area is um not very useful. It's the same problem we saw in the pie charts. We don't judge the area of the wedges very accurately here. is a bubble chart. Um, it shows carbon per capita for different countries. I challenge anyone to figure out what these numbers represent um when represented by bubbles. Proofread graphs. This is a bubble chart where clearly the data is encoded by the diameter. The numbers are simple enough that we can see that uh this is this diameter is roughly half of that one. This one is half of that. Um the problem here besides being a bubble with diameters is that two of the labels point to the same circle. Nothing points to this one. This data is simple enough that we can all tell what it should be, but that's not always the case. Here's another graph that wasn't proof read. Um, all the labels on the vertical axis are zero. There obviously were other numbers that didn't show up. But it's as important to proofread your graphs it is to proofread your text. This is a blatantly deceptive graph. We have a change in one dimension. In this case, it's earnings or dividends, but they show visually they show it by two dimensions. And it's clear that it doesn't start at zero. If you um so that we have lengths not starting at zero and increasing in two dimensions which gives a very very inflated view of the growth. I took the bottom chart and plotted it accurately and you can see uh the dividends are growing but very gradually nothing like what you see in that um in the last figure. All right. Equally spaced tick marks for unequal intervals. This figure comes from a sociology textbook and we have the number of persons sentenced to death from 1953 to 2004 which is over 50 years. But notice they give the first 41 years less than half the space. Then they go up by one, two, three. You can't do that. You can't just arbitrarily put tick marks where you want. Let's look at the vertical axis. 0 500,000 2356. What are they doing? Well, this says that the data comes from the US Department of Justice Statistics. So, I went to the Justice Statistics website and I found this chart. I don't love it. Uh, we see 1968 here. I can't tell whether 1968 is here or anywhere to there. Um, I would rather tick marks, but I'm being petty. We have even five-year intervals here. We have even 500 prisoners here and everything is drawn to scale properly. Let's compare the two. If I were to ask you what percentage of the time there were relatively few prisoners on death row here, you would say about half the time. Here you would say less than a quarter of the time. If I were to ask you how does it grow when it does grow here? You would say pretty linear. Here you'd say there's a big hump. So the way this was drawn is totally distorted. You cannot just arbitrarily put the data where you want. Another example 10 year 10 year 2 year 6 years. You can't do that. Um, here's a bar chart. This is a range of $10,000. Um, 10,000, 25,000, 50,000. Uh, these bars look bigger, but this has a bigger range. So again, if you're drawing this, there should be the same range for each one. This has the opposite problem. Uh we have even 10-year intervals, but they're not equally spaced. In this case, I can figure out what they're probably doing. Uh, this chart is heavily annotated and I think the spacing is to fit the annotations in, but the effect is the same as the last one. It distorts the data if you don't have equally spaced intervals. Here we see a scale break and they they draw a line through the break as if this line had some meaning. That's another no. You cannot draw a line through a scale break. Here is a chart from none other than the prestigious New England Journal of Medicine. Uh what we're looking at here is men with different um body mass indices and uh the relative risk of death. And look here. This is one and a half. One and a half. Um, here's two, three, five. It's distorted. You can't do that. How should you do it? If your data is not equally spaced, show the spacing that exists. This is handled correctly. We have data after 1 2 4 8 12 weeks. So we space them at 1 2 4 8 12 weeks and um okay people make comparisons with different scales. This is the same chart as before it show be uh before you saw just the men enlarged. Here you see how it appeared in the journal which was the men and women um on the same page. We've already discussed the horizontal axis. Let's look at the vertical one. The way it is presented on the same page invites comparison. Uh the women go from 6 to 2.2. The men go from 6 to 2.8. In addition, the spacing say between 6 and8 here is much bigger than it is there. So, we're making comparisons, but they're not fair comparisons. Elements of a graph should be used consistently. Um, in this case, we have color. And if we look at the right in all three figures here, green is 2010, blue is 2012. So I have learned green is 2010, blue is 2012. I come here. Wait, now blue is 2010 and green is 2012. people will learn from an early chart and assume that the same uh parameters um continue and that's just not the case here. This is from a different sociology book. Uh in this book many many of the examples use different countries. So although I don't think uh one should use each bar a different color in most cases in something like this it does make sense if you want to associate a color with the country so people can find it more easily like for here the United States is blue you go to this chart United States is blue it makes it easy to find the United States well let's look Here, Sweden is gold. I don't see gold here. Oh, wait. Here's Sweden. It's purple. Well, why use colors if you're just going to change them and confuse the person? Would you think of writing sentences with each word a different color? Then why make each number a different color? Why not show the same respect for numbers that we show for words? The next mistake we're going to discuss is bar charts not with no zero. Uh I am not saying that all graphs must stay must start with zero. I don't believe that all graphs must start with zero. But all graphs that judge length, um, let me change that to most graphs that judge length should begin at zero. Why did I say most and not all? Well, if you're looking at ratios, say the ratio of men to women, if they're even, you'd be starting at one. There are some cases that you can think of. But basically um if we're looking at length, lengths begin at zero. And so a chart like this is a visual lie. Uh here we see temperature. This is going from 70 to 77. Um, look how much longer that is. 77 than seven. They say their forecast is smarter, but I say their graph is dumber. Great lines don't help. Here we have a graph not starting uh I mean it's really starting at 45 but they tell you zero and there's a break. The fact is that when you look at the cylinders you cannot tell from the cylinders that there's a break. So, I mean, it helps that they're telling you that it's not starting at zero, but visually it's as much of a lie. Another case, they're telling you it's not zero, but it's still a visual lie. Figures not to scale. This is one that I found in my local library. It shows the police officers in different counties in New Jersey. I want you to look at Pake and Salem. Pake had just over a thousand. Salem had just over a hundred. Does this line look 10 times that line to you? Here it is drawn to scale. Now, Pake is 10 times Salem. This figure shows the number of tourists that come from different countries coming to America. Mexico had just over 4 million uh whatever year this was. Uh Canada had just over 14 million. This line doesn't this rectangle doesn't look over three times that one to me. Again, Germany 1 million, Japan 2 million. That's not double that. If we're going to use something based on length, it should be proportional. Another one, uh, this shows summer medal Olympics. Um, Germany had two and that's $4.99. So I figure two metals represent around 500. So six metals should be about 1,500, but it's 19. They're just not to scale. This came in my mail with from a very worthwhile charity asking for money, but I did not like their figure. Um, the yellow says program expenses. This yellow green says development. And this tiny sliver says expenses. Expenses is red. The sliver is red. So you think that that's their only expense. However, let's look at it. Development 10%, management general 7%. This is certainly not 70% of this. Also, the development is almost the same color as the program, so you don't notice it. Um, so visually it's a sliver that's nonprogram, whereas if you read it carefully, it's 17%. Um, very deceptive. Another problem we regularly see in graphs is that error bars are not explained. The error bar could represent the standard deviation of the data. It could represent the standard deviation of a summary statistic which is called stand if it's the standard deviation of the mean, it's the standard error. It could be a confidence interval. And yet regularly, okay, some distributions the error bars are equivalent to 68% confidence intervals. Others the confidence interval is not based on the standard deviation. But are we really even interested in a 68% confidence interval? That's very useful in tables because it lets you create your own confidence interval. So if you tell people one standard error, they can uh multiply to do however many standard errors as they want. But the graph is a finished product and that's not particularly useful. Here we see error bars. We have no idea what they stand for. Another error bars, no idea what they are. I hope you're getting to see that this is a common problem. Color is another way to mislead with graphs. Uh color has three dimensions. Uh hue. There are a number of ways of defining them. We're going to use hue, saturation, and lightness. Uh the hue is like blue, orange, whatever. Saturation, as it's less saturated, it's grayer. As it's more saturated, it's more pure. Uh this chart is uh from a talk Judy Olsen gave and she was kind enough to give it to me. Lightness, darkness, you know what it means. Cindy Brewer, Cindy Brewer has a wonderful website, colorbower.org, or which I highly recommend for uh getting ideas of uh color schemes. Sequential schemes which is the same hue but you're varying lightness and saturation are useful for quantitative data. Uh qualitative schemes where it's different hues of the same saturation and lightness are useful for qualitative data. And diverging schemes where you have neutral in the middle and then going out sequentially with two different hues are useful for things like liquor scales where you go from say very dissatisfied to very satisfied. Here's a map using a qualitative scheme. And it looks as if there's a big divide, a big change here. I'm just enlarging this so you can see what the colors are. So you can see that we have sequential labels with colors uh with these colors. If I do it in black and white, you can see that we're going up and down this. We get lighter, darker, darker, lighter. It doesn't make much sense. Uh Kenneth Morland redid it with a sequential scheme and now the divide here completely goes away. Um, it was all a result of a poor choice of colors. We should consider people with color vision deficiencies. Uh, many people will tell you you're okay if you avoid red and green. That's a gross oversimplification. I recommend putting graphs through a color vision simulator to see how they would look uh for people with common forms of color vision deficiency. Um popular colors I often see orange and green. If I put it through a simulator, this is what you will see. They will have a lot of trouble distinguishing the orange and green. Another example, gradient backgrounds are a problem. Uh, all four squares in this are this color, but they look darker on a light background than they do on a dark background. These um here we see a dot plot. I really like dot plots, but uh what they show here are the size of the 50 US states in alphabetical order. Alphabetical is rarely best. If I reorder them by size, it's much more informative. Suppose we wanted the median size of the state. Um it would be very difficult here to find it. Whereas here you just go up and down 25 and you see the median size. This is probably the worst graph I am showing you. This one I drew but it looks exactly like what one of my clients had done. And this time I am going to ask you to put the problem what you think the worst problem of this graph is in the chat. Look at it tell me what's wrong and put your answer in the chat. Uh while you're doing that I'll point out some minor problems with it. I don't really like percent in every row. I think it clutters it. I can show you examples where numbers are clearer if they're not cluttered by percent signs. I'd rather say performance in percent and then just have the number. Let's go to the goal column again. I'd rather have percent in the heading and not have it in every column. We certainly don't need uh Okay, somebody wrote the data fields are not sequential. I'm not sure I understand what you mean. Uh oh, the date fields. Terrific. Um Okay. Uh you're smart. You're getting it. Um okay. Uh, you all said the right thing that the dates are wrong, but Adalfo has what I really wanted, which is the months are alphabetical. uh the um in other words uh people are saying uh that so um you've hit on a lot of the things uh the goal I don't see the need for two decimal places um just 90 another problem is that the first thing your my notices are these traffic lights and arrows which are the least important thing here. And somebody pointed out that we've met the goal. I spent a long time trying to figure out why red was red and green was green. And then I realized it was programmed so that you had to exceed the goal. It probably said if your goal if your performance is greater than 90, it's green. Otherwise, it's red. So, if you meet the goal, you get a red. Um, if you meet the goal, you get red, which is stupid. Um, so, um, again, too many percent signs, too many decimal places, putting the attention in the wrong place, having a stupid thing for red and green. But the really serious problem is what I said on the last chart. Alphabetical is rarely the best order. And certainly if you're talking about months of the year, you do not want it alphabetical. You want it January, February, March. And several of you have seen that. Um here is the same data in a graph. We show the goal. We show the performance. And you can see over time when you've just met here and here, we've just met the goal. Uh other months we've exceeded it. We're never below. Um my last example we have here a bunch of pie charts um which show world car production from 77 to 80. Uh first of all it goes back. We're used to reading graphs from left to right. This goes from right to left which most people will miss. It's very difficult to follow the wedges of the pie chart over time. Uh it's funny that this came in a book on better visualizations and I do not consider this to be a better visualization. I drew the data as small multiples where I've held each country fixed and then plotted the data. My time is up. So, um, I will just end by saying, um, we we've got plenty of time for questions. So, um, if you want to carry on for a minute or two, you're okay. I was just going to end up by saying that if you'd like more information, um, I've written a book. The emphasis here is more on creating good graphs as many of the examples in this talk are in the book many are not but the book is more positive whereas today we emphasize mistakes um the paperback is out of print but it's available as a PDF uh from bitly dot it's in the Um it's in the uh comments uh bit.ly MBR book. If you want to get in touch with me, here is my contact information. And I just want to end by the two figures that we started with. We started with an agenda with silly stupid fonts. We also showed a graph that I consider silly stupid. Both of them are designed to attract attention as opposed to communicating clearly and accurately. Both of them are, "Hey, ma, look what I can do with my computer." Well, in 2024, nobody is impressed that you know how to change fonts on a computer. or they might have been when computers first came out. Both of these are very stupid. But if I had to pick the lesser of the evils, I would pick the fonts because once you figure out what it says, you know what it says. Most people will never know what the graph says. I thank you. All right. Super. Uh, thank you, Naomi. That was uh incredibly informative. Um so we've um I've got some questions for you and for anyone in the audience if you've got questions for Naomi, we've got plenty of time. Please do add your questions to the chat. One thing I will say, I really like the way you had examples of a terrible plot and then showing the better version afterwards. And I know we've got quite a lot of people in the audience who are kind of looking to um promote themselves, part of their career. I think a really good exercise to do is just take a terrible plot, write down why it's terrible, and then draw a better version. You can put that in your data portfolio. That's going to be a really good way to show off that you actually know what you're talking about with respect to data visualization. And um in terms of places to find terrible plots, I'm a big fan of the data is ugly subreddit. It's full of monstrosities just like the kind of thing Naomi has been showing. So if you want some examples to play around with just to uh see how you can improve them, that's a good place to go. Um all right. So uh Naomi, uh my first question for you is we talked a lot about the some of the examples of specific problems with plots. Um can you talk a bit about what the impact of this might be at work? So what are the consequences of drawing terrible plots? People won't understand what you're saying. People will get the wrong impression. People will make poor decisions based on inaccurate data. I mean inaccurate information. Uh people lose credibility. The very first talk I gave 25 years ago, somebody contacted me afterwards and said based on this one-hour presentation she heard, she picked up a serious problem in one of the graphs they were about to send out. and um she said just by hearing the one-hour lecture she saved the company from a great deal of embarrassment. So um it's loss of credibility, it's bad decisions and inaccurate understanding. Yeah. Yeah. So, this has real serious consequences for both in internal blocks and then if you're putting something out into the public where your customers are going to see it, that can cause even bigger problems, it seems like. Um, okay. And I'm wondering, so you focused a bit on problems with Excel and PowerPoint, which are just notoriously awful for um for generating blocks. I know sort of a lot of database software has moved on um in the last couple of decades. Do you have any recommendations for better tools for drawing plots? I had a long career at Bell Laboratories where the S language was originated. And I knew S. So when R came out, I used R. And I really don't want to recommend I mean R is certainly well known for drawing accurate graphs for um considering perception in graphs and that but I'm not familiar enough with all software to be able to say use this and don't use that. I can recommend R, but I'm not criticizing other software. Okay. Yeah. Uh certainly um if you're into R then we're big fans of the ggplot 2 package u here at data camp. So that's I I can certainly recommend that as a good choice. And for Python it's a bit more mixed. There are a lot of different packages. Um I think plotly express is is a is a pretty good one. And then Seabour is also a solid choice. So, you've got a few options and they're going to be better than um uh they're going to be better than Excel for sure. Let me say something in Excel's defense. Um I learned very early on not to say you can't do that in Excel because Excel can do all it's amazing what can be done in Excel if you know how to use it and don't limit yourself to its defaults. There's work. There are several books um I'd have to turn around uh that where they even have reproduced Napoleon's march um in Excel. It Excel can be extremely powerful. Uh when my book came out, I showed dot plots and trellis plots and things that uh were not well known outside of the at the time it was S and then our community. And a number of Excel experts then did websites on how to do these plots in Excel. And um I highly recommend utilities by John Peltier which lets you do all of um uh dot plots. So uh what he calls panel plots in Excel. Um uh the book I was thinking of um on doing amazing things in Excel is by George C oes. I don't I'm not sure I pronounce his name correctly, but um he wrote a book which shows what Excel is capable of. Um yeah, maybe that's fair. So you can do almost anything you want in Excel. is just don't accept the defaults or assume that they're the best thing um that you that you're going to want to use. Um okay. So um one of the things you talked about was the use of color and how it can be used to distract people or it can um cause problems with color blindness for example. Now in a corporate setting there's often some kind of official corporate color scheme. I was wondering where whether you have any advice um on how to draw plots that are both going to be effective and also honor some kind of corporate colors um because it's often a a point of tension between data people and marketing people. So um do you have any advice here? Make sure you run your graphs through a color vision simulator to make sure that everybody can use it. And since I think that one should limit the use of color, you as you saw my chart, would you do each word a different color? We don't want to overuse color. So whatever your corporate color scheme is, you should be able to um do effective graphs with them. Um, okay. Yeah. So, hopefully your corporate color scheme isn't so ugly that um every graph's going to be terrible, but you might need to do might need to go and have a work with the design team or marketing team and uh see if you can get something that's going to work for database. Um, all right. So, there's a question from the audience now. So, Adulo asks, "How does one balance the evolution of data visualization in the sense of reproducibility?" So, if I can't code the perfect graph, but I can draw it, which method should I use? Um, oh, so I guess yeah, if if you're struggling to make a good plot in software, but you can draw what you want, do you have any advice there on what to do? I am much better at what I want the finished graph to look like than how to get your computer to do it. and I work with a number of people who are much better coding than I am. Uh so that um I I recommend something that's coded so that you have reproducibility as opposed to point andclick where you don't have the reproducibility. But um again uh I'm going to recommend R just because that's what I'm aware of. Uh I'm sure there are others as well. Oh um Richie, you probably can answer that question better than I can. Yeah. So I think um the the point Adovo is making is that um if you have if you've written code to draw your plot then it's easier when you come back to it six months later because you you can just click run again and it's going to reproduce what you've done. Whereas if you've had to point and click to create the plot then you don't have that you you you've got to do the pointing and clicking again to create something like it. So um if you want to be able to um create a plot using code but you're not that strong at coding there is actually a great solution this last year in that AI can now help you write code. So um for example, you can use chat GPT to help write code or if you're using data camp products, then we have data camp workspace. That's our productivity platform and that's got AI assistance built in. So you can just say what you want the plot to look like and it's going to write the Python code or the R code for you in order to draw the plot. So yeah, AI is your friend here. While we speak of AI, uh the New York data visualization meetup is having a um session on how well AI reads plots and um it's going oh I think it's February. I have to get my calendar but it will be virtual. So, anybody can attend. It'll be six o'clock. Oh, goodness. Um, well, we can perhaps send out a link to that event. Uh, for anyone who's registered for this webinar, we'll when we send out the recording, we'll send out a link to this event. So, if you want uh February 21st at 6 o'clock. Okay. Super. So, if anyone wants to attend that, then uh yeah, uh uh we we'll get you the link for that. Um, all right. One more question from the audience. So, Preston asks, "Oh, have you encountered issues or mistakes with AI generated graphs?" Um, uh, Naomi, do you want to take this or I have not I have not tried to generate graphs with AI, so I'm not qualified to answer. I will say that if you attend this um virtual meetup on the 21st you will see a lot of mistakes with AI reading graphs but um I cannot I have not tried to draw graphs with AI. Okay. Yeah. So there there are two options here. So one is you can uh get AI to generate code in order to draw the plot. So you can get it to write Python code or our code and there once it's drawn the plot you should be able to look at it and see is this nonsense or not. The other option is you try and get it to draw an image because you know you got all these image generators like Darly and Midjourney and stable diffusion. These are almost certainly going to generate nonsense plots. So I wouldn't recommend that approach. The first approach is what you want where you want to get it to generate code. And then once you look at the plot, just look very carefully and see does this actually look like what you were expecting or not. Um all right. So we're coming up to time now. So before you all dash off, since Naobi was hyping future webinars, I'm going to do the same. Tomorrow we've got a session on uh data transformation in the pharmaceutical industry. So we've got three very senior people from like chief data officers from the pharmaceutical industry. Uh so it's going to be a very big session. It's going to be um some interesting discussion. They're all good speakers. Uh next week we've got a codealong on uh getting started with data pipelines. Also got a session on Wednesday on how to get a job in data. And then on Thursday, we've got a session on uh building a data literate workforce. So, lots of uh great webinars coming up. Go to datacamp.com/weinars to register for those. Uh I look forward to seeing you there. Naomi, just once again, thank you. That was a really great session. My eyes are kind of bleeding from all the terrible plots you showed, but uh that was a that was a lot of fun. Thank you. Um all right, super. And uh Ree, thank you for moderating. Thank you to everyone who asked a question. Thank you to everyone who showed up today. And I hope to see you all again in future sessions.

Original Description

Data visualization is one of the most important data skills. By improving your abilities to understand and draw plots, you dramatically increase your ability to communicate with data. In this session, you'll learn key principles of data visualization, from understanding which plot to draw in common situations, to design techniques to improve your audience's comprehension. The instructor is Naomi B Robbins, a legend in the world of data visualization. Naomi has trained thousands of people in data visualization skills (including your host, Richie Cotton). Key Takeaways: - The most common graph forms are not necessarily the most effective. - Consider readers with color vision deficiencies. - Dot plots are often a useful alternative. - Small multiples are another useful alternative.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from DataCamp · DataCamp · 0 of 60

← Previous Next →
1 SQL Server Tutorial: Date manipulation
SQL Server Tutorial: Date manipulation
DataCamp
2 R Tutorial: Intermediate Interactive Data Visualization with plotly in R
R Tutorial: Intermediate Interactive Data Visualization with plotly in R
DataCamp
3 R Tutorial: Adding aesthetics to represent a variable
R Tutorial: Adding aesthetics to represent a variable
DataCamp
4 R Tutorial: Moving Beyond Simple Interactivity
R Tutorial: Moving Beyond Simple Interactivity
DataCamp
5 Python Tutorial: Why use ML for marketing? Strategies and use cases
Python Tutorial: Why use ML for marketing? Strategies and use cases
DataCamp
6 Python Tutorial: Preparation for modeling
Python Tutorial: Preparation for modeling
DataCamp
7 Python Tutorial: Machine Learning modeling steps
Python Tutorial: Machine Learning modeling steps
DataCamp
8 R Tutorial: The prior model
R Tutorial: The prior model
DataCamp
9 R Tutorial: Data & the likelihood
R Tutorial: Data & the likelihood
DataCamp
10 R Tutorial: The posterior model
R Tutorial: The posterior model
DataCamp
11 R Tutorial: An Introduction to plotly
R Tutorial: An Introduction to plotly
DataCamp
12 R Tutorial: Plotting a single variable
R Tutorial: Plotting a single variable
DataCamp
13 R Tutorial: Bivariate graphics
R Tutorial: Bivariate graphics
DataCamp
14 Python Tutorial: Customer Segmentation in Python
Python Tutorial: Customer Segmentation in Python
DataCamp
15 Python Tutorial: Time cohorts
Python Tutorial: Time cohorts
DataCamp
16 Python Tutorial: Calculate cohort metrics
Python Tutorial: Calculate cohort metrics
DataCamp
17 Python Tutorial: Cohort analysis visualization
Python Tutorial: Cohort analysis visualization
DataCamp
18 R Tutorial: Building Dashboards with flexdashboard
R Tutorial: Building Dashboards with flexdashboard
DataCamp
19 R Tutorial: Anatomy of a flexdashboard
R Tutorial: Anatomy of a flexdashboard
DataCamp
20 R Tutorial: Layout basics
R Tutorial: Layout basics
DataCamp
21 R Tutorial: Advanced layouts
R Tutorial: Advanced layouts
DataCamp
22 Python Tutorial: Time Series Analysis in Python
Python Tutorial: Time Series Analysis in Python
DataCamp
23 Python Tutorial: Correlation of Two Time Series
Python Tutorial: Correlation of Two Time Series
DataCamp
24 Python Tutorial: Simple Linear Regressions
Python Tutorial: Simple Linear Regressions
DataCamp
25 Python Tutorial: Autocorrelation
Python Tutorial: Autocorrelation
DataCamp
26 R Tutorial: The gapminder dataset
R Tutorial: The gapminder dataset
DataCamp
27 R Tutorial: The filter verb
R Tutorial: The filter verb
DataCamp
28 R Tutorial: The arrange verb
R Tutorial: The arrange verb
DataCamp
29 R Tutorial: The mutate verb
R Tutorial: The mutate verb
DataCamp
30 R Tutorial: What is cluster analysis?
R Tutorial: What is cluster analysis?
DataCamp
31 R Tutorial: Distance between two observations
R Tutorial: Distance between two observations
DataCamp
32 R Tutorial: The importance of scale
R Tutorial: The importance of scale
DataCamp
33 R Tutorial: Measuring distance for categorical data
R Tutorial: Measuring distance for categorical data
DataCamp
34 Python Tutorial: Plotting multiple graphs
Python Tutorial: Plotting multiple graphs
DataCamp
35 Python Tutorial: Customizing axes
Python Tutorial: Customizing axes
DataCamp
36 Python Tutorial: Legends, annotations, & styles
Python Tutorial: Legends, annotations, & styles
DataCamp
37 Python Tutorial: Introduction to iterators
Python Tutorial: Introduction to iterators
DataCamp
38 Python Tutorial: Playing with iterators
Python Tutorial: Playing with iterators
DataCamp
39 Python Tutorial: Using iterators to load large files into memory
Python Tutorial: Using iterators to load large files into memory
DataCamp
40 SQL Tutorial: Introduction to Relational Databases in SQL
SQL Tutorial: Introduction to Relational Databases in SQL
DataCamp
41 SQL Tutorial: Tables: At the core of every database
SQL Tutorial: Tables: At the core of every database
DataCamp
42 SQL Tutorial: Update your database as the structure changes
SQL Tutorial: Update your database as the structure changes
DataCamp
43 Python Tutorial: Classification-Tree Learning
Python Tutorial: Classification-Tree Learning
DataCamp
44 Python Tutorial: Decision-Tree for Classification
Python Tutorial: Decision-Tree for Classification
DataCamp
45 Python Tutorial: Decision-Tree for Regression
Python Tutorial: Decision-Tree for Regression
DataCamp
46 Python Tutorial: Census Subject Tables
Python Tutorial: Census Subject Tables
DataCamp
47 Python Tutorial: Census Geography
Python Tutorial: Census Geography
DataCamp
48 Python Tutorial: Using the Census API
Python Tutorial: Using the Census API
DataCamp
49 R Tutorial: A/B Testing in R
R Tutorial: A/B Testing in R
DataCamp
50 R Tutorial: Baseline Conversion Rates
R Tutorial: Baseline Conversion Rates
DataCamp
51 R Tutorial: Designing an Experiment - Power Analysis
R Tutorial: Designing an Experiment - Power Analysis
DataCamp
52 R Tutorial: Introduction to qualitative data
R Tutorial: Introduction to qualitative data
DataCamp
53 R Tutorial: Understanding your qualitative variables
R Tutorial: Understanding your qualitative variables
DataCamp
54 R Tutorial: Making Better Plots
R Tutorial: Making Better Plots
DataCamp
55 SQL Tutorial: OLTP and OLAP
SQL Tutorial: OLTP and OLAP
DataCamp
56 SQL Tutorial: Storing data
SQL Tutorial: Storing data
DataCamp
57 SQL Tutorial: Database design
SQL Tutorial: Database design
DataCamp
58 Python Tutorial: Introduction to spaCy
Python Tutorial: Introduction to spaCy
DataCamp
59 Python Tutorial: Statistical Models
Python Tutorial: Statistical Models
DataCamp
60 Python Tutorial: Rule-based Matching
Python Tutorial: Rule-based Matching
DataCamp

This video teaches key principles of data visualization to help create better graphs and avoid common mistakes, enabling effective data communication. By applying these principles, viewers can improve their data skills and convey insights more efficiently. The session covers graph forms, design techniques, and considerations for color vision deficiencies.

Key Takeaways
  1. Understand common graph forms and their effectiveness
  2. Consider readers with color vision deficiencies
  3. Use dot plots as a useful alternative
  4. Apply small multiples for improved comprehension
  5. Design graphs with the audience in mind
💡 The most common graph forms are not necessarily the most effective, and considering the audience's needs is crucial for effective data communication.

Related Reads

Up next
How to Prompt Your LLM Directly from SQL
Ian Wootten
Watch →