Conjugate Priors : Data Science Basics
Skills:
ML Maths Basics90%
Key Takeaways
The video covers conjugate priors in Bayesian statistics, a fundamental concept in machine learning, and their importance in data science, referencing the Beta Distribution for further understanding.
Full Transcript
Hey everyone, welcome back. Today we're going to be talking about a concept in basian statistics called conjugate prior. Now, I'll be honest, this concept was really hard for me to grasp at first, but what made it way easier is just starting with a story and then using that story to build up this concept. And so, that's exactly what we're going to be doing today. The story is that let's say one lazy Sunday afternoon, you find yourself standing on some land here, and then there's an island on the other side. And between you and the island there's some body of water. And so naturally what you start doing is just start skipping rocks across this body of water and seeing if you can make those rocks land on the island. Sometimes they do land on the island and other times they don't. And your mind starts wondering what would it take to model the probability P that any one of my rocks is going to reach that island. What we have is some data of throws that we've done through the afternoon. So x1, x2 all the way to xn. Our n throws that we have attempted. And the value of each of these is just a binary variable. So it's going to be zero if that throw sank, did not reach the island. And it's going to be one if it did reach the island. We've also been studying basian statistics in our school. So we wonder how we can model this probability P that any one of our throws would reach the island using basian statistics. Now the easiest concept to get started with is actually the likelihood which remember is the probability of observing the data we observe in the real world. So this capital X here is just shorthand for all of these throws that we've done through the afternoon. What is the probability of observing that data given some setting of the parameter of interest which again is P. Now framed in this way which is we have n trials n throws let's consider them all as independent and each of them is a binary variable and we want to know what's the probability of observing k successes or k times where we did reach the island and then n minus k failures framed in that way that is exactly the definition of the binomial distribution and so we can already say the likelihood the probability of observing the data given some setting of the parameter of interest is going to be binomial is going to be a binomial distribution. So we know the binomial distribution probability of X given P is going to be N choose K where K again is the number of these throws that were successful. This probability P that we have assumed here to the power of K 1 minus P to the N minus K. So we already have one step out of the way. But of course in basian statistics the other quantity we care a lot about is going to be the prior. So what is a natural choice of prior to set for P? which remember again is a probability which has to be bounded between zero and one and that's the thing we are trying to model. So if you think back to your courses or if you have watched previous videos on this channel and even if neither of those is true a very natural choice of the prior here is going to be the beta distribution. The beta distribution if you're unfamiliar with it first I'll link a video that I've made uh to the beta distribution explaining it in the description below. But it is a very common and a very natural choice we make when the thing we are trying to model is a probability itself because the beta distribution is defined between 0ero and one exactly as probabilities are. So the probability density function of the beta distribution is p to the alpha minus one. I'll come back to what this alpha is in a second. 1 - p the^ of beta minus one all divided by some normalizing constant which is the beta function. Don't need to worry too much about that just yet of alpha and beta. Now what are these alpha and beta in this context? They are called hyperparameters and they are values that we set beforehand. Values that we are going to set based on our prior belief of which of these probabilities are likely and which of them are unlikely before observing any data. The usual interpretation of these alphas and betas is that alpha minus one is going to be the number of pseudo successes that we are using to form up this prior distribution and beta minus 1 is going to be the number of pseudo failures that we are using to build up this prior uh distribution on p. So to make that concrete let's say alpha is equal to 3 and beta is equal to four. Then what that means by setting alpha equals 3 and beta equals 4 is that we are kind of seeding or initializing our probability P by having 3 - 1 or two pseudo successes and then we are initializing that with four minus one or three pseudo failures. Now this interpretation is crucial because it has two big implications. One is that it's exactly the ratio between this alpha and beta that is going to let us encode certain values of this p being probable and other values being not probable. And the other is that the magnitude of these alphas and betas are going to encode how strongly we believe in that prior and how much evidence x it would take to move away from that prior. All this will be much more clear by the time we get to the final sheet, but that's just something that I want you to keep in the back of your mind as we talk through the next steps. So just to recap where we're at so far, we're trying to model this probability P. We are going to have a likelihood which is the probability of observing this data given some setting of P. And we have some prior which is in the absence of any of that data before observing any data, what is our prior belief on different probabilities P. And as we do in basian statistics, the next thing we care about, and you probably saw this coming a mile away, is what is the posterior distribution, which is a distribution of the probability of this parameter of interest P given the data we do observe. And that is exactly what we're after. So let's go ahead and try to work that out mathematically. So we know by bay theorem that the probability of the probability we're trying to model given the evidence that we see is going to be proportional to the prior. So this guy right here is going to be our prior times the likelihood. So probability of the evidence given some setting of the parameter of interest. So prior times likelihood that's proportional to the posterior. And now we can just stick in the mathematical forms that we have. So we know the beta distribution which is our prior has this probability density function and the binomial distribution which is our likelihood has this probability mass function. So just to be clear, this is our likelihood and this guy right here is going to be our prior. And now there's a couple things to notice here. Even though this is starting to look a little bit complicated, there's a really easy way to clean it up, which is that this n choose k thing right here, does that have any bearing on p? Is that related to p at all? No, that's just based on the evidence that we have. And this beta function of alpha and beta, is that related to the p? No, that's just based on the alpha and beta that we chose for our prior. So those can be pulled out as constants. And so we can say this whole thing here is proportional to just the pieces that are in here that are actually related to P which is the thing that we are trying to model. So it's just going to be P alpha minus one and then we add this K. So if it's not clear what we did here is P alpha minus one * P to the^ of K. Exponents add when they have the same base. So that's exactly what happened here. And we do the same exact thing for the 1 minus P. So we get this beta minus one from the prior and then add that to n minus k from the likelihood. And so that's the form we have right here. And I put a proportional to symbol here because we pulled out these constants which again have nothing to do with p. So we can say our posterior is proportional to this quantity here. Now we want the true form of the posterior here. We know it's proportional to this, which means there's some normalizing constant C, which if we multiply these terms by is going to allow this whole thing to be a proper probability distribution. It's going to allow it to actually integrate to one. And we can find that by just carrying out that integral. It's going to look a little scary, but I promise it's going to work out very cleanly. So we're going to integrate over all possible values of P which technically is going to be between 0 and 1 because P is a probability of our posterior probability of this P we're trying to model given X DP and that needs to equal one for this to be a actual probability density function. And so we know the form of this. We just worked it out here. So this piece right here is exactly this component right here. and this C this constant we pulled out and we know this needs to equal one. So if we just call this integral here I for shortand I for integral then we can easily see we have C * I equals 1 or in other words this constant is equal to 1 / I. And the way this all gets much simpler is if you look at this form of the integral right here you're going to find that is exactly the form of the beta function whose arguments are alpha plus k and beta plus n minus k. Now I'm fully aware that this part of the video might be a little bit confusing. It kind of feels like I just handwaved complexity into the form of a function. But basically what's happening is that the beta distribution was already defined to be a probability distribution by having this normalizing constant down here. So that this thing is going to integrate to one. And so all we're doing is just noticing we're just noticing that this integral we have formed here has exactly the form of that normalizing constant except that the arguments are going to be different. They're going to be alpha + k and beta plus n minus k. But let's zoom out for a second and realize what this is all telling us. We have found that the posterior distribution, so probability of this P given X, so probability of this P given X is going to be equal to, so previously said it was proportional to this piece right here. So it's going to be equal to that piece right there. And I don't have room here, so I'm just going to call that piece T for terms. So forgive my short hand there. And we found that the normalizing constant is going to be one over this integral where this integral we've noticed is exactly the form of the beta function whose arguments are alpha + k and beta + n minus k. So hopefully you can see that here. And if we look at this, if you look at this, plug in these terms here, the posterior is also a beta distribution. And this is the big aha moment of the video. But it's not clear why it's an aha moment yet. But right now, all I'm asking you to do is notice is notice. And we have proven mathematically that if our likelihood function is this binomial and we have chosen this beta distribution for our prior then the math works out such that our posterior distribution which is the thing we are using to model this probability that we're going to skip a rock that lands on the island is also a beta distribution with different parameters. It's a beta distribution with parameters alpha plus k and beta plus n minus k. And now what I need to convince you of is we we believe this is true. We have shown it mathematically in fact. Why is it significant? That is what I need to convince you of and that's what we're going to do on the next sheet here. So to recap, we chose a prior distribution P on this probability we're trying to model as beta alpha beta. Again, alpha and beta are just chosen based on our prior beliefs. The likelihood function, our data given some setting of the parameter p is going to be binomial with n and p. n being the number of trials and p being this assumed uh probability a rock's going to land on the other end. And our posterior we found mathematically this p given our data is also going to be beta with parameters alpha plus k and beta plus n minus k. And now here is when we come back to the interpretation of the beta distribution. Remember before we talked about a beta distribution with parameters alpha and beta as having alpha minus one pseudo successes. That's why I've put successes in quotes and beta minus one pseudo failures. Again failures in quotes. And remember that we said the ratio between the alpha and the beta. let us inform what values of P we thought were likely and unlikely before seeing any data and the magnitudes of alpha and beta. So if alpha and beta very large magnitude, it would take a lot of evidence for us to move away from that. If they were small likelihood, we were going to let even just a little bit of evidence X allow us to move away from that. And that is exactly the story that is being told when we get to the posterior because remember if that's the interpretation of this beta distribution, this also being a beta distribution must have the same interpretation. It's just that now we have alpha minus one plus k successes and we have beta minus one plus n minus k failures. And where have we seen this quantity of k successes and n minus k failures? It's exactly in our evidence x. Remember when we skipped all those rocks and we said we had k successes among those n trials. Therefore we had n minus k failures. So literally literally by picking a prior distribution that is beta distributed with alpha and beta we are allowing ourselves by the time we get to the posterior distribution to interpret that in the same way. We are telling a story that before we had alpha minus one and beta minus one successes and failures respectively. We saw some evidence X which contained k actual successes and n minus k actual failures. And so what we're going to do is model our posterior again as a beta distribution. It's just that we're going to add in those k successes in the first term and we're going to add in those n minus k failures in the second term. We're able to form this very intuitive, nice, neat, very interpretable story by picking our prior as beta if our likelihood is binomial. And that's where we're going to get into the actual terms now that we understand the significance of the story here which is that this prior and this posterior are conjugate distributions are conjugate distributions. So this prior and this posterior are conjugate distributions with respect to our likelihood with respect to our likelihood. And that's exactly the language you use when your prior and the posterior have the same distributional form which they both do. They don't have the same parameters but they are both beta distributions. So therefore our prior and posterior are conjugate distributions with respect to our likelihood. Another term people use in the title of this uh video itself was that this prior is called a conjugate prior for this likelihood function. So it's all three of these working together. It's given you picked a certain likelihood. If you pick certain priors, those can be conjugate prior if the posterior also has that exact same distributional form. And we care about this because it's not like basian statistics breaks. If this is not true, this is not true for every single prior likelihood and posterior you're going to run into out in the world. This is nice for two reasons. One which is a little less important in this day and age and the other which I think is very important in this day and age. So the one that's a little bit less important is that uh if our prior and posterior are both the same distribution, they have these closed forms and that makes computational efficiency a little bit easier. That is true, but we have a lot of computing power these days. So it's not the worst thing in the world if our posterior does not have a closed form. We can just use numerical method. So that's not a make or break. What I think is the actual important piece of this whole concept is that it leans into the inherent interpretability of basian statistics because if the prior and posterior have the same distributional form with different parameters, then we're able to compare apples to apples. Imagine the posterior had a different distributional form. Well, that's fine. We can work that out and it doesn't make it wrong. But it means that when we go talk about the story of updating our prior using the likelihood to get to the posterior, it's a lot more difficult. and in most cases impossible to talk about what is actually happening and how do these observations and the likelihood update the posterior and what is the actual physical interpretation of that. Here we're able to exactly tell that story. We're saying we had alpha minus 1 beta 1 pseudo successes and pseudo failures with the prior. We saw some real successes and real failures and we're going to use those in order to update our posterior. So now we have a total of alpha - 1 + k successes and beta - 1 plus n minus k failures. And the story can go even further there. Notice that if k and n minus k are large, we skipped a lot of rocks. Then the dominant terms are going to be the k not the alpha and the n minus k, not the beta. So our beta distribution is going to be fully dominated by the evidence and less so by the prior. And on the other hand, if our alpha and beta were very large, we had a very confident prior going into this and we observe just a little bit of data, then the dominant terms are going to be alpha and they're going to be beta. So it's going to be those pseudo successes and pseudo failures that are going to dominate our posterior distribution. That's a story we don't get to tell if we don't have this conjugate distribution phenomenon going on. So, if I was going to say that in a much less rambly way, what the concept of conjugate distributions allows us to do is that it lets us better link the prior and the posterior together through the link of the likelihood function itself, as we've done here. And I've just shown one rather simple example here, but there's a whole table of prior and posterior conjugate distributions given a certain likelihood. Many of them discrete, many of them continuous. And I encourage you to check those all out. And if you want me to go through any other ones in future videos, do let me know. But hopefully in this video, you learned about the concept of conjugate priors and why they're so powerful. It was confusing for me at first because I was taught it as, hey, it's when your prior and your posterior have the same distributional form. But my next question was, oh, okay, like, who cares about that? What does that matter? But I think when you build it up through a story and are able to say that we're explicitly able to have the interpretation of our prior be the same as the interpretation of our posterior just with the benefit of extra data. In my mind, that became a whole lot more clear and I'm hoping it did for you, too. So, any questions or comments are always welcome in the section below. Please like and subscribe for more videos just like this.
Original Description
All about the importance of conjugate priors in Bayesian statistics!
Beta Distribution Video : https://www.youtube.com/watch?v=1k8lF3BriXM
0:00 Intro
5:13 The Math
9:48 The Big Idea
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from ritvikmath · ritvikmath · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Math Team Update
ritvikmath
Single Variable Calculus Volume of a Sphere - Proof 1
ritvikmath
Single Variable Calculus Volume of a Sphere - Proof 2
ritvikmath
Multivariable Calculus Volume of a Sphere Proof - Triple Integrals
ritvikmath
Multivariable Calculus Volume of a Sphere Proof - Double Integrals
ritvikmath
The Euclidian Algorithm
ritvikmath
Proving the Chain Rule
ritvikmath
Proving the Fundamental Theorem of Calculus Part 1
ritvikmath
Proving the Fundamental Theorem of Calculus Part 2
ritvikmath
Math Puzzle - Poison Perplexity
ritvikmath
Math Puzzle - Poison Perplexity - Solution
ritvikmath
Expected Value and Variance of Continuous Random Variables (Calculus)
ritvikmath
Expected Value and Variance of Discrete Random Variables (No Calculus)
ritvikmath
Array Method
ritvikmath
Complex Power Series and their Derivatives
ritvikmath
Distributions - Intro
ritvikmath
The Poisson Distribution
ritvikmath
The Bernoulli Distribution
ritvikmath
The Binomial Distribution
ritvikmath
The Continuous Uniform Distribution
ritvikmath
The Geometric Distribution
ritvikmath
The Triangular Distribution
ritvikmath
The Exponential Distribution
ritvikmath
The Borel Distribution + Notes on Poisson Distribution
ritvikmath
The Gamma Distribution
ritvikmath
The Normal Distribution
ritvikmath
The Laplace Distribution
ritvikmath
The Chi - Squared Distribution
ritvikmath
Overfitting
ritvikmath
Vector Norms
ritvikmath
Truths Behind the Titanic : K-Nearest Neighbor
ritvikmath
The Mathematics of Breakups
ritvikmath
Sillyfish
ritvikmath
Finding Optimal Paths - Dynamic Programming
ritvikmath
HowToDataScience : Scraping Twitter Data
ritvikmath
Decision Trees
ritvikmath
Perceptron
ritvikmath
Naive Bayes
ritvikmath
K-Nearest Neighbor
ritvikmath
Evaluating Machine Learning Models
ritvikmath
Decision Tree Pruning
ritvikmath
K-Means Clustering
ritvikmath
Gaussian Mixture Model
ritvikmath
Data Science - Fuzzy Record Matching
ritvikmath
Time Series Talk : Autocorrelation and Partial Autocorrelation
ritvikmath
Time Series Talk : Autoregressive Model
ritvikmath
Time Series Talk : Moving Average Model
ritvikmath
Time Series Talk : ARMA Model
ritvikmath
Time Series Talk : ARCH Model
ritvikmath
Time Series Talk : White Noise
ritvikmath
Time Series Talk : Stationarity
ritvikmath
Time Series Talk : ARIMA Model
ritvikmath
Time Series Talk : Lag Operator
ritvikmath
Time Series Talk : What is Seasonality ?
ritvikmath
Time Series Talk : Seasonal ARIMA Model
ritvikmath
So ... What Actually is a Matrix ? : Data Science Basics
ritvikmath
Derivative of a Matrix : Data Science Basics
ritvikmath
Basics of PCA (Principal Component Analysis) : Data Science Concepts
ritvikmath
Eigenvalues & Eigenvectors : Data Science Basics
ritvikmath
The Covariance Matrix : Data Science Basics
ritvikmath
More on: ML Maths Basics
View skill →Related Reads
Chapters (3)
Intro
5:13
The Math
9:48
The Big Idea
🎓
Tutor Explanation
DeepCamp AI