Truths Behind the Titanic : K-Nearest Neighbor
Key Takeaways
The video explains the K-Nearest Neighbor algorithm and its application to the Titanic dataset, using techniques such as normalization, standardization, and Gaussian norm to improve accuracy. The algorithm's performance is sensitive to the choice of K, and tweaking parameters can improve accuracy.
Full Transcript
in this video we'll be continue our exploration of the Titanic data set and the exciting thing in this video is that we're going to be going deeper into the machine learning concepts beyond the regressions and the basic data analysis we did in the last two videos so we're gonna be using in this video a one of the more simple machine learning techniques called K nearest neighbor and how that kind of works at a visual level is that we're gonna have some similarity measure between each two passengers so let's say I'm passenger I and your passenger J so we're gonna have some number called s subscript IJ which is between 0 & 1 which tells us how similar we are as passengers so if it's close to 0 that means we're not very similar if it's very close to 1 that means we're very very similar passengers so let's say for a second and we'll be will be constructing the function which tells us a similarity but let's say for a second we have that function already so we have the numerical similarities between each two passengers so given that let's say now we want to find out if passenger I survived the Titanic disaster or not so we can look at this kind of graphically let's say that passenger I is this represented by this dot right here so now out of all the other end passengers on Titanic or out of the other end passengers or I guess n minus 1 passengers who are looking at in the data set they're gonna have we're gonna look at the ones that are closest to this passenger so if this is passenger I let's say graphically the other passengers lie like this so let's say there's another passenger this is another passenger this is another passenger and so on and the distances represent what's the similarity between them so the closest ones actually have a similarity close to one and the ones that are all the way out here have a similarity very close to zero they're not very similar to this passenger now the K in K nearest neighbor is user-defined we can define as whatever we want and that just tells us how many neighbors we're gonna be looking at let's say for the sake of this simple example we're just looking at three neighbors so we're gonna look at the three nearest neighbors to this person I write here so it looks like that's this person it looks like it's this person and maybe it is this person so these are the three nearest neighbors to person I now we look at is what's the status of the three people let's say this person survived the disaster though say this person did not survive so put and and this person survived since two out of the three survived we're gonna estimate we're gonna guess that person I also survived because two out of the three neighbors survived it's likely that this person also survived and the rationale behind that is since this person is similar to these other three people these are the most similar people to this person we're gonna say that his his or her fate was also similar to the fate of these other three people so that's kind of the rationale behind it and logically it does make sense but again one of the caveats is that K is user-defined which means that if we looked at a different sub if we took a different K we might get different results for example let's say instead of K equals three we looked at K equals five so now we're looking at the five nearest neighbors so maybe were looking at this person as well maybe we're looking at this person as well and let's say for these two people these people did not survive so now we've changed the outcome because now we have three non survivors versus two survivors so we'll say this person did not survive so the best thing one of the better things can do is if we have a lot of people then we can pick K kind of big so that we get a better sample than just three people and also some kind of strategy you can use is let K be odd so there's never a tie because we're just looking at bivariate cases it can be if there's a tie it's kind of unclear what to do so we'll be keeping K as odd in all of these cases so I'll let you know the exact value of K we use a little bit later in the video but right now let's see how we construct this similarity measures so this w is called a similarity matrix it's an N by n matrix so it's n by n is the height and n is the width and n can be everyone in the sample or it can be a sub sample of the people and the reason we wouldn't choose everybody is that this K nearest neighbor it grows quadratically with the number of people you choose for example if you double the number of people that you're using the sample then the number of computations is going to quadruple because you're gonna have doubling the height doubling the width so it's going to be times four so that's one drawback of this right away is that it's order N squared so it takes n square time to run versus some of the other methods we might have used before so we're gonna be using just an arbitrary n maybe it's the whole sample maybe it's some subset whatever now what do you do so this w is the similar tricks this tells is similar between the first person and himself this tells you the similarity between the first person and the second person is the similarly bit in the second person the first person and so on so s IJ so s IJ as we said right here represents a numerical similarity between 0 and 1 between the I and the J person let's observe some things about this matrix before we determine exactly how to calculate these similarity measures right here what is s11 what should it be kind of logically what's a similarity between a person and him or herself it should be 1 right because you can't get any more similar than having a copy of yourself so that means this should be 1 also for the same reason s 2 2 should be 1 for all these s and n should be 1 so this diagonal should be full ones now what s 1 2 well it could be anything we don't be a second person could be any similarity with the first person but whatever it is what should s 2 1 be so what's the similarity of the first person and the second person versus the similarity the second person the first person there should be the same right because it doesn't matter what order you take them in they should have the same similarity so in addition to having ones on the diagonal this W matrix the similarity matrix is also symmetric because this is the same as this this is the same as this and just like that every component s IJ basically equals s J that's what symmetric matrices me so that's kind of all there is to know about the similarity matrix it's a matrix of the similarity measures between each group of people it's symmetric and it has ones on the diagonal so now let's get to the meat of what we're talking about how do we determine the similarity between two people well let's start with this we're going to go about it iteratively so we can kind of build on previous knowledge we're gonna go ahead and look at three characteristics of people to kind of keep things a little bit simple for right now and kind of taking inspiration from the regression cases we were looking at before we're gonna be looking at someone's sex their fare and their age so that's that's kind of like a triple so we can represent as a vector so let's say for a person I let's say VI is a vector which has their sex their age and their fare that they pay to be on the Titanic now similarly for V V J is the same set of values for a person J it's there sex their age and their fair now what's one of the most crude things we can do to figure out how similar these people are we can subtract these two vectors right so let's let's just do that let's say we do VI minus VJ and then we get another vector another triple which represents the difference between their values so we have si minus SJ AI minus AJ fi minus FJ and we see that if this if each of these is small that means that si is similar to SJ AI Soulard AJ F has similar FJ then if each of these is small then we can say these people are pretty similar because they have similar sex similar age similar they pages somewhere around to be on the 10x10 similar Affairs right so we can go ahead and say that but let's do a little fix right here because we we don't want the order to matter because we don't want this if we took a VJ minus VI then this would be negative rather than positive if it was positive in this case so we don't the order to matter so when we're compressing all this down to a single number rather than just adding up these three components we're gonna be adding up their squares so what we're gonna be doing is I'm gonna do this in a different color we're gonna be doing si minus SJ squared plus AI minus AJ squared plus fi minus FJ squared and then you can take the square root to come get it back into the units we were working in before and this is known as you probably know it as the norm of a vector the Euclidean or l2 norm so that's written as VI minus VJ with a subscript two you can write the subscript to most people if you don't put a subscript to it they'll just know it as the norm but again there's the l1 norm and the l2 norm so just clear this is the l2 norm so to be clear what we're doing here is we're taking the subtraction of the sexes remember sex 1 is female zeros mayl taking the subtraction of their sexes squaring that subtraction of their age is squaring it subtraction of their fares squaring it taking the square root and the square is just so the order doesn't matter so the negatives don't really matter and we have a nice value right here now we still have a little bit of ways to go because the problem one one initial issue here is that the magnitude of these things are pretty different let me show you what I mean so let's say one person is male and one person is female that means we'll have a 1 minus 0 or a 0 minus 1 or whatever in the end this is going to be 1 squared which is equal to 1 so it's gonna contribute that much to the difference now that's actually a huge difference because that's the biggest difference you can have in sex you can't get any bigger or difference than that so this should be actually a big deal now is it a big deal compared to these other components let's look at fare so for example if someone paid let's say 18 dollars to be on the Titanic and someone else paid let's say 22 dollars to be on the Titanic that's not a big difference when you look at the range of values that people actually paid there's some as high as 512 there's there's some pretty high values so but if we take the subtraction of these two we get we get 4 and we square that and that becomes 16 that's huge in comparison to this one that pales in comparison comparison to this one this one is very small so this one kind of gets swamped by this which is not very big deal so we kind of want to in a way normalize all these values so that they kind of have the same pull the same impact on the overall difference between people so how are we gonna do that we're going to actually you can do many different things you can kind of use whatever method you want that you think will kind of get them in the same units for example one that we're not going to use but which will mention is that you can take instance you can take instead of taking age I and HJ you can take AI as the age of a person minus the minimum of all the ages divided by the maximum of the ages over minutes let me just under formula so instead of taking AI let's say you can take MI which is the modified age so you can take modified age I is equal to the age of a person minus the min of the ages over the max of the ages minus the min of the ages so this kind of normalizes because of Max this is gonna be a number between 0 & 1 right because you can't have age I be any bigger than max of the ages so this is going to be between 0 & 1 which kind of brings into the same ballpark as the sexes which are between zero and one so we can do that but what we're going to do is we're quite literally going to normalize so we're going to the normalization metric which involves was so if you normalize something but just say ni is equal to if you take the values so we're going to do H I minus the mean of the ages over the standard deviation of the ages so you might have seen this in your stats class if not it's just a way to get the mean of all your n is to 0 and the standard deviation of your all your n is to be equal to one so now this kind of really only works if the data you're looking at is already normal and and normal means that it kind of looks like a normal distribution or very close to revolution like that so let's look at how do the ages look so we're gonna go ahead and look at a plot of the ages right now so here it is from this histogram of the ages we see the ages look pretty normally distributed they look pretty close to a normal distribution so we're qualified we're validated in using our standardization metric from the previous page now let's compare that with the fares instead if we look at a histogram of the fares shown here it looks kind of like the right half of a normal distribution or it has a long tail going in the right hand direction so right away it doesn't seem like we're qualified in using the standardization metric because this isn't it doesn't really look like a normal distribution now there's a little trick we can do when we have a distribution like this where we have a tail going off in the right direction and high values up front we can take the natural log of it to get something that looks more normal so let's go ahead and do that and then we take the natural log of these fare values we get a histogram which looks like this so this is now natural log fair versus frequency so this looks more normal maybe we're more qualified using this in our standardization metrics are taking the normalized of the natural log fares instead of just affairs so that's what we'll do so now I also turn to the previous work and using these transform metric so now that we did all those data transformations we have quite a different metric set than we had before so remember before on the back of the page we were defining VI as someone's a triple of their sex their age and their fare now since we add to those data transformations to kind of get things in the same ballpark and make sure each thing had a similar impact right now currently we're looking at is VI is now defined as sex den change that's still si now instead of age we're looking at normalized age so we'll put us na there a normalized age of the earth person and instead of fair we're looking at normalized natural log affair so we're looking at normalized natural log affair of the ice person so this is now our triple and similarly for VJ it'll be all the analogs with the Jade person so now this does that work so if we take these values and so the previous values and we put them into this metric is that better in a sense it is a little bit better because we now have things kinda in the same ballpark as this sex variable so we have we at least we know that their mean is 0 and the standard standard deviation is 1 so they're kind of in the same ballpark but that still doesn't guarantee that they're gonna have the same impact so now the next thing we did to kind of tweak it is something we just did kind of empirically so what I did was I'll just write a VJ here so we can compare SJ + aj + ln f ji promises fugly notation we'll go in a little while but so yeah we had VI and VJ so now what I did was I took the subtraction and I did that so I did VI minus VJ and I took the average of these differences so I took Si minus SJ and I took the average of that difference over a bunch of different pairs of people in the sub sample I was looking at and I did the same thing for nain AJ yeah the difference between Na and the difference and LNF and i found that the average of the differences for the sexes turned out to be 0.4 T 6 the average of the difference for the natural the normalized ages was 1.1 and the average of the difference for the normalized natural logs of the fairs was 0.3 so that's that's okay it's a little better at least they're kind of in the same ballpark but still we see that this is more than double of this more than almost triple or more than triple of this so the simple thing we can do is kind of just use these empirical values to assign coefficients to to each of these differences to kind of put them make sure their averages are the same or very close to the same so what I mean by that is instead of doing SI minus SJ we're going to do we're gonna put a K s multiplied by house I minus SJ we're gonna have a K a multiplier AI minus AJ and a KF x fi minus FJ where K so this would be for the ages so we're gonna leave we're going to say K ay is equal to 1 because we're just gonna we're gonna baseline everything at one point 1 we're gonna say that K s is equal to 1 point 1 over 0.4 t6 and we're gonna say K F is equal to 1 point 1 over 0.3 and the reason I'm doing that so let's demonstrate with for example KS so the reason I'm doing KS as 1 point 1 over point 46 is that if we do that so if we do KS is 1 point 1 over point 46 then that multiplies by SI minus SJ right and we saw that si minus s J's mean was point 46 so if we systematically multiply those differences by one point one over point 46 then their average in the end will come out to one point one and the same logic applies for the natural normalized natural log of the Feres that happens since they're systematically their mean is point three if we systematically apply that 1.1 over point three then their mean will come out to the difference with the difference of the mean of the difference sorry will come out to one point one so in the way we've tweaked it so we've initially tweaked it by making sure that we normalize it in this case taking the natural log to get in the same ballpark and now we've done the final tweaking so now you ask are we done now the answer is still no we're almost done let's look at why we're not fully done so let's write down what we have so far so that we can see where we have to go from that so we'll write KS si minus SJ squared plus K ck8 this is n AI minus n a J squared and finally plus K F this is n ln fi minus n ln J squared so this is our current this is our current trial metric does this live up to the expectations we want the answer is almost so the problem with this is that the more similar two people are which means the more the closer that their eye and J values are for each of these this will be closer to zero right because if people are very very in the in the most extreme case if we're looking at the same person with himself or herself then this is going to be 0 this will be 0 this will be 0 so this entire thing is 0 but we want kind of the opposite we want the more similar they are we want it to be 1 because that's going to help us in our calculations later on so something we can do very easy thing this is called the it's kind of it's called the Gaussian norm is doing e the Euler's number to the negative all of this initially you're saying why am i doing this that looks really ugly but why we're doing that is think about it so if we have a person with herself then we're gonna have 0 here 0 here 0 here this is 0 negative 0 is still 0 each of 0 is 1 so we've satisfied the condition we're but that we wanted where if we have a person with himself then they have a similarity of 1 now what if we have two people who are very dissimilar so that we're gonna say what that means is that these numbers are very very big so this is very big this is very big this is very big just ignore these coefficients let's just say these are very big so that this entire square root is very very big negative of a very big thing is a negative is a very small thing so very huge in the negative direction and e to the power of something that is very very negative is going to be close to zero right because if you look at a graph of the exponential function e to the X looks like this so if we're looking at values all the way back here is gonna be really close to zero which is what we want so this is the similarity function we're going to use in the end so now I'd like to really quickly recap everything we did because I know this was a very iterative process and it might have lost some of you along the way I know it can be confusing and after we finished recapping will give the conclusion what the results were so just a recap we want to define a measure of similarity between each two people each two passengers and we said our baseline the things we're going to use to do this are some one sex their age and their fare how much they paid to be on the Titanic now we said that we can put each of those in a vector and we can subtract them the initial problem here was that we didn't want to do something minus something because if we flipped it then we would get a negative here and we didn't want that to be any different so instead we took the square of each of these components each of these differences and then we took the square root to get it back in the same units so our initial guess was that we took the square root of each of these differences squared now we said that the problem with that was that the the magnitudes of these were very different so that if we had a sex difference of 1 that's the biggest you can yet that should be a big deal but that's kind of swamped if you have even a fair difference of 4 because that ends up having a 16 when you square it so we said to kind of account for this we're gonna have to normalize our different metrics so we're gonna keep sex where it is because it's already between a 0 & 1 range but we want to kind of bring the ages and the fares to that kind of range as well so the ages when we looked at the graph it was already kind of normal so we were able to use the normalization metric here on it so we were able to take each age minus the mean of all the ages over the standard deviation of ages which guarantees that our new metric which is normalized age I will have a mean of 0 and a standard deviation mean of 0 and a standard deviation of 1 and that kind of helps it move towards the same range that sex is in and have the same kind of impact now we wanted to do a similar thing on the fairs now the problem was when we looked at the fairs they weren't looking very normal in fact they were looking like this but we said that in explanation if something looks like this if you take the natural log of it it kind of looks more normal so we did take the natural log and then we do the same treatment as we did two ages we took the normalization metric with affairs instead of ages here so now we helped each of these metrics right here be in a zero to one ish range and I'm being kind of fuzzy here because we saw that with this with these values they weren't exactly having the same impact so let's get to that now so after we did all the transformations here to these two and left the sex as it was we found that the mean difference of sexes was 0.4 to 6 the mean difference of normalized ages was 1.1 and the mean difference of normalized natural log affairs was point 3 we wanted to get these mean differences I do it to be about the same so we just pick one of them arbitrarily we said this one we're gonna keep the normalized ages as it is and we're gonna try to move the sex difference and the normalized natural log fares to move towards that one point one how we did that was simply we multiply the differences in our function here we want all the differences by a coefficient each of them by coefficient sin or since we're keeping the age is where it was we just left the age coefficient has one because we wanted to leave it or was and we took the other coefficients and we multiplied by this number divided by respectively either of these numbers so that when we took the mean of those differences it would go to one point one as well and after we did that so we had currently we were here we took the coefficients now we said that if we just had that as it was it was a bit of a problem because if we had some person with herself then their similarity would be zero if we had people who were very very different their similarity would be huge unbounded actually so this e to the negative that that we just found before helps keeps things bounded between 0 and 1 and it and make sure that people with 1 are exactly similar people near zero are very very very dissimilar which is exactly one and just had a name we can say it was the Gaussian norm it's often used in the K nearest neighbor as the UH as a similarity function so this is it this is our finalized similarity function now let's go back to what we said the beginning of the video what do we do with this now this similarity function is what's used to generate this similarity matrix so let's say we want to find the similar between the first person the second person which is business so we take the first person's matrix which we there their information vector which looks like this and the second person's which looks like this and we take each of these components and run it through this formula and we get a number and that number is what goes here and here and that's how we fill in the entire table now let's say we want to figure out we want to try to predict the fate of the first person right so how do we do it so we're gonna look at all n of these people in this table let's say there's a hundred people in this table so we took a subsample of 100 we're going to look at all hundred of those people and now I told you I'll tell you the K I used k equals 15 you can experiment different case if you want but I use K equals 15 so we look at which of these hundred numbers which 15 of these hundred numbers is the biggest we're going to of course exclude this first one because obviously this one's gonna be the biggest because the first person with himself has the biggest similarity but it you can't you don't know the fate of this person so since you're trying to figure that out so we're gonna exclude the first one so out of these ninety-nine other numbers you pick the 15 biggest ones because those are the most similar to the first person let's say you have a list of those 15 people's so you have that list and now you look at the fate of each of those 15 people so let's say we have survived survived let's say we had nine survivors out of those 15 people and we had six nonsurvivors since more people survived we're gonna predict that the first person also survived it's as simple as that and the W you having an odd numbers that you never have a tie so it's either going to be swaying in the survival direction or the non survival direction okay so and you do that for each of the hundred people you just go through their rows and it's really nice to be in a matrix it's really convenient you go through each of those rows make sure you exclude themselves because you have no information about them you're trying to figure that out and you look at the other 99 people and you look at the top 15 and you see what their fate was of those people and you prescribed this person the same fate under the assumption that since they're similar to those other people they should have a similar fate so we do exactly that and the accuracy that we're getting the the accuracy they would get with this is about 60 sorry seventy six point five percent which is around the same ballpark as we were getting with our regressions so we can tell that this method is working it's doing its job pretty well so let's talk about a few of the caveats before we end the video of course one of the caveats we already mentioned was that if you double the size of the data set since if we have a hundred people if you do 200 you're not doubling your work you're quadrupling your work because this matrix gets four times bigger it gets twice as long twice as twice as wide so it gets four times bigger so that will be a strain on your computer when you do all this then other thing is that these coefficients are pretty much arbitrary you can change them to put more emphasis on any of the factors you want for example I if I made KS like 1000 then sex would be the dominating factor here because these would be very small in comparison to this bigger one if I made this KS very small like point zeros or one or even zero then there's pretty much vanishes and it's only about these two so you want to make sure if you want them to be weighted equally then you have to kind of do the analysis we did back here if you purposely want for one of them to be weighted more than you can go ahead and do that and change your coefficients accordingly it's just that you have to know what you're doing you have to know what kind of things are going to into here also one of the biggest caveats and the last one talked about is a similarity function itself nobody says in K nearest neighbor that you have to use this similarity function here you can really use whatever you want we we try to explain one earlier where you can normalize things with this kind of easier norm than the Gaussian then the this standardization only we're using here so there's a lot of things you can define you can be creative here and maybe you even get higher accuracy as if you do things your way and of course one of the other things that you're defining is your K you can have K be as I chose 15 but what of what's stopping me from choosing 21 or something like you know 45 whatever or even something smaller than 15 like 7 so it's really up to you it's a lot of tweaking that's what I found using K nearest neighbor but that's just kind of how much in learning is as you get deeper into it there's a lot of parameter tweaking but hopefully you learn something and you're you're you're convinced that this method does have a pretty good accuracy 76.5 that's the same ballpark we're getting before and it's definitely better than the baseline that we were trying to the beat beat the crude probability so until next time we'll go more to machine learning
Original Description
K-nearest Neighbor Machine Learning Technique
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from ritvikmath · ritvikmath · 31 of 60
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
▶
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Math Team Update
ritvikmath
Single Variable Calculus Volume of a Sphere - Proof 1
ritvikmath
Single Variable Calculus Volume of a Sphere - Proof 2
ritvikmath
Multivariable Calculus Volume of a Sphere Proof - Triple Integrals
ritvikmath
Multivariable Calculus Volume of a Sphere Proof - Double Integrals
ritvikmath
The Euclidian Algorithm
ritvikmath
Proving the Chain Rule
ritvikmath
Proving the Fundamental Theorem of Calculus Part 1
ritvikmath
Proving the Fundamental Theorem of Calculus Part 2
ritvikmath
Math Puzzle - Poison Perplexity
ritvikmath
Math Puzzle - Poison Perplexity - Solution
ritvikmath
Expected Value and Variance of Continuous Random Variables (Calculus)
ritvikmath
Expected Value and Variance of Discrete Random Variables (No Calculus)
ritvikmath
Array Method
ritvikmath
Complex Power Series and their Derivatives
ritvikmath
Distributions - Intro
ritvikmath
The Poisson Distribution
ritvikmath
The Bernoulli Distribution
ritvikmath
The Binomial Distribution
ritvikmath
The Continuous Uniform Distribution
ritvikmath
The Geometric Distribution
ritvikmath
The Triangular Distribution
ritvikmath
The Exponential Distribution
ritvikmath
The Borel Distribution + Notes on Poisson Distribution
ritvikmath
The Gamma Distribution
ritvikmath
The Normal Distribution
ritvikmath
The Laplace Distribution
ritvikmath
The Chi - Squared Distribution
ritvikmath
Overfitting
ritvikmath
Vector Norms
ritvikmath
Truths Behind the Titanic : K-Nearest Neighbor
ritvikmath
The Mathematics of Breakups
ritvikmath
Sillyfish
ritvikmath
Finding Optimal Paths - Dynamic Programming
ritvikmath
HowToDataScience : Scraping Twitter Data
ritvikmath
Decision Trees
ritvikmath
Perceptron
ritvikmath
Naive Bayes
ritvikmath
K-Nearest Neighbor
ritvikmath
Evaluating Machine Learning Models
ritvikmath
Decision Tree Pruning
ritvikmath
K-Means Clustering
ritvikmath
Gaussian Mixture Model
ritvikmath
Data Science - Fuzzy Record Matching
ritvikmath
Time Series Talk : Autocorrelation and Partial Autocorrelation
ritvikmath
Time Series Talk : Autoregressive Model
ritvikmath
Time Series Talk : Moving Average Model
ritvikmath
Time Series Talk : ARMA Model
ritvikmath
Time Series Talk : ARCH Model
ritvikmath
Time Series Talk : White Noise
ritvikmath
Time Series Talk : Stationarity
ritvikmath
Time Series Talk : ARIMA Model
ritvikmath
Time Series Talk : Lag Operator
ritvikmath
Time Series Talk : What is Seasonality ?
ritvikmath
Time Series Talk : Seasonal ARIMA Model
ritvikmath
So ... What Actually is a Matrix ? : Data Science Basics
ritvikmath
Derivative of a Matrix : Data Science Basics
ritvikmath
Basics of PCA (Principal Component Analysis) : Data Science Concepts
ritvikmath
Eigenvalues & Eigenvectors : Data Science Basics
ritvikmath
The Covariance Matrix : Data Science Basics
ritvikmath
More on: ML Maths Basics
View skill →Related Reads
📰
📰
📰
📰
“Los Movimientos”: The Routing Problem That Nearly Broke My Spirit
Towards Data Science
How Guardoc transforms medical document processing with Amazon Nova models
AWS Machine Learning
The Reward Calibrator That Learns the Shape of Its Own Judgment
Dev.to · Daniel Romitelli
Reducing Human Annotation with ML Active Learning
Towards Data Science
🎓
Tutor Explanation
DeepCamp AI