Building Machine Learning Model from Scratch | Community Webinar

Data Science Dojo · Beginner ·🧠 Large Language Models ·8y ago

Key Takeaways

Building a machine learning model from scratch is explained in a community webinar

Full Transcript

I'm the chief data scientist for this sum the me a company called data science dojo as well so I'm the chief data scientist for that company we started around almost three years ago my own background actually has been working in this space for more than a decade now I have worked in in bought detection machine learning online ads relevance on an experimentation maybe testing so I've been doing this for some time my background prior to joining industry I have a PhD in computer science and my focus was computer vision and data mining and really this this whole data science buzzword really it is a buzzword that we cannot deny now right so it is but computer vision actually is data science you are doing data science on images actuaries they were doing data science and insurance industry right so some of you may be may be may have seen that may have a degree in statistics and you already are doing you have been doing data science really but it's a it's a buzzword really variable you apply machine learning and starts and this space in whatever space you are flying it in that's really what data science is so why why this talk so from my experience I see this unnecessary emphasis on on machine learning the Machine if the specifically the machine learning aspect of things the machine learning modeling part of it all the deep learning and support vector machines and boosting and gradient booster and grand and forest right so I'm throwing terms at you but some of you may have heard of this and the the problem that those of you who have who have practiced or our protect practitioners in machine learning you would probably realize that machine learning the actual model building it's a very tiny part of the actual problem and people expect miracles to happen right so I will I will get this library this deep learning library or this support vector machine library and the life is going to be good right I was just take all my data and throw data at this model or this learning algorithm and it will solve all the problems for me but those of you who are practitioners you know that it doesn't work that way it's only a very tiny part of the problem even though there is that unnecessary emphasis on the learning or modeling part of it machine learning is hard work you have to spend time on your data you're choosing the right matrix having the right understanding of the business so what I will emphasize this is more of a beginner level tutorial I would try not to avoid the technicalities of it because one hour might not be enough for to get into the deeper discussions here but what I would like to emphasize here is that you don't need to actually you don't need to know a lot of machine learning algorithms to actually do machine learning that's the first first thing that I will talk about and what kind of problems can data can have if your data has problems no matter what you do your model is not going to be good if you choose the wrong metric for building the model again no matter what you do your model is not going to be good right so so it's a general best practices TED talk and we'll we will we'll start with an example and we'll see where we are at the end of the presentation so let's take this example so what is customer churn customer churn is when when you have a you are a subscription-based business and you don't want people to leave your service and now people start leaving and the question is it's a very important applications out there for machine learning a predictive modeling so can you predict really that this customer is about to leave that's your customer churn problem I mean I'm a t-mobile customer and can people can can they guess that Raja is about to leave t-mobile and let's say I come up with a model that predicts customer churn with 70% accuracy my question for all of you is do you think it's it's a good model a decent model good enough model it's better than 50% right so it looks better than 50% so here is the problem that we are facing right I see these things so people you know in your part of different LinkedIn and Facebook groups and we are everybody is sharing hey I built this model it is 85% accurate what does it even mean to have a model that is 85% accurate why are we why why am I so happy about building a model that is a developer eighty-five percent accurate so I will talk about this how do you dissect how do you dissect this if someone tells you that they built a model that is that has this much accuracy or this much precision recall what sort of questions would you ask them or at least if you don't ask them do you know can you can you spot what's wrong there maybe there's nothing wrong there but at least in the absence of all the is there any information that is missing so in this case the problem is that data literacy it's it's not very common we we look at some stats someone reports the numbers and we would think that yeah I mean it sounds about I mean this model looks awesome in this this model looks like horrible right so so we'll talk about this so let's ask some fundamental questions about this problem the question is but then someone reports that they had this model that is 70% accurate the question is what data was considered what time frame the data was how big was the data set was it only 10 examples or hundred or a thousand or 10,000 examples the next thing is what was the skewing balance I will talk about this in a moment how many in your data set how many customers were actually this is supervisor running problem right so how many cuz of these points in your data set or the rows or records in your data set how many of them were actually customers who churned and how many of them were who did not churn and did anything else change in the business the business was already in a turmoil right main so the business was going through this and via predicting churn but the people were churning anyways right so so you know so these are the questions that we would ask then the other thing is the metric question what if the data set is from period when the 70 when 70% customers were already churning so you pick any hundred customers 70% would leave my model said everyone is leaving am i 70% accurate am i useful it's nice so I mean if you look at it yeah out of every thousand customers 700 are already leaving right and my model is saying everyone is leaving my calculate the accuracy yes 70% of the time my model is correct but what is the model actually actually doing so did you use the right metric what is that sorry I could not hear you yeah yeah yeah so so I mean and we will not get into the precision recall business but what I'm trying to say here is that sometimes I mean you might choose a metric that might make you look good right so in this case accuracy is an amazing metric right so it makes my eye look it looks like that I'm better than random guess but there's not much value that I'm adding here right so I'm just not doing anything really the next question is the repeatability question how many times did you repeat this is it that you handcrafted some sample to make your model look good you would think it doesn't happen right it happens all the time especially in academia when you are trying to when you're trying to report the results so you came up with this novel algorithm right so you have this new algorithm and what do you do you come up with this new algorithm and this new algorithm there are some benchmark datasets and you will on those benchmark datasets you will say okay I will keep running my my own algorithm until it looks good and for the for the other existing state of the art I will keep them coming coming up with shuffle different shuffles for the data set until the state of the art looks inferior to my learning algorithm and I go and report the results this is my my algorithm is better than that better than that state of the art and it happens all the time so the question is if you repeat if you repeat you're given any random data set because on a given data set you might be looking good you might be looking bad but if you repeat this ten times or 20 times on different random data sets can you reproduce these results or not 70 percent does not mean anything really one was for the data reason the other one was for the metric reasons so there was a reason around data right so we talked about that the second one was the the metric reason even even if you are correct on both the data and from the data and the metric angle how we are measuring the performance still if you cannot reproduce it consistently that means your model is not a good model give the your predictions it may be a fluke it you just accidentally got something yes you would actually do some sort of and I will talk about only one technique at the end cross-validation you can do a bootstrap sampling you can actually you can you will try to shuffle and demonstrate that it is repeatable it is not an accident okay this is the most important that the question that is overlooked this is the the customer churn model that we build it is one of the best that we have ever built but you have built this customer churn model and let's start calling them let's start calling them to stop these customers okay so if you look at this so this is an actual example actually this is this company called Telenor they actually built this model that reduced the customer turnover customer churn by 36% now let's look at it let's let's visualize the situation I'm very accurate in predicting customer churn I hand this model to my business team and tell them I'm very confident 98% of the people that you call or 90% of the people that my model says are planning to leave they are planning create your business tip team picks up the phone and starts calling these people and the other person you customer picks up the phone and says yeah I've been waiting for your call I want to leave okay so you basically you you woke up this sleeping giant right so your customer was planning to leave and they were lazy and they were not they wanted to leave but they were delaying that decision right you call them you say yeah I mean I want to leave thank you and this business teams come and now what happens you lose you lose a bunch of customers you expedite their the departure from your for your company so now this business team guy comes back and says you told me that this is a good model and now this guy says did they leave yes they did leave I told you right so they are going to leave and the business team and now they are arguing right so but think about this that I mean if if this they knew how to use this model correctly the situation might not have been like this so this company actually two of these people they they attended one of our trainings in Singapore so the two employees from this company they said that they had a hierarchical model the model was such that the first first they will predict who's good likely to leave or who's going to churn and then there was another model that was run on all the people who were predicted to be churning and then he will predict who's going to stay if called do you see this problem right so even though if it's a it's a good model built with this amazing accuracy or precision recall whatever that metric may be but if you don't use it in the right business context that can be a problem as well is this making sense to everyone are we following this so far okay yeah I mean you would be able to always have a control group right so that's the best practice but even after the control group you mean you saw that I mean you the people you called they you expedited their departure earlier because they were already upset and you predicted that they are unhappy and you basically just call them and then expedited their departure maybe all of us I think most of us we are unhappy with some service i sat on my Netflix account for I don't know three years until I cancelled it if they had called me I would have said yeah I mean cancel it right now I am sitting on another my bondage account I don't use it nothing but paying thirty five forty dollars a month I'm not canceling it if they call me absolutely I will just tell them I mean just just I mean I don't need you so it's it's basically a machine learning used in the wrong without a business context it's actually it will hurt the business and instead of helping it right so even if you have a control group and a treatment repair so we can do all of that but if that's a bad model and yeah if it's a good model used incorrectly you will see the problem yeah so yeah but the question here is that would they would they stay because some customers they might not stay no matter what because coming up with an offer that that is yet another predictive modeling problem so because you don't want to start giving away one year of free service to everyone right you will say if someone can stay with one month of free service let's offer them one month of free service and nothing more than that because you cannot really do a one size fits all type of proach right so and and this is the kind of business context that we are talking about here right so if you don't have that understanding of the business so I might say yeah I mean if any everyone who is staying if you give them one year of free credit to the service they will stay but what is the cost to the business for that is it even worth stopping or keeping the customer yeah yeah I mean so you may want to worry more about customers that are that you make more money from versus low value customer right so that's that's very true any other thoughts comments and if you look at this this has been I mean we have not we are not even talking about machine learning here really per se right so we did not talk about where the deep learning was used or the boosting was used or random forest with use we did not talk about anything we're talking about everything outside of that that black box that is used for building that model yes mm-hmm mm-hmm yeah yeah I mean it depends upon how you how you how you lay your problem out right so maybe maybe if it if you set this as a churn versus no churn you will call this a classification problem in a machine learning context right if you give them some sort of scores that absolutely have a happy with the service versus neutral versus absolutely disgusted right if you have this it's a ranking problem right so you will rank people or maybe a regression problem where you assign them some sort of scores right so it depends how you lay a problem out but and and you will lay your problem out in a you will set the problem up in a way this again overly used an abused word of actionable insight right so anything that is actionable because you can get any insight but is it really actionable so you will come up with something in a way that when you give it to the business team or when you give it to some other partner team can they actually act upon it or not so that's that's the high level idea so you will you will make sure that the model that you're building it is consistent with but with the how the business wants to use it otherwise there would be a disconnect okay any other thoughts questions yes [Music] so I don't think so I emphasize that the model is not a bad model I never said that model is a bad model it is basically the person building the model if he does not he or she does not understand how the model is going to be how the model is going to be used eventually the model can backfire it can hurt more than help and that's why I was saying that I mean this company is the specific company app I mean people who were related to this they I was talking to them they saw that they said that they had a hierarchical approach first you detect who is going to leave and the second second time around you will say ok out of these who's very likely to stay and only it may be a resource issue as well right if you have ten thousand your model saying that ten thousand people are going to leave would you go and pursue every one you have where you could your call center can only make thousand calls possibly you will actually say okay let's find the customers who are going to stay who are very likely to stay or most likely to stay let's call them up there is far more business value in that versus versus just randomly calling everyone okay so the conclusion here is that a good model with bad business judgment is a bad model this is I think no rocket science here right so because as much as a lot of people would want to think when your business does not care how sophisticated your model is and how how clever the you how cleverly you prove the error bounds and all of that in your I mean theory is aside if it doesn't bring any value to the business it is a bad model no matter regardless of how good the MA the model actually was so that's that's common sense but not common practice and that is why anyone anyone who's building a model must have a very good understanding of business as well the domain and the business and I would actually go on to say that even a dev who has a better understanding of business but it's slightly less technical than a dev who's a dev really a good death but does not understand business I would rather go for someone who has slightly to come to super solid dev less super solid dev but this one has a better understanding of business I will go for this one and not for this one because this guy may come up with any bill he would insist yeah I mean look at the library that I wrote I mean what's the business use for this library for me right I mean sir it's the same example here right so you want to you want to have that understanding of business whenever you're building a predictive model sorry yeah yeah it's overall I mean so in the end what I'm trying to say here is it is the business value that you extract out of your data science learning process it is that but what matters and not specifically how clever your implementation was and what kind of library that you're using yes actually I think I think yeah yeah I mean so if if you don't understand the problem space if you don't understand your domain there's no way you can build a good model because there you you cannot extract good features out of your data sometimes your feature means features may not be even obvious to you they might not be sitting right there in front of you you may have to do some sort of transformation and so many things so I think anyone who has a better understanding of domain is going to be far more useful than someone who just understands theory inside out and does not understand the domain yes mm-hmm mm-hmm hmm that's that's a great question right so the question is in the traditional bi or traditional descriptive analytics of someone with a understanding of the domain can go and build a dashboard upon request and against we have something in the predictive or machine learning context so even in that scenario I would say that if the person is not data literate I mean so they may have understanding of the domain but if they don't understand the intricacies of the biases and everything around data that are possible European reports might not be accurate even then right so you're I mean they can be sampling issues you are taking the wrong sample you're taking data from only one day whereas the representation actually for getting a representative sample you should consider at least one week worth of data right you are ignoring certain certain segments so things can happen but but you have a point but in a predictive contact you are actually you are handing this businessperson a working model and you have made some assumptions when you were building the models so it's so the analogy would be is that the business person knows how to build a new model from scratch because the end product is that plot that graph and that report and essentially the the business person is reusing the all the data to create a new report right in this case if they the business person knows how to build a model from scratch well yeah I mean they can do it but they cannot just simply tweak the model to actually just get a different outcome from that maybe maybe not but in general know might not be possible sorry can you repeat that so you're talking about to a machine learning for machine learning right so machine learning can that can do machine learning right who's going to train train that part and what is that data centers going to use it so yours mm-hmm okay so you're saying that yeah so that there is a desire right so I know that we want to we want there is a desire that the machines can learn on their own as much as the industry and the press and journalists they want to make you believe that machines can learn on their own machines are dumb machine learning algorithms are as smart as the person training these machine learning models right so I mean there have been some efforts around that to some extent they can figure things out but in the end some human has to actually go and and intervene so not yet yes absolutely in every sense of the word I did not get there yes absolutely absolutely absolutely so yeah that's that's a great point that David brings in he's a practitioner so he knows what what it means right so the the repeatability was our saying was that if you have a given data set on that data set if I pick up okay let's take this random variation and see if my bottle does well take another random variation from the existing data set and then whether it does well or not but what David is pointing out is what if the whatever your data essentially the distribution in more technical terms what if the distribution of the data is changing right so you had your Amazon and you had only you did not even have a Footwear category I'm making it up right and suddenly a new category gets introduced on your e-commerce website and which is a Footwear category and then you bring this category in would the data be the same no but the behavior be the same no whatever model that you learn it is going to degrade over time as more customers come in as more categories come in and as the ground realities change so that is another source of variation that is absolutely possible yeah and that's the and that's why these companies they have people on their payroll otherwise they will just hire them and then they will say I mean your job is done right so but the the business realities are changing you have more data you're you're going out and seeking more features and improving your models right and sometimes your models will automatically degrade just by virtue of due to the fact that you have different kind of data now you have maybe you had on a given day you had more BOTS those don't represent your actual actual real customers what's going to happen the model will degrade you have now you have marketing traffic that is coming in so all of these things do impact your model and that is why these these companies they have these armies of machine learning engineers and data scientists who actually keep working on these models and keep fixing these models the more the fundamental assumption behind a model is that the distribution of future data that you've the model is going to be used on is the same as the data that it was trained on and this more this assumption is really accurate I can also tell you it's a it's a big assumption we assume that this is going to be the case but in most practical real-world business problems this doesn't this doesn't happen it keeps changing and that's why you have to you're you're really you are in a catch-up mode you are always fixing and fixing your models and fix you fix it something changes you fix it something changes and so on okay so modern machine learning libraries they are actually extremely easy if how many of you have used are or have used our Python socket learn pandas ok tensorflow ok so and you would actually agree with me that building a machine learning model is it's not even a big deal and I've shown this our example just as a sample you split your data into training test samples this is a random forest model literally you tell that you want to predict survival tilde dot dot means all the columns and you pass it the titanic data set and now when you get want to get predictions you use this predict function and done this is this is what it is and it may sound like something yeah I mean I can do machine learning but the reality is a lot of it happens before before you bring this data into it into this library a lot of feature cleaning of features and and extracting features bringing more data and then after that evaluation are you evaluating your model correctly or not did you make sure that you your ma the model that you're building the performance is repeatable or not right so yes I don't actually this is just an example here mm-hmm yes they are both blessing and occurs at the same time yeah because it they they these the modern libraries they make it look so easy I call it I call them a blessing and a curse at the same time because if they make things easy for you if you know what you're doing and if you don't know what you're doing they create an illusion or understanding you might think that you understand not you as in you but general anyway like one one might think that they understand machine learning but the reality is that they don't right it's just that they figured out the library somehow and that's it okay so how how do you build mmm machine learning models and what is the right approach so here is here is what we have here you start with a business question you don't start with you don't actually start with the machine learning model because a foreign for a newbie person who just started they might say yeah I will go and use this this amazing library that I found found out but the deep learning library or tensorflow and just something like that but the reality is that you will ask the business question first then the next thing would be some janitorial work a lot of cleaning data is never clean and if you have clean data you're not working on a problem that matters real-world data sets are never clean they are messy multiple so duplicate clicks or missing clicks or duplicate queries or missing queries and bought traffic I mean so if real world data sets are messy then the next thing would be is building the model you did build the model but do you know how to evaluate the model do you know how to tune the model do you know how to how to set the choose the right parameters did you is your tree at the is a five level deep tree the right one or 10 level deep tree is there a right one or not if you're using regularization are you using the right amount of penalty so the model is complex enough or not complex enough right so all of those trade-offs you have to really be aware of that and you this is what we call the parameter tuning you would tune the model in a way that it it gets the best it gets the best performance out of your circumstances whatever data you have and then eventually you will do some more validation on blind holdout data and online experimentation a/b testing and we will talk about these steps one by one so the first question what I said was ask the business question you should always start with the business question first when you're building a predictive model right where is what is it that will bring value to the to the business you are always going to be doing something that increases your revenue improve your profit margins reduced cost improves customer satisfaction and really adds business value in general and you if you are you've spent some time in industry you will see that someone very smart brilliant they will come up with and suggest something yeah I mean looks great but by should we do this what's the point doing this yeah I mean I I love it it's a great idea what value is it bringing to the business so the question always starts with what will help the business so how do you start with this so if you have some so you you know machine learning now you you have some idea how to build machine learning models how do we identify identify what what would be a good model to build or what would be what should we do in machine learning given our situation you will ask these questions right so we have too many frauds that are happening if we did this it will result in this business metric if we could predict customer churn what if we could forecast sales if we could forecast sales this will happen so you always will actually keep thinking about this wishful thinking let me put it this way right so something is not happening right now in your business and you will ask this question I wish if we were able to do this that would add business value the next thing is has anyone heard of this this thing it's popular right so so you find a lot you you come up we came up with this business question you decided that you want to predict you want to forecast sales what is the next thing that you will do the next thing that you will do is you will say it go and ask this question what are what are certain things what are some pointers and in machine learning language we call them features or predictors right what are some predictors of features that can help me predict sales you will go and start looking at yeah I know that table in that database that has some information around sales and maybe that other team in my company they have some information and maybe this third party data if we can just bring it in join it with our existing data set that will help us forecast sales better so the first thing here is going to be getting the right data if you don't have the right data garbage in garbage out that's that's the rule right so you must get the representation representative data and then build the model but there are so the generally this the idea this database algorithm is that if you have if you have good data even a simple and okay machine learning algorithm it will still help you but if you have the bad data even a better relatively better machine learning algorithm it might not be able to actually help you so it starts with the right data the next thing is you want a lots of data how much data do you need in general what would be your decision how would you decide how much data for a given problem I'm showing a very big question at all of you what would be some considerations okay sure that would be one consideration yeah but let's yeah definitely okay any other consideration mm-hmm so if if the data is still representative or not right so that's probably what you're saying writes up the data is older so you want some sort of representation right so I would freshness is only one aspect of representation but maybe women versus men elderly versus young people living in this zip code versus that zip code yes yeah yeah yeah yeah so yes number of dimensions can you elaborate a bit so is you're concerned that if they are too many dimensions as you're learning going to be difficult on too many dimensions or okay okay yeah so you would actually be so I'm glad that you brought this up because traditional statisticians they still call this estimation right so we call it predictive modeling but it is it is an estimation problem right so I'm going to be 80% accurate it's an estimate it's not a guarantee that it is going to be 80% accurate it's just an estimate just like if I if I go and a toy problem right so our problem would be is that what is the average age of everyone living in Redmond how would you go about estimating that you'll probably yeah what confidence level what's the distribution right so you will go and say yeah maybe if I had if I went for and I think for certain things most people actually could obviously tell that is it wrong for instance if I told you yeah I will step out and ask the first five people who are passing by and take their average and say yeah this is the average age what's wrong with this approach yeah it's yeah it's it's not big enough sample so just come up with just draw a number what one hundred five hundred yeah I'm just I mean because we cannot go into the all the technicalities but maybe someone said yeah we need 500 samples so and I so someone said 500 I said yeah I mean I can go to this elementary school I mean 500 kids I mean as they come out I will ask them their age what's wrong with this approach it's not random enough or maybe so if you said no if you said no not random I will say okay I will pick every 10th kid so what's wrong with that it is random it is still random okay so I mean if a kid comes by and I will flip a coin and decide whether I should ask this kid or not it is it is it is not representative enough right it is not represented in effect so so it is ingrained in our like in our brains really you don't have to be a statistician to understand this but somehow when we do machine learning we forget this so the same thing applies when you're building machine learning model you have to have the right sample size the right sample size that can get come up with a good estimate and a representative sample and going back to the same problem right so if you if you go and ask people their age they might lie data quality issue you might want to verify I mean if you want the estimate to be correct you might want to verify show us your ID possibly right so so all of those things they matter right so data quality and variety and and besides you have to you have to actually extract useful features out of your out of your data so that's sometimes that the the the features of predictors that are needed they might not be right there in front of you you may have to do some sort of transformation to actually get the features out of it there is this common method so some sometimes people would think that there is a yeah we'll take this black box and put every no matter what data is given to us I mean this this algorithm is amazing right so we'll just give it anything it will just come up with a model it doesn't happen this way it doesn't work that way so this is this is a map and the other thing is that you have to really make sure that you spend I call it the 8020 rule spend a lot of time on your data that will save you a lot of trouble later on when you build the model if you spend do your due diligence while handling data acquiring data cleaning data engineering of features it will save you a lot of trouble when you have actually built the model your model is going to be far more robust and better if it if you do your due diligence in the data in the data stage let's call it the whole data stage from collection to all the way to feature engineering and cleaning so on there is the this idea acquire as much data as you can sometimes the data may be sitting somewhere nearby and you may not know that the data even exists some some customer data you are someone came and told you yeah I mean we need to predict this and I think go to that table and you are just stuck with that table in a given database but maybe there is another table in the same database maybe your partner team has some data set maybe the third party has some data set go and get as much data as possible more data is always better than less data okay the question is can everything be predicted that's also a very interesting problem right so I mean so can you predict customer churn yeah I mean companies use it but it will again depend upon what data you have outcome of election results can that be predicted yeah it can be predicted but how accurate is the prediction so I was I saw a lot of interesting tweets and blog posts around data science failed and the recent presidential elections has anyone does anyone have any any thoughts on that did data science actually fail or something else happened maybe failed maybe didn't yeah that's a great point yes but they have worked in the past everything they have largely worked in most elections they are both but I mean there was something different about this election that they did not work right so hmm a few pre model did work right but and many most notably I think Nate Silver's model did not work right so that was okay so I mean that you bring up another point right so I mean so yeah I mean it it was within that range of confidence right so yeah I mean yeah okay so just the fundamental assumption when you build a model like this a predictive model is that the distribution of future data is going to be similar to the past but what if the voters change their mind because of some news coming in it's not the same data anymore right what if people were did not actually did not actually they were not they were not really candid about who they are going to vote for right they may be the men said something and they said it otherwise what if the people who said that they will vote for a certain candidate did not show up at the polling station data science cannot take care of that so so all of those things are there because these are very well established techniques and yes your predictions can be wrong but there are so many factors that again revolve around your data if your data changes the predictions are going to be messed up [Music] [Music] yeah I mean so a lot of things right I mean and that's that's what is tricky about when you're working on data that's what is tricky about working on data that yeah I mean so I mean they have been building this for a long time and now they did not factor certain things and and maybe they could not have factored everything in right so it's it's very difficult sometimes to predict everything that is possible yeah yeah I mean whatever whatever approach you took I mean it will result in different outcomes yes yes we are referring to the data attributes yes so in case of plain text what are the features or the dimensions so you use some techniques natural language processing some text analytics to actually convert data into an equivalent representation so there are ways to actually convert so you might be seeing some something as a tweet whereas the machine learning war model they only works on columns and rows right so you will convert every tweet into a bunch of a row with a bunch of columns and if you have 10,000 tweets all of them will be a bunch of rows and columns and so there are actually techniques to handle that yes I mean they are actually a model that predicts Crock stock prices but they are probably not for not for everyone right so and because they so these companies that do are in this business they hire the best and the brightest from the top most colleges the best coders and best machine learning right so best of everything and to the point that they actually have their servers very close to where the piece information is because even a few seconds a few milliseconds of lag that can actually impact the decision making right so yes it can be predicted the next market crash is coming okay yeah I mean it's interesting right that if if this this becomes more commoditized these high-speed trading models then what would what would it look like what would shock market look like because it's almost like it depends upon which service is bidding on my behalf right because humans are going to lose because machines can process so much data so quickly so that's okay that's a great question let me try to come up with an example here yeah yeah I think for some problems having a model that is 90% accurate it might not be enough because it's a very simple problem right it's almost like I mean what's the big deal right so anyone can do it right but for some problems even getting a model that is 60% accurate it might be incredibly hard so there is no fixed constant number for a metric and it might not even be accuracy that you're looking for you might be completely totally looking for something different so there's no there is no common metric across all the problem domains so in some cases you might look for accuracy in some [Music] yeah you will compare them but there's more to it and I think it did that maybe as somewhat of an involved discussion but I think the way we should look at this is that I think your question was is is 70% a good accuracy or not is 80% a good accuracy or not and my answer is I cannot comment on it unless I know the problem domain I know what data was available and I'd or I know what was the state of the art I mean what has someone else built already if I know that ten other people have built a model on the same data set that is consistent showing 90% accuracy anything that the less than 90 percent is bad but if nobody has even reached 60% no matter what they did 60% is a great accuracy right so or are even more fundamental question might be is is accuracy the right way to even measure that metrics are not that metric or not just like what we saw here right so I think the answer would be it depends upon the problem space you had some coming [Music] yep yep absolutely absolutely that's a good example right yeah and you might go to recall or a procession based on what the problem is space is right yes yep that's that's a great point yep and if you are I think also be aware of this thing if you consistently get accuracies in the range of 98 99 percent or something along those lines you're probably solving a problem that nobody cares about because if it was that easy you would not even need a model for that real-world models they would actually not have that high performances right so I mean in general you would need a predictive model for something that's unpredictable right and if it is so predictable what's the big deal so be aware of the the these kind of pitfalls are there let me see I think I will quickly try to wrap up here it's already 7:30 let me let me quickly talk about this right so in this case what what features would he use there is this example I found this somewhere that the Facebook recognizes faces with a 97% accuracy and they claim that and I'm talking about recognition not detection right so you recognize who this person is with a 97% accuracy and it is performing better than humans what is going on how can an alga than me better than humans is it possible okay so is it a level playing field yeah I mean I was here but what kind of magic I mean so I mean it's almost like I mean so I keep hearing them deep learning and Watson these are two things that you keep hearing right so Watson one jeopardy right so is it really Watson that one jeopardy or something else yes yeah no but I mean we are talking about now in the if the Train test sets are the same this is a wrong practice fundamentally so I'm assuming Facebook Facebook did not do that right they did not report results in that fashion right so I'm assuming that they were following all the best practices and these are actual test sets they were never used in training and so on right so but they still showed these results and here's the thing do you think that humans are at a disadvantage in this case because humans are only looking at these based on visual similarity here yeah so this is Bob and this is Joe and this is Jen right of course of course I mean so I mean based on whether I know them or not I'm with of course I mean them I would expect them to have run tests like this right so yeah yeah so I know that I know everyone in the picture whether I can recognize them or not of what advantage does do I have an advantage or the algorithm mm-hmm okay great yeah you had something [Music] yeah okay yeah so here's the thing right so mostly when humans do it humans are going to the only queue that they have is more of a visual queue right so they look at it and maybe I recognize this person yeah I mean if this is this person then this must be the person next to that person right but in the case of a machine learning algorithm or the Facebook algorithm they know they have a lot of other context they know who liked this image or who shared who commented who tagged it who viewed it and so they have far more information than moti verdicts that you would excite so it is necessary inherently humans are at a disadvantage we call it features right so a Facebook algorithm that is recognizing faces has far more information than a human to actually come up with this conclusion so yeah I mean this is no surprises it's not that magically deep learning is figuring this out deep learning actually is relying on something that actually is figuring this out I mean so it's just like Watson thing right so iBM has done an amazing amazing job in in evangelizing Watson but you you you if you had given any other machine learning a library to the same group of people who built those algorithms they would have come up with the same outcome it is it is it is not Watson that did it but somehow based on my experience with the number of people they think that it is actually IBM Watson that did all of this without realizing that it's not Watson okay so so what are some data related activities that you will be involved in acquiring data from all possible sources sampling transforming cleansing data exploration visualization feature engineering tons of things that you would do here and this is the most painful step in your machine in your machine learning model building cycle yes yes I would say I would say that right yeah yeah yeah yeah they have they actually they employed the best people to work on that project and they got the results that they wanted the next thing is let's let's wrap it up very quickly what the success looked like I think we have to we have to actually be very clear about what is it that you want out of your model and I think some of you already mentioned this example if I have in a model that actually predicts this is the situation right I'm predicting frauds you have number of non frauds as nine thousand nine and ninety n number of frauds there's only ten my model says nothing is a fraud life is good is my model accurate it is very accurate actually in a machine learning context we are we are not talking about accuracy in in English right so we're talking with a machine learning context accuracy is accuracy is your total number of correct predictions when the prediction is when you say it is a fraud and it was actually a fraud we call them true positives were are in true negative so you basically your correct predictions divided by your total number of predictions and in that case looks great and in those scenarios we have some metrics which are call we call them precision and recall and F F score F F measure so all of those metrics exist but this is not it it is not only going to be is it it's not going to be just it is not always going to be that you are in a situation where it's yes and no and fraud and non fraud or bought and non-god right so sometimes you might need how far off are you like Zillow Zillow would not say whether you're accurate you're you're accurate or not I mean you would not measure a Zillow prediction predictor model whether it is correct or not or what is the accuracy you're going to say how much does it deviate from the actual price if the house price was a million did it predict 900,000 or 1.1 million and then and basically you will sum it up which is called the mean absolute error or mean squared error root mean square error so there are the point here is to actually give you an idea that there are other metrics for other problems if you're in a ranking setting you might not want to use mean absolute error you might want to use a metric called NDC G so you choose the right metric for your problem if you don't choose the right metric for your problem if you don't measure it correctly you don't know what's going on okay so the next thing is what if I built my model and I think it is awesome and in production it is doesn't look that awesome is it possible it happens all the time okay so and how do you find out I mean so my model shows an accuracy of 90% in the training environment but the question is would my model be equally accurate when I deploy it I'm predicting customer churn 90% of the 90% correctly in my training set my training set is my historical data but you don't build a predictive model to do well on the training data training data is just a proxy for what is to come in future right so if you build a predictive model that looks 90% accurate and your training data what is the guarantee that it is going to look equally accurate in the future 0 yeah I mean depends upon I mean how how skeptical you are but and there are actually there are formal ways actually of verifying this so there is this concept of generalization and in machine learning generalization is extremely extremely important you have to build a model that is generalizable and what is a generalizable model well then the name sounds right so you want to build a model that is no matter what data set you bring in of course from the same domain the same distribution no matter what data set you bring in i roughly perform the same way if your model can exhibit these characteristics it is a generalizable model and most of you may have heard of this term overfitting what is overfitting overfitting is lack of generalization you're you're very you're very focused on a specific data set you only worried about that specific data set any different variation that is brought to you from that is slightly here slightly there you don't do well that is what overfitting is and most people actually have heard of this term if they if they work on or they're learning machine learning or predictive modeling but not many people actually understand the ideas around generalization overfitting and these are important concepts to know what else we have there's this idea of I mean one might be tempted to say yeah I will train the model and test on the same data but there is this practice of training the data I split the training data into partitioning 80/20 70/30 50-50 and then what you do is you train your data on this 70 percent and doesn't have to be right so sometimes I mean people have take it very like it's almost like anything other than 70% and 30% 30 is its mm-hmm it is not allowed right so I mean you can do 60/40 it is totally fine and you can do 50/50 that's also ok you can depending on what your situation is you can use any split there but the very here is that this is not enough there is this idea of a blind holdout data set and the blind holdout data set is something that is set aside more than your test set so you will actually create this 70/30 on this and this data set is I mean in a good a good practice would be is that this data set is locked away from you you if you are the one building this model you have no access to this data set and essentially it's all most likely a future a future data because that's what your model is going to be tested on so someone else is controlling this data you are going to build your model on on this data set using the 70/30 split and then someone else will just go and validate and then tell you how good is your model just like real life right so you get only one shot one chance whether you predict fraud correctly or you don't predict it correctly you cannot just go and keep checking yeah am i correct no I mean let's fix it I make correct no go let's fix it right - it doesn't work that way so even if you do 70/30 split chances are that you might actually mess up and that's why you have this the idea of blind holdout data set okay I talked about this I will only at high level talk about this this technique we call it cross validation and cross validation is actually a technique where you actually split your data into into ten random ten random partitions and instead of training and testing it once you train and test it ten times over and once you have when you do it ten times over and if now your model looks good every single time that means you have come up with a good model because you're showing both that your that your performance of the model is good and on average it is good and it doesn't fluctuate a lot so that's that's the whole idea about cross validation and what I would for those of you who are serious about building machine learning models and you don't know cross validation I would actually strongly strongly recommend that you learn cross validation as as soon as possible yes if you are using a completely separate the blind holdout you mean if you do it this way you have is there still a risk of overfitting yes but you're minimizing the risk of overfitting if you you will get some sort of statistical guarantee if you're getting your estimated with a 95% confidence right so one in 20 times you might mess up but 95% of the time you will be okay with many assumptions being there there are a lot of assumptions that we make and once you do this you're going to actually tweak your model you are going to tweak your model deeper tree less deep tree higher penalty for model low penalty for model more trees less trees so you will twin tune your model and you would tune it to the point that it gives you a consistently good consistently good results and now there are other techniques that can be used as long as you create these random multiple models on random datasets that is important you have to randomize and see if your model consistently on different random data sets it gives you a good train test error and it doesn't fluctuate a lot that is what you're looking for okay also anything that anything can can change in data many things I mean data has maybe your business has changing I've talked about this right so how do you ensure that your model is indeed good you will actually go and run a controlled experiment on a small sample and gradually increase adoption we call it the a/b testing or online experimentation you are going to actually go and maybe try up on a point five percent traffic or 1% traffic if you are running an online service or maybe on 1% customers or two-person customers to minimize the impact on your business and events you have verified that this works you will actually start increasing the adoption and keep going ok I think that that is it I will stop here I have skipped a few slides but in general I think most has been covered so any questions about anything yes I don't have any experience we are using deep learning for any real scenarios no and so I'm just curious out curiosity [Music] mmm-hmm but I mean how to what extent right so I I see that deep learning might be a good one for images and natural languages speech and so on but it's not we should not look at this as a one size fits all right but in certain domains deep learning might be a good tool to use because it has the ability to actually generate all the other features generate features on its own but it may not be universally applicable for across all the domains as long as we understand the distinction yes I mean it's yeah yeah AB not sorry yeah yeah yeah and that's that's what my intuition is for a real-world problem computer vision I'm sure any image processing anything that's I would I shouldn't say that has a that doesn't have that wildly varying structure faces yeah I mean how many variations kind of phase have or object human object right okay great yeah yeah but it's a great question right oh yeah it's yeah but this there's hype and there's reality right so and if you understand all of these best practices I mean no matter what machine learning algorithm you pick it is going to be a decent model right so that's and that's why I think I never mentioned what machine learning algorithm we will be using right that's all it was everything but machine learning if you put it some people might say is the machine learning right so yeah this is what machine learning is powerful I would not actually here even try to if you knew what are your odds of winning then you would not play actually or maybe play and pray all right I don't I don't know this is again a feature engineering problem right so because maybe you you collapsed them too soon you'd combine them into certain categories that they should you shouldn't have or maybe you should have grouped them through categories because individually there were way too many categories right so all of these things really there is no silver bullet I mean you have to really try and fail and see whatever makes sense they said this is how you build good machine learning models yes I I will try to actually find something and if I can find something I will share but it will it will depend upon your domain though right so I mean and say take the example of a telco like tellin alright so what would be some of the features that can cause that can be factored in customer churn can you any good common age of the account maybe if it is up for renewal or not did they make any recent customer service call did they make multiple customer service calls what for the length of the call and do you guys what was the tone of the call where did it look satisfied or not right so really I mean if you look at it I think if you know the domain feature engineering I think if you think hard and you know the domain I think you can do it yes so basically you will if it what is the unit of prediction you're predicting churn at the level of a customer so you will represent each customer as a row in your data set and each customer may be you know the age of the account so two years right so customer one field would be two another one is the median length of customer service calls maybe I'm making it up right the third one was maybe how often in the last week they have called has some one how many has someone there who they call frequently has has that number been switched to a different competitor or not so you will create these columns for each of these customers and then you will start building a predictor model whether they will join or not so that's how you'll do it you might want to actually extract features in time soon time series fashion but this is this may not be actually a time series problem or say time series is more for just more continuous D changing data yes I mean ideally as much as you can practically practically yeah practically a linear algebra is important okay because a lot of a lot of regression modelling is revolves around linear algebra principal component analysis and dimensionality reduction reduction techniques they revolve around algebra so there's probability in stats of course I think that's a no-brainer I think most people understand that and then beyond that I would go for calculus of course I know in calculus is important optimization yeah so I mean but but do you need to know all of this no you can it depends upon what level of sophistication do you want right so to be an expert in academic expert you need to know all of these inside out but to be a practitioner a little bit here and there and that will actually serve the purpose okay okay so I will still be here but thank you thanks for coming over and [Applause] you

Original Description

An overview of how to build a machine learning model from scratch. Modern machine learning libraries make model building look deceptively easy. An unnecessary emphasis (admittedly, annoying to the speaker) on tools like R, Python, SparkML, and techniques like deep learning is prevalent. Relying on tools and techniques while ignoring the fundamentals is the wrong approach to model building. Real-world machine learning requires hard work, discipline, and rigor. The development of robust models requires due diligence during the data acquisition phase and an obsession with data quality. Feature engineering, choice of evaluation metrics, and an understanding of the model bias/variance trade-off are often more important than the choice of tools. Experienced machine learning engineers spend most of their time dealing with data-related issues, model evaluation, and parameter tuning while spending only a fraction of their time in actual model building. This is the 80/20 rule. Unlike most talks these days, this talk is not about deep learning. We will ignore the hype and strictly focus on the fundamentals of building robust machine learning models. Table of Contents: 0:00 Introduction 6:26 Data question 8:44 Metric question 10:51 Repeatability question 13:42 Customer churn 21:48 Business question 33:07 Modern ML libraries 36:59 Lifecycle of an ML model 40:07 Data beats algorithm 48:02 Beg, borrow, and steal 1:05:59 Measuring model performance 1:10:29 Overfitting 1:11:30 Train/test 1:13:54 Cross-validation 1:16:34 Real production data 1:18:30 Questions -- At Data Science Dojo, we believe data science is for everyone. Our data science trainings have been attended by more than 10,000 employees from over 2,500 companies globally, including many leaders in tech like Microsoft, Google, and Facebook. For more information please visit: https://hubs.la/Q01Z-13k0 💼 Learn to build LLM-powered apps in just 40 hours with our Large Language Models bootcamp: https://hubs.la/Q01ZZ
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Data Science Dojo · Data Science Dojo · 0 of 60

← Previous Next →
1 Feature Engineering and Predictive Modeling | Data Analytics with R and Azure ML | Community Webinar
Feature Engineering and Predictive Modeling | Data Analytics with R and Azure ML | Community Webinar
Data Science Dojo
2 Data Exploration and Visualization | Beginning Azure ML | Part 3
Data Exploration and Visualization | Beginning Azure ML | Part 3
Data Science Dojo
3 Reading External Data Sources | Beginning Azure ML | Part 2
Reading External Data Sources | Beginning Azure ML | Part 2
Data Science Dojo
4 Importing Data, Accessing, & Creating a New Experiment | Beginning Azure ML | Part 1
Importing Data, Accessing, & Creating a New Experiment | Beginning Azure ML | Part 1
Data Science Dojo
5 Casting Columns & Renaming Columns | Beginning Azure ML | Part 4
Casting Columns & Renaming Columns | Beginning Azure ML | Part 4
Data Science Dojo
6 Scrub Missing Values & Project Columns | Beginning Azure ML | Part 5
Scrub Missing Values & Project Columns | Beginning Azure ML | Part 5
Data Science Dojo
7 Feature Engineering & R Script | Beginning Azure ML | Part 6
Feature Engineering & R Script | Beginning Azure ML | Part 6
Data Science Dojo
8 Building Your First Model | Beginning Azure ML |  Part 7
Building Your First Model | Beginning Azure ML | Part 7
Data Science Dojo
9 Run and Fine-Tune Multiple Models | Beginning Azure ML | Part 8
Run and Fine-Tune Multiple Models | Beginning Azure ML | Part 8
Data Science Dojo
10 Deploying Your First Predictive Model As a Web Service | Beginning Azure ML | Part 9
Deploying Your First Predictive Model As a Web Service | Beginning Azure ML | Part 9
Data Science Dojo
11 Using R API to Obtain Predictions From Your Web Service Beginning Azure ML | Part 10
Using R API to Obtain Predictions From Your Web Service Beginning Azure ML | Part 10
Data Science Dojo
12 Using Python API to Obtain Predictions From Your Web Service | Beginning Azure ML | Part 11
Using Python API to Obtain Predictions From Your Web Service | Beginning Azure ML | Part 11
Data Science Dojo
13 Twitter Sentiment Analysis | Natural Language Processing | Community Webinar
Twitter Sentiment Analysis | Natural Language Processing | Community Webinar
Data Science Dojo
14 Listening to the Melody of the Universe (LIGO Gravitational Waves Presentation) | Community Webinar
Listening to the Melody of the Universe (LIGO Gravitational Waves Presentation) | Community Webinar
Data Science Dojo
15 David Wechsler on the Impact of Data Science Bootcamp
David Wechsler on the Impact of Data Science Bootcamp
Data Science Dojo
16 Andrew Choi on the Impact of Data Science Bootcamp
Andrew Choi on the Impact of Data Science Bootcamp
Data Science Dojo
17 Microsoft's Software Engineer Shares Her Experience with Data Science Bootcamp
Microsoft's Software Engineer Shares Her Experience with Data Science Bootcamp
Data Science Dojo
18 Michael DAndrea on the Impact of Data Science Bootcamp
Michael DAndrea on the Impact of Data Science Bootcamp
Data Science Dojo
19 Data Driven Decision-Making with Data Science Bootcamp: Artem Kopelev's Revelation
Data Driven Decision-Making with Data Science Bootcamp: Artem Kopelev's Revelation
Data Science Dojo
20 Learn the Fundamentals of Data Science: Srinivas Rao's Experience with Data Science Bootcamp
Learn the Fundamentals of Data Science: Srinivas Rao's Experience with Data Science Bootcamp
Data Science Dojo
21 Re-Learning Data Science with Data Science Bootcamp: Analyst's Revelation
Re-Learning Data Science with Data Science Bootcamp: Analyst's Revelation
Data Science Dojo
22 Scale R to Big Data with Hadoop & Spark | Community Webinar
Scale R to Big Data with Hadoop & Spark | Community Webinar
Data Science Dojo
23 Enhancing Skills with Data Science Bootcamp: Sharon Lane-Getaz's Revelation
Enhancing Skills with Data Science Bootcamp: Sharon Lane-Getaz's Revelation
Data Science Dojo
24 Ryan DeMartino on the Impact of Data Science Bootcamp
Ryan DeMartino on the Impact of Data Science Bootcamp
Data Science Dojo
25 Software Engineer at Microsoft Reveals About His Experience with Data Science Bootcamp
Software Engineer at Microsoft Reveals About His Experience with Data Science Bootcamp
Data Science Dojo
26 Wade Wimer on the Impact of Data Science Bootcamp
Wade Wimer on the Impact of Data Science Bootcamp
Data Science Dojo
27 Analyzing Data with Data Science Bootcamp: Hannah Richta's Revelation
Analyzing Data with Data Science Bootcamp: Hannah Richta's Revelation
Data Science Dojo
28 Applying Data Science Skills to The Current Role with Bootcamp: Marcos Lacayo's Revelation
Applying Data Science Skills to The Current Role with Bootcamp: Marcos Lacayo's Revelation
Data Science Dojo
29 Lance Milner on the Impact of Data Science Bootcamp
Lance Milner on the Impact of Data Science Bootcamp
Data Science Dojo
30 Deloitte's Data Scientist Revelation: Learning Predictive Analytics with Data Science Bootcamp
Deloitte's Data Scientist Revelation: Learning Predictive Analytics with Data Science Bootcamp
Data Science Dojo
31 Rajesh Patil's Experience at Data Science Bootcamp As an Enterprise Architect
Rajesh Patil's Experience at Data Science Bootcamp As an Enterprise Architect
Data Science Dojo
32 Michael Atlin on the Impact of Data Science Bootcamp
Michael Atlin on the Impact of Data Science Bootcamp
Data Science Dojo
33 Amina Tariq's In-Person Experience at Data Science Bootcamp
Amina Tariq's In-Person Experience at Data Science Bootcamp
Data Science Dojo
34 Ceo's Revelation about Data Science Bootcamp
Ceo's Revelation about Data Science Bootcamp
Data Science Dojo
35 Stephen Miller Describes His Experience at Data Science Dojo's Bootcamp
Stephen Miller Describes His Experience at Data Science Dojo's Bootcamp
Data Science Dojo
36 Kevin Hillaker on the Impact of Data Science Bootcamp
Kevin Hillaker on the Impact of Data Science Bootcamp
Data Science Dojo
37 Marko Topalovic's Experience with Data Science Bootcamp
Marko Topalovic's Experience with Data Science Bootcamp
Data Science Dojo
38 Text Analytics With Python, Cognitive Services & PowerBI | Data Analytics | Community Webinar
Text Analytics With Python, Cognitive Services & PowerBI | Data Analytics | Community Webinar
Data Science Dojo
39 Unisys Manager's Revelation: Visualizing Real Time Data with Data Science Bootcamp
Unisys Manager's Revelation: Visualizing Real Time Data with Data Science Bootcamp
Data Science Dojo
40 Learn Data Mining with Data Science Bootcamp: Ryan LaBrie's Revelation
Learn Data Mining with Data Science Bootcamp: Ryan LaBrie's Revelation
Data Science Dojo
41 Vang Xiong on the Impact of Data Science Bootcamp
Vang Xiong on the Impact of Data Science Bootcamp
Data Science Dojo
42 Data Scientist's Experience at Our Data Science Bootcamp
Data Scientist's Experience at Our Data Science Bootcamp
Data Science Dojo
43 Alejandro Wolf Yadlin on the Impact of Data Science Bootcamp
Alejandro Wolf Yadlin on the Impact of Data Science Bootcamp
Data Science Dojo
44 Introduction To Titanic Kaggle Competition | Part 1
Introduction To Titanic Kaggle Competition | Part 1
Data Science Dojo
45 Learning How to Code in R with Data Science Bootcamp: Priscilla Mannuel's Revelation
Learning How to Code in R with Data Science Bootcamp: Priscilla Mannuel's Revelation
Data Science Dojo
46 Andrew Berman On Why Data Science Bootcamp Is Better Fit for Him
Andrew Berman On Why Data Science Bootcamp Is Better Fit for Him
Data Science Dojo
47 How To Do Titanic Kaggle Competition in R | Part 3.1
How To Do Titanic Kaggle Competition in R | Part 3.1
Data Science Dojo
48 How to do the Titanic Kaggle competition in R | Part 3.1
How to do the Titanic Kaggle competition in R | Part 3.1
Data Science Dojo
49 Delve Deeper into Data Science with Data Science Bootcamp
Delve Deeper into Data Science with Data Science Bootcamp
Data Science Dojo
50 Bank of America Data Scientist Reveals His Experience of Data Science Bootcamp
Bank of America Data Scientist Reveals His Experience of Data Science Bootcamp
Data Science Dojo
51 Shaena Montanari on the Impact of Data Science Bootcamp
Shaena Montanari on the Impact of Data Science Bootcamp
Data Science Dojo
52 Types of Sampling | Introduction to Data Mining | Part 12
Types of Sampling | Introduction to Data Mining | Part 12
Data Science Dojo
53 Sampling for Data Selection | Introduction to Data Mining | Part 11
Sampling for Data Selection | Introduction to Data Mining | Part 11
Data Science Dojo
54 Data Aggregation | Introduction to Data Mining | Part 10
Data Aggregation | Introduction to Data Mining | Part 10
Data Science Dojo
55 Data Cleaning | Introduction to Data Mining | Part 9
Data Cleaning | Introduction to Data Mining | Part 9
Data Science Dojo
56 Missing & Duplicated Data | Introduction to Data Mining | Part 8
Missing & Duplicated Data | Introduction to Data Mining | Part 8
Data Science Dojo
57 Data Noise | Introduction to Data Mining | Part 7
Data Noise | Introduction to Data Mining | Part 7
Data Science Dojo
58 Graph and Ordered Data | Introduction to Data Mining | Part 5
Graph and Ordered Data | Introduction to Data Mining | Part 5
Data Science Dojo
59 Document Data & Transaction Data | Introduction to Data Mining | Part 4
Document Data & Transaction Data | Introduction to Data Mining | Part 4
Data Science Dojo
60 Data Quality | Introduction to Data Mining | Part 6
Data Quality | Introduction to Data Mining | Part 6
Data Science Dojo

Related Reads

📰
Qwen2 is here. It’s time to re-evaluate your default model choices.
Explore Qwen2, Alibaba Cloud's new open-source models, as a high-performing alternative to traditional choices for multilingual and long-context tasks
Dev.to · albe_sf
📰
The Brain and Machines: What It Really Means to Say AI Is “Inspired by the Brain”
Discover the true meaning of AI being inspired by the brain and its implications
Medium · AI
📰
The Brain and Machines: What It Really Means to Say AI Is “Inspired by the Brain”
Discover what it means for AI to be inspired by the brain and the limitations of this concept
Medium · Machine Learning
📰
How to Use Chat GPT to Make Money Online (Complete Beginner’s Guide for 2026)
Learn how to leverage ChatGPT to generate online income with this beginner's guide
Medium · AI

Chapters (16)

Introduction
6:26 Data question
8:44 Metric question
10:51 Repeatability question
13:42 Customer churn
21:48 Business question
33:07 Modern ML libraries
36:59 Lifecycle of an ML model
40:07 Data beats algorithm
48:02 Beg, borrow, and steal
1:05:59 Measuring model performance
1:10:29 Overfitting
1:11:30 Train/test
1:13:54 Cross-validation
1:16:34 Real production data
1:18:30 Questions
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →