Live Virtual Mock Interview Of Statistician IIT Kanpur For Data Science

Krish Naik · Intermediate ·🏗️ Systems Design & Architecture ·5y ago

Key Takeaways

Conducts a mock interview for a statistician position at IIT Kanpur, covering data science and related topics

Full Transcript

um okay so i'm sure we are live i'm just going to give you the url of youtube yes so we'll wait for another one two minutes max you know so that people will join and then they will be able to see usually it takes some time for notification now four people have joined are you sharma what is this cocky op hello guys okay now people are joining uh the notification is gone uh so once you can unmute a style you can also unmute okay yeah sure this uh are some of your friends uh pinging in the chat i guess what's the name of the girl yeah yeah they are your friends right so they have a nickname for you [Music] [Laughter] so guys hit like as you're joining i can see many people joining yeah it's gonna be fun today love from another statistician from kolkata amazing hello everyone let's start so okay nice nicknames are coming i think you should not read some messages in short if you open again okay because i i swear they saw right why they are saying like that okay and okay okay okay uh so welcome guys uh we have with us sahil is doing this uh m master of statistics uh in iit kanpur he's in his final year so again this is just a mock interview to he he basically wanted to attend the interview and let's see what all kind of questions we will ask and as usual we will ask questions whatever sahil will say again we have our on the host sudan sukumar again uh today again so there are some people they're saying like krishna you also ask some question always so the answer is asking so if sudan sudans are giving me some time to ask i'll definitely ask but it's okay i can i will be the more like you know coordinator like kind of thing and then she can ask any number of questions right um okay many people are sending you wishes from do you uh sahil okay so let's begin without wasting any time um first of all sahil uh uh introduce about yourself uh you know and what all things you're good at uh about your uh projects that you want to do you know what you're doing some interesting things about your life and many more things yes good sure so hello everyone i am finally a postgraduate student at iit kanpur i did my bachelors from ramallah and college delhi university so this summer i interned with this fox foundation as data science and analytics at that time i was also working on my project my special project i should say which was the solar flare prediction and apart from that i am currently i'm department placement coordinator of department of mathematics and strategics at iit kanpur apart from that i love playing chess i was also the chess captain at my investor's footage okay uh just a second uh silent can you just uh put your headphones near your mouth so that i know yeah it is good uh yeah okay so any interesting thing that you have done in statistics will start with statistics on me any interesting thing that you found out know something new since you're doing your masters probably in your thesis or something like any kind of projects what is what is the final a project that you've selected uh actually in time series model i'm working on time series project which is not yet completed and apart from that and the most interesting project i have done till now is the solar fair prediction project which i have mentioned in my cv okay uh so as soon as the time series when you're saying right so don't show yours are doing like this so he'll start firing your questions so that you please guess proceed okay fine sahil so like uh as you can see right so you have done uh like uh msc statistics you are doing as of now your cgpa is 7.1 cpi right that you have so uh like uh fine so let's uh try to talk about statistics in the first place yeah yeah so uh let's let's start with the very common question right the simple one so generally what happens is whenever we talk about like a sample right yeah so at that point of time uh generally we used to take n minus one in a denominator so if we are talking about let's suppose variances if we are talking about uh let's suppose like a standard deviation right so what is a reason behind that what is the logic behind that one okay so uh basically when we pretty pretty cool common question right yeah we can start from there yeah so when we calculate the sample variance we as you said we usually take n minus 1 in the denominator the reason behind that we say that while calculating the sample variance it is it generally underestimates the variance somehow and this and this has been proved using some experiments so that's why we are reducing the value of the denominator to somehow that it will become closer to the population variance but how you are going to prove it basically like how this one is going to like make this kind of affect so do you have do you know something which can prove these things like okay fine so with the help of like this and this uh like a sample so we can try to prove it uh do you know some algorithm do you know some kind of approach approach the approach behind it i'm not familiar with this but i just know here it somehow reduces uh underestimates the value and that's why we okay have you heard of like a vessel uh correction no no okay i haven't heard it okay okay fine okay uh so uh second question uh for you uh like uh okay so as you were talking about it's n minus one and just to like uh normalize these things okay uh so the second question um for you is again from the statistics like uh generally we we try to like uh use i will say uh central limit theorem right so can you please tell me a situation can you just tell me a scenario where i can apply a central limit theorem effectively and i will be able to get some kind of uh like correct statistical result out of it okay so there are so many things interrelated with that like first what is center limit theorem the central limit theorem is basically suppose we have a big population if we are drawing some samples from them and take the mean and if we take the continuously taking the samples and finding the mean so that sampling distribution of the sample mean basically it will it uh it will become normal distribution it will become normally distributed yeah so this is what central limit theorem says basically yes and that's why the normal distribution is sort of popular okay so in in what in all places so can you please tell me the uses of central limit theorem because definition wise uh like most of us knows the definition of central limit theorem right so if we are going to draw uh like a sufficient amount of sample and if we are going to take a mean of it so for sure a mean of the samples is going to show me a normal distribution curve that is fine but like uh what is the uses so how i am going to utilize uh those things so suppose uh like uh like uh if if i have to apply these things into a data science along with the data so can you please tell me uh some kind of references where i will be able to apply and i will be able to get a significant outcome from central limit theorem yes yeah so we have a concept but for sure like uh we are supposed to apply those concepts we used to study this central limit concept right but where we are going to apply it okay can i get a couple of seconds to think actually i've never thought about it yeah because uh before like uh understanding any kind of a mathematical expression we have to think about it like where we can apply those things right yeah otherwise what is the meaning of like mugging up this entire i'll say like a definition of central limit theorem there is no end of central limit theorem in that case okay fine forget about it so uh let's let's talk about solancho let him think i don't have an issue yeah okay so suppose there are so many statistical tests we can use on on on the on on the data which follows normal distribution somehow if then we can't use those statistical tests so i think uh that's where we can use central limit theorem right what are the outcomes of the statistical test what kind of statistical test do you do like supposing for the parametric test we assume that the the data we are working on should be follow some distribution normally distributed so if the data is not not normally distributed then all many uh tests automatically not in use if the data is not normally distributed so i think uh central limit theorem can work in this field okay what are the outcomes of normal distribution outcomes mean when you say that data is normally distributed what are the outcomes what are the final things you are making out of or assuming out of that specific data hello can you hear me hello hi can you hear okay so uh listen to me suppose if you have a data which is normally distributed yes what are the assumptions do you usually make from that particular data tell me something some so you mean properties of normal distribution yeah assumptions like okay my data is normally distributed so these are the assumptions that i'm considering with respect to that data okay so a large large portion of the data points should lie near the center near the me okay okay and near the mean and there should be symmetric uh symmetric data uh all the very of the values left side equal to the right side okay and these are the properties you are saying but yeah assumption from that data okay you are saying symmetric sign uh okay i'm not getting about from the assumptions assumptions basically means that uh okay i have this data which is normally distributed that i will definitely consider this four points some of the four to five points about that particular data set okay so can you please tell me one point and all right so i'll tell you one point uh all my data will be most of the more than 96 percentage of my data will be ranging between three standard deviations okay oh so this okay so that's the rule about 68 95 99 now you are saying it so you understand no you cannot ask an interviewer you tell me what sorry i'm sorry sorry about that okay so fine uh one more question for you right so like uh let's suppose if if i have like a data is given to me right and uh statistically if someone is telling me that okay fine so let's uh do some kind of a statistical analysis and uh let's try to find out anomalies inside the data right so like uh what will be my approach in that case so can you please list down uh like a five to six uh like approach different different kind of things what we mean so outliers you can say yeah so outliers if i have to find out that okay fine so like what is available to you science yeah yeah yeah so let's suppose if like uh some data set is given to me and uh some person is trying to ask me that okay fine so let's do a statistical analysis and with the help of a statistical analysis only with the help of statistical analysis we have to find out that okay fine so whether we have like a outliers or we have like uh some uh kind of a data which is not acceptable one so what will be your approach and what is the approach that you are going to take to make those data as a normally distributed data set so again so like a approach to find out like outliers as well as approach to handle those outliers so what will be like your test what kind of a test one test we can use is z test okay so z test basically suppose for all the data points we subtract the mean of all the data points and divide by standard deviation and then we see whether the data that the value is greater than three or less than three as crystal said if it is less than greater than three so it means it is it doesn't lie within the third standard deviation of the data that means a 99 point it is not one from the 99.7 three percent of the data that you are talking about actually so z is equal to data minus mu divided by a standard deviation so basically formula that is statistics we can use to like a sift us data or maybe i can say that to like uh change a data set from a normal distribution to a standard uh like a normal distribution we are finding whether the data is uh what is outside the range of what is okay that that is fine so even without converting this data i will be able to do it right so we can we can find uh iq by iqr method the first approach that you talked about right so basically even if i'm not going to take that approach that is fine because i'm not doing anything except changing a scale of the data right yeah yeah okay tell me yeah so what are the other approach that you can think of to finding the outlier right yes yeah so i have to like find out layers uh inside the data set and then i have to handle it so approach to find out and approach to handle the outliers hello yeah yes i'll go sorry for the next selection so uh iq is wrong uh like i said about the iq so i it's not okay so what about we can for finding the outliers we can use box plot or iq method or z-score there are three we can use okay anything else that you can think of in terms of statistics okay apart from iqr z-score and box plot yes yeah they're for statistical technique uh i can't go further than these skewness yes test data can be positively scored negatively skewed and maybe those outliers is significant so that is the case okay okay and uh so fine so you have listed down couple of the approach by which i will be able to find it now tell me what is the approach to handle these outliers what i should do so when i will be able to get the outliers so what i am supposed to apply on top of that we can remove those outliers how like you are saying that like i am supposed to discard those data this is what you see yes so do you think that like in every cases this is something which is going to work so simple discarding the data so am i not going to lose a like i'm not going to lose like a the data set which may represent some kind of for relations yes if the data is important then we shouldn't discard those outliers those outliers will be important so yeah so just tell me some of the approach where i'll not have to like discard the data but still i will be able to handle it yeah without discarding how will you handle that okay so suppose like i am supposing that the data is skewed and that's why we are having an outlier suppose in a salary suppose the person's salary of a person so it can be positively student there can be outlier which is significant yes i'm saying that how will you treat so the answer question is that how will how you will treat that outliers okay just listen to the question again so the answer question is that how will you treat the outliers without removing it this is a simple question okay okay so we can change the data so we can normalize the data yes normalization is one technique and somewhat actually this time i can only think of normalization yeah so science just tell me one thing so when you say normalization right so can you please give me a particular example by which you can explain me a meaning of normalization so what do you understand by exact term normalization because like normalization term is a big term so many people like blindly just use a term called as normalization right so what is your understanding about this term normalization i'm just looking for some example and with that like like in celery case okay we have the salary of a person of uh suppose we have 100 data points and that's that data point trigger and salary okay so if the data sorry i forget your question again please sorry i'm just asking that what is uh like um what is your understanding it's a normalization yeah sorry so uh suppose we have data point which is whose range is so big and we want to put the data in a certain range suppose like in normalization we put the data between zero and one so putting the data and then what is the difference between normalization and standardization so standardization in in standardization we basically say that our data will follow uh have me hello yeah so data will have mean zero and variance one that uh both techniques can be used in separate cases so in suppose normalization as i said if you want our data to be in range suppose you and one in a particular range we use normalization and standardization if our data points has at different scales standardization or data point is in different scale i think you should recheck those definitions that you are talking about actually fine moving ahead so like uh let's let's talk about uh i'll say like a difference between uh so again a very simple question a very common one right so now like uh what is uh so difference between a jet statistics and a t statistics and in which case which one i'm supposed to use and which one i'm not supposed to use actually they both do the same work but the difference between where we can use what if if the data points is large suppose the data point is and data points and n is greater than 30 so basically we can use the test if the data is greater than 30 and t test if the data is less than 30 we can always use t test in this case okay so any other differences sorry i can't think of any other difference can you please talk about like okay fine so in terms of number of sample you have given this definition but can you please talk about the difference between z and t in terms of uh standard deviations uh like let's suppose if i'm giving you standard deviations of populations or something like that so can you please talk about that can you please try to compare this z and t in terms of a standard do you have any idea about that fine in terms of uh like a sample like a number of the samples so yes people used to like compare these things lesser than 30 or greater than 30. so in this way this is a very normal comparison right so like do you understand anything in terms of standard deviations how we can like differentiate these two in terms of standard deviation yes yeah i think in t test we calculate the there are different kinds of data like a pair t test for pair of means your two pairs of samples and in standard deviation i think we use the calculator to the uh standard deviation of temple i think i mean uh sorry i hello okay okay yeah go good feel free to like to answer yeah hello it's it's always good to not say anything i can't recall properly okay okay that's fine that's fine it's it's like always fine to like accept that uh you can you can revise and then you can come back okay so that is uh like uh one thing uh from the statistics now um i'll again i'll ask couple of more questions from maybe a puzzles and statistics as well but uh let's let's try to understand your uh machine learning uh concept yeah so generally in a machine learning right so now let's suppose if i'm talking about a regression so like we used to consider a root mean squared error yeah squared error so why like a root mean square is uh sometime called as a worst uh i'll say like uh calculations error calculations uh in terms of regressions any idea so rms is called its first scenario yeah so basically it's called as a worst error calculation sometimes in in case of regression so why do you have any idea oh okay so i can try like in my project in time series project i forecasted some values and calculate the rmse for the test set and for the forecasted values okay but the data under consideration was the gold prices okay so gold prices can vary suppose today's gold price is 32 000 and tomorrow's gold price will be 34 000 or 32 000 so in forecasting there will be some difference uh in the forecasting and the test cases so if we calculate the rmse there will not be there will not be any certain range or that we should consider that it is good or that is bad so i think that might be the case okay script is again like uh give me some other example and then try to explain me the same thing okay so i said suppose uh the data under consideration plays important role if our data is uh suppose 0 to 100 or 0 to 5 vary between 0 to 5 another another data is 0 to 10 zero to ten thousand to zero or ten thousand to one lakh okay and then we predicted our value we predict our values in both the scenarios the rfc will vary a lot and they signify for the same thing suppose in first case i am having a rmse around two or five and in the second case i am having rmse of 1500 or thousand okay we can't say looking at the rmse whether it is good or bad okay yeah so in that case what i'm supposed to do so what is the probable solutions for like this particular situations where you have a one data variation in some other skill and then another data variation in some different scale so what what are the solutions what i should do in that case so basically we can feature our skills oh sorry sorry what you guys feeling features can you just talking about right yeah okay so maybe i can try to talk about the standardizations or maybe like so-called normalizations that that area uh you just has a very common word right yeah okay fine now so the next question is uh like uh let's suppose if i'm if i'm talking about maybe uh logistic right okay so can you please like i talk about uh like a error functions or loss function in logistic so i have used a square red log flow function uh which one was squared error loss function okay it's quite a loss function into a logic it is logistic not in logistic it can be used right logistic is a classification yeah sorry sorry sorry it can't be using logistic okay so yeah what is what is the like a loss function error function that you have used in logistic okay yeah i know just give me a second to think yeah so loss function can be so the predicted value can be the predicted value minus suppose our logistic regression is we use sigmoid i think sigmoid function and the output of it i can't recall sorry i think it should be it must be some kind of a probability function right because we are trying to like talk about the classifications over here and for sure so we are supposed to find out that okay fine so like uh whether classes it is able to predict correct or incorrect right so basically we are supposed to like think about that okay fine so we can use the different accuracy measures like accuracy prediction species and recall for those purposes but but that is a accuracy metric says right i was talking about the error function i was talking about the loss function actually so there is a difference between accuracy right right precision recall f1 is cool so these are the things which i can use for uh finding out accuracy not for not as error function hope this is clear right yeah yeah okay uh now so like uh okay fine just tell me what is your favorite algorithm just tell me about that algorithm um machine learning deep learning whatever whatever you like i like with your best okay so let's go with random forest everyone goes with the random words i have seen it's okay it's okay we can also we can also go with the logistic regression linear regression or you know again okay so suppose if i'm a kid right so let's suppose i i don't know like uh i'm just trying to like uh learn this uh entire logistics or linear okay and someone is telling me that okay fine so this logistic or linear basically what it does is like in case of regression it will be able to give you some number in case of i'll say classification if i'm going to use logistics right so it will try to like separate a certain classes fine okay now so if you have to like a convince me if you have to like uh teach me that okay fine so like uh this is how uh like a prediction happens into a linear like a linear eq linear regression and this is how like classification happens in case of a logistic how you are going to explain the string to me so your basic question i just have one idea i i just know about the couple of things like okay finds a prediction for casting we can do it right so prediction of forecasting like uh is possible if i'm going to use or classification is like a possible right but you have to convince me you have to explain me maybe like within two to three minutes right so what will be your approach in that case how does the prediction happen yeah basically like how it is able to forecast something so okay what it does what it does internally by which it will be able to do give me some kind of a classes in case of logistics or in values in case of regulation okay so suppose we are working with univariate model we have only one independent variable independent variable and one independent variable and we want to predict whether it will be zero or one plus so binary classifier okay so in in this case [Music] how prediction happens okay so we will take uh it's starting from a equation suppose beta naught plus beta 1 x plus beta 2 x square so i asked you to explain these things through a layman right okay so we have certain suppose we have a feature uh by a feature i mean we have certain values suppose x1 x2 xn x up to x n see i'm looking for a story over here right i'm not looking for your mathematical equation i'm not in a position to understand your mathematical equation okay so in by prediction we mean what we have some used case in the previous time and from the used case we know that if that situation happens we have that output if that situation happens we have that output and using those scenarios or those data points we are predicting something which happens with your own answer and not really then then how do you think that i will be like convinced and i will be able to understand it okay fine you said that like a logistics so we can use a binary uh like we can use for binary classification we can use logistics right generally it's been said that uh for a multi-class classification we are not supposed to use logistics although we can use it right so but uh it's not an ideal uh like i'll say algorithm a logistic regression is not an ideal equation or ideal i'll say algorithm uh to use in case of a multi-class classification for binary that is fine so why and what is the reason behind that actually i haven't done any use case in multi-class classification i am only familiar with binary classification is it okay so do you know that that multi-class classification is possible with the help of logistics yeah it should be possible in entrance against uh a course i did in december january and i think i have read about it that it can be used in multi-class classification okay so then my question was like why it is said that you are not supposed to use logistics in case of multi-class for binary that is fine but for multi we are not supposed to use it any thoughts about it i think the output of the very the output in multi-class specific in logistic between probability zero to one and suppose we have uh suppose k classes or three classes four classes then i think the output will be 0.3 and 0.1 for each case which will give the better in logistics have you heard of one versus all ovr one versus all one versus rest no okay not a problem i i just have one question can i ask yes sure please go ahead yeah so one question uh what i see in your project is that you have to work on finance right so uh i just sorry not finance uh you have worked on some time series and forecasting just tell me the type of train test split how do you do the train test in this kind of problem statement so basically we can't use random splitting in time series so we have to take the for test case we have to take the last values so that's the only possibility but why just tell me why why do you go ahead with that approach so suppose if we take random data points no not about random i understood about random yeah about this why do you take the last one as a test yeah so okay go ahead go ahead with uh what you are saying yes go ahead so i was thinking suppose okay we can't use random observations because because the data points of certain time will be missed and we can't be able we can't be able to predict the future future data points so that's why we take only the last data points or we can also take the uh so the test cases we want we want the data for test cases should be in sequence sequentially arranged i think uh why sequential energy not only i'm asking why why sequentially arranged you you are telling it right okay suppose i have data set from 1950 to 2020 suppose from 1950 to 2000 i have taken as my training data set and 2020 to 2015 or 2020 i'm trying to take as a test data but why why do we go in that way why not randomly because in the then they will be missing values in the data and we don't have any information like for uh in times this we take uh a our model we rather see it is a r or m a so what is the parameter so we can't be able to uh relate whether that whether the value of the given time will be ready to the direct effect or the indirect effect of that previous time that information will be missed out yes so whenever you say this kind of answers is that whatever technique you you are using right with respect to time series let it be rnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnn nnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnnn or arima model or sarimax anything what it requires it requires sequences of data so suppose if you have your data you the training data set you have to consider as a sequences you cannot just randomly pick up somewhere similarly test data will be that but now tell me will there be too much impact if the data movement is too much with respect to time series will you have a very mod big model impact can you please repeat the question suppose i have from 1950s to 1970s i have some data okay price of gold suppose or just forget about this year only jan to december price of gold let's see yeah okay initially it was very less but suddenly it got a spike at one point of time now when we are doing train test split for that particular kind of data will there will be a huge model impact because now here the sequence of data has lot of variations yes so how do you fix that part sorry i don't okay not a problem so uh i'm just seeing your resume uh you have actually written or you have applied smart okay let's go with clustering algorithms you have written k-means clustering how do you decide the number of clusters uh actually the data point the data the data which i am working on was the polar fair prediction where the zero class represents there will be no emission of solar flares and one represent there will be addition of solar please so the class of zero will be you already know there are two classes you already know that yeah because i'm considering it uh whether solar flare will be emitted or not okay so you are solving a supervised machine learning problem statement initially you read a clustering and you know that okay it is belonging to two groups this is what you have done actually so clustering i used as a technique to under sample the data under sample the data yeah what do you mean by under sample the data in that case how did you understand like the one one class okay if you understand the data what is the disadvantage a loss of information then why did you do that so actually the thing is that was my first project and why i applied this is they were i have two options whether i go for over sampling the data but if i go for over sampling the data i had only 100 100 data points of one class and if i uh suppose do this site will increase the noise in the data so i have read somewhere that it will be better or sampling will increase the noise who told you that actually i used to read blog that time and i read that what does what does over sampling do just tell me what does over sampling do it will increase the synthetic data points synthetic data points why you are saying it as noise then noise is different synthetic data points is different so that i read that instead of using only under sampling or over sampling it will be better if you use both or a little over sampling and little under sampling that what do you mean so you are saying that you have used smart technique for that right yeah for over sampling okay have you tried anything like class weights what algorithm did you what algorithm did you apply for that particular problem statement after clustering uh forward see you did clustering algorithm right yeah you said which algorithm did you apply for supervised i'm saying um for training the model for training trained logistic regulations yeah logistics for training i use logistics logistics okay did he not try with any other algorithms are there any algorithms that will not get impacted by imbalanced dataset not tell me this imbalance data uh if we can go with hello yeah yeah i can hear you go ahead please hello yeah yeah go ahead uh we can hear you i think if i go with random forest then as the random forest takes you know bootstrap samples so it will it will be the case that the class of one class samples will not be included in the data so there might be a case i think okay so i like uh uh fine uh so i think uh let's let's talk about some of the scenarios situation right so as i can see in your resume that you are like our department placement coordinator student placement officer spo iit kanpur right yeah this is what you have mentioned so august 2020 ongoing yeah so you have basically planned and executed a placement drive 2020 uh in a three-tier team of 1200 plus student yeah so this is what you have uh mentioned so basically you are a placement coordinator and you used to maintain manage on uh do all of those things so i have a problem statement for you actually right based on the same work that you used to do right so you are basically a placement coordinator uh that is uh fine but you are facing a challenges actually so there are like uh many companies which used to visit your uh for sure like a colleges right there are thousands of companies who used to visit your colleges and uh every company is having their own like a set of the requirements so some people are looking for a candidate who is a very good into a statistics right some companies are looking for a candidate who is very good in the programming uh some companies are looking for like uh i'll say a candidate who is who is very good in terms of like a research paper or somewhat like that right or maybe a solution designing and different different like expectation they have basically right and as a placement coordinator right so you get stuck in between so you get stuck in between a company and a student so there are let's suppose like uh there are thousands and thousands of resumes you have from one side right you are the placement coordinator and on other side so there is a company right so you have to map or you have to align a correct resume to a company so let's suppose like uh if i'm looking for a guy who is very good in terms of algorithm a guy who is very good in terms of and again inside the algorithm so we have a different different thing algorithm uh can belongs from our data science as well as data algorithm right from the company we have we have a category right we have our two category right so your job is to provide so your job is to build a system actually right your job is to build a system so where people will be able to upload a resume right everyone will be able to just dump the resume right so you can't give a pain to a student that okay fine so like a create or fill a form you can't ask those things right that is your problem so a student will be able to dump a resume so we have a thousands of resume in a single drive okay we have a thousand resumes a single drive right now a problem with you is how you are going to align those thousands of resumes to which company you can't allow all the thousands uh like a resume right all the thousand candidates to sit with one company so i have to select the resume which is relevant to that company irrelevant to the companies like relevant to a company who is looking for a particular skill yeah i'll say so relevant to the requirement so company used to publish a jd job description right let's suppose right that i'm looking for this is this kind of a candidate right so you have to map so that company will not have to waste even a single second of time right to in do a interview or get a candidate who is not uh like a required for their particular requirement yeah so how you are going to build this entire system so i think that is a problem of something nlp type with uh let's suppose uh like uh let's suppose uh you don't have an nlp let's suppose you just know machine learning and maybe just based on the machine learning you have to solve it so how you will do it in that case okay forget about nlp i don't know about nlp solution so in this in this problem we are mainly interested in selecting or extracting the resumes and based on the information by information we are extracting those resumes and sending to the company yeah so we are doing something we are doing something we are trying to apply some kind of a machine learning technique some kind of maybe a statistical technique we are going to apply right but i have to solve this problem just based out of machine learning or statistical technique because as you are the student of the statistics and as you just know machine learning so you are not supposed to use anything else no deep learning nothing no nlp nothing you are supposed to use just a clarifying question that uh in this the objective of the problem is to select the resumes yeah allowing a resume to a company's requirement okay allowing a resume so that is a problem statement right and you have to design a system so that as a company right so if i'll come to your college right so i'll just go and write down my requirement okay fine so i'm looking for this kind of this kind of this kind of connect must be having this skill this skill the skill i'll pull my resume i'll be able to get like i can resume so we can do one thing as i'm thinking so we can create features suppose there are requirement of suppose machine learning or something designer there are some separate words or we can take it as a feature and the information contained in the resume i can take number suppose in first resume it is uh it is about machine learning projects so i can rate the feature suppose one two three four five six seven eight nine ten scale basically 10 means they are more relevant to that specific thing or one means they are very less relevant to that so we can uh make features suppose 10 features for 10 different profiles and based on that we can use solid using regression problem so we can again again you are going in one direction see as a company i'll come to your college i'll just give you a job description agree okay right i'll give you job descriptions now let's try to simplify this problem statement even on a different level right so i'll just give you a jd job descriptions right in that job descriptions uh like uh whatever is my requirement that record will be mentioned right from a company side right from a student side right from a student side so you have a resume and every resume like all the stories are mentioned about their details about project and about the technical skill everything is mentioned right now tell me like how i am going to interface uh let's suppose there is 10 company right there are 10 company and there are thousands resume right so how i am going to distribute this thousand resume to this 10 company according to their requirement do you have any idea any approach anything that you can think of in in real time so the previous solution i provided that was that was wrong yeah basically like uh that will create a huge complexity and you will not be able to align uh like a particular skill set with a particular like a requirement this is what i think okay so so we can use classification problem where the output feature is the number of companies suppose 10 companies and the variable variable is a student suppose data point is a student and it will the output will be whether it will it will be in the first company or it will go in the second company or third company like that okay do you think of so how like a logically can you can you improve these things that okay fine so effectively at least with the 80 percent accuracy i will be able to or even with the 50 percent accuracy i will be able to align those resume what classification algorithm you are going to use how you are going to classify a text which is available inside the resume how you are going to understand a word which i have given you inside a job descriptions how you are going to match both of these things because these are the problems that is going to come into picture right yeah and we are assuming that there is no case of nlp right no i'm like i don't even know nlp right so how i can apply nlp simple i just know statuses or machine learning okay so we can use decision tree right i think in this problem and one two minute time okay so first we can use this entry and in the first split we can split the data whether it will be core profile or non non core profile no yeah yeah so first note will be whether it will not core profile or not profile second split will be based on the broader category suppose uh you know analyst job type of data science job or the hardware type job something like that and then we will raise to the leaf not something and leave not will be the company so where the student will go from how you are going to first of all create a classes can you please explain me that particular part because you are talking about like a decision tree and it's a classifier basically it's keep on like creating a branches based on the values that you have given right so how you are going to decide that particular part here please explain okay so because i've just given you resume i have not given you any kind of classes so you have to work on your uh data like a preparation part as well so how what what is what do you think that how you're going to work on that part actually i'm confused how can we extract the information from the jd it's fine so let's suppose i'm i've given you all those things in a word format right so simple you can use a word reader and then you can just extract it so or else even if i've given you all the jd into a pdf format so fine there is a like a python library which is available for the pdf reader so you can read it and then extract all the information so suppose you are able to extract it this is the assumption this is simple uh like a part of programming i'll say so you will write couple of line of code and then you will be able to extract it now what now we can like we can use in decision tree we can use first feature like suppose one feature is there is a class where is a class so do you think that like a thousand like a student's resume you are going to consider as a thousand classes this actually works this entry works i agree with that but where is the class i have not given you the classes at all i just given you a number resume you have you are able to extract all the information from the resume where is the class you are saying that the decision will apply so placement coordinator right so you have to solve this problem so that uh like you can you can work effectively you're not supposed to like run after a student and a company again and again and again and again you have you don't have to involve someone just one system and your problem is solved this is what i'm looking for okay okay so one last solution i'll try suppose uh we are considering features first okay for the data preparation first feature is profile i as you see i can i said i can extract the information it is a floor profile on code profile so first feature is categorical variable which contains two two classes first is core and non-core yeah please go and do it i'm just i'm trying to interpret your solution yeah yeah so so second feature second feature contains another category whether the person will whether it is a job of something hardware related or software related the third feature will contain the information whether the classes or sorry categorical uh variables like uh whether it be a data and list data scientist or something like that and use we can see for those categorical variables we can use okay so this is it from my end so hope you have enjoyed case studies the solution give the solution let him let him think of right so he's a placement coordinator so for sure he's even last time i've given the solution right so whatever question i've asked i'll just try to talk about these things in the next session so yeah he's a placement coordinator right let him think about it in just in terms of machine learning the restriction is uh you're not supposed to use an lp you're not supposed to use deep learning anything else just a machine learning i'm not sure this is a very interesting question yes think through it as you have already mentioned in your resume that you are a placement coordinator so for sure you must have faced this problem and you have to solve it yeah so fine question i'm just done uh uh like uh yeah so if you haven't should i know that that part was right or not in the future one of the approach basically so uh like uh i'll say like yeah partially it was right and again if you can explain me the entire approach then uh we can think there can be like a four to five approach it's not like there is only one of us that you were talking about right so yes one of those approach i think uh you were trying to explain because yeah that is fine so in this way also we can uh like i try to achieve uh like some sort of i'll say like alignment of the resume okay okay one question from my side uh i think suraj you told him to think about it right yeah yeah so like he can you can think about it so like like maybe in a next interview session okay so i will be like uh talking about solutions and okay i'll i'll i'll i'll ask one question uh final question from my side side you perform one hot i see i'll ask very simple questions don't uh you know so you you're written in your resume you have done one not encoding to choice to change category features into numerical variables right yes using dummy what if you have a category like pin code okay how do you handle that okay so if i use one hot encoding this in that feature there will be so many classes and using those dummy variables will be bad then how do you solve that can we use label encoding which i think we can use label encoding that by default it has label no pin code has a unique label only 92 percent accuracy i have categorical so many have people have written the right answer in the chat yeah i can see that yeah [Laughter] think of it okay just tell me why do we convert category features into numerical features because our model is unable to interpret them because they are num they are categories like linear regression is unable to interpret categorical categorical features so if you make it as one zero zero zero one zero zero one zero then model will be able to interpret if there are categories then suppose there are two classes then it will model will interpret okay like that you are thinking right yeah okay think over the pin code thing i'll give you two minutes sudan should we be singing suppose if i take in bangalore you'll be having 10 000 different different based on location okay can we can we use suppose we have different pin codes for different cities and can we use if we have support cities which are nearby to each other then we can cluster them and use as a single variable something like that you are grouping them you are grouping them i'm saying handling category features you're grouping them you're applying some clustering algorithm and you're doing some of the things right i told that how you're handling like this right by using dummy variables or like that so how you're doing it again i'm just checking your resume i'll find some more things right till then you answer this question okay now you have written in bold one whenever you write like this in bold no that basically means you're good at the thing right and i was prepared for this that question but not this question see people are saying if you follow krishna except this video you will get the answer easily i follow but it's around only two to three months okay so okay you're saying that you've built a neural network model also right to classify images yeah that coursera project how did it decide like how many number of neural networks or how many number of layers and hidden neurons needs to be used if we use many hidden neurons in a particular layer it will overfit the data so now how did you decide like how many number of layers or how many number of hidden neurons you have to use we can use uh parameter tuning suppose we can use different parameters which parameters we can use k fold close relation and suppose i am using k for course validation and for each oh it does this is why k fold cross validation is used just tell me that okay so in k fold close validation suppose i have k value of 10 and we have i am think i am taking 10 different values of neurons in each layer careful cross validation what does cross revalidation basically say hello yeah just uh are you able to hear me yeah what is cross validation cos radiation is a fit available which will which will have the better accuracy okay which okay how do you select the internal parameters like how many number of layers how many number of neurons i'm not talking about input whenever you have your input data in cross validation you'll be able to divide that into train and test split yeah based on k folds k number of folds okay yeah sorry i'm using actually actually i'm new in you know deep neural network things that's why i see your resume i know it is written i told you you know that was the online guided project that's why i am not i have not had a complete understanding okay so that's true do you want to ask anything uh no i think uh like uh now i'm done for like today yeah so uh maybe we can talk about the solutions uh so are you able to think of any kind of solutions for the problem that i've given to you i think you have heard for the first time if you are going with coursera you may have not heard i don't know because we are using that in real world application so keras tuner is the answer from that okay so with the help of keras tuner you can say that how many number of layers you can use what should be the range of number of neurons in the hidden layer and many things okay so yeah like a question so like now are you able to think of some approach any any approach to build that particular solution okay so i'll try yeah so hello yeah so last time so after that i'll try to give my answer yeah hello hello yeah okay so m first can i assume that there are only ten companies okay uh yeah yeah you can assume ten companies okay so i'm assuming there are only ten companies and i am here using decision tree fine okay so first i am in first node i am calculating suppose skin impurity to check the split to check whether to check the node which i prefer at the first place and suppose the impurity says that i should prefer first node that says whether the profile is of cour profile or non profile okay okay then likewise i i will further going deep into that and on the leaf node i will get the companies so whether which resume suppose i am inputted resume and which resume fall into which company okay so is that correct but so generally in a decision tree so on a leaf node so we used to keep our classes right this is what it used to do so like do you think that the resume which you are going to get so in that label of the companies will be mentioned so how do you think that i am assuming that the output will be name of the company whether it is company a company b company c okay so in that case you have to label a first of all a data right yeah in dc entry so let's suppose whenever you try to build a distance can i say that you will be having x as well as y yeah yeah so where is your y so you are saying that now output will be y so at the time of like training your model so from where you will be able to get y so how you will be able to get y at the time of testing of the model okay so have you heard of like a distance based approach no distance so maybe like i clean distance hammering distance that's kind of a resistance management distance so can you think of anything based on that yeah i have no i haven't included yeah i have a resume i have a jd so now can you think of like any of the like approach based on that there's nobody maybe cosine similarity manhattan distance hamming distance curating distance right so can you please think of any of the approach okay so yeah what this distance does actually so can i say that like uh so it is possible to find out a distance between even a word or even a sentences agree even sorry even a word or even a sentences if i'll talk about a hamming distance what hamming distance does so can i say that hamming distance will try to check of what similarity and then it will give you a binary result one or zero actually i'm pretty family i'm familiar with only manhattan distance and including distance okay so just a numerical based distances yeah so there is something called as a hamming distance there is something called the cosine similarities there are many like a distance based approach that you will be able to find out even in statistics or even in machine learning right so the only thing that you have to find out over here is so first of all whenever you are reading a data and then storing a data of a resume this is one of the approach i'll tell you there are like i said so there are five to six other approaches also so clustering based approach classification based approach distance based approach right so there are many solutions that you will be able to design so let me talk about one approach which is a distance based approach so here what you can do is so first of all you can pass the entire resume which is just a couple of line of code you can take any language right so after that so you can try to filter out uh maybe you can try to like filter out all the like a words which is going to occur again and again right so just you you need some some of the words maybe you can say that okay fine i'm removing your stop words and all the common words all the punctuation and everything i'm just trying to keep a word which is important important or maybe like which is not repeating itself again and again a kind of a unique word i'm trying to keep right inside my resume again so company has given you a jd right so whatever so here is your resume and from other sites so you are able to see what you are able to receive a jd right you will do a same thing for jd right so now you have a jd right and you have a different different different different i'll say a resume which you have already passed in your databases so can i say that one jd versus all i will try to map so one jd will try to find a distance with every resume right and for sure it will try to give you some of some sort of the distances at that point of a time so whichever is giving you a lesser distance so can i say that that that particular resume content is closer to the jd right so now if your interview is asking you that okay fine so this is the job description now i'm looking for a five resume right so if you are a distance between five yeah so first five first five you will try to select which is having a lesser distance right and then you will be or maybe based on the probability you will be able to do that so similarly cluster based approach is also possible classification based approach is also possible to solve this particular problem by which like uh this again i'm not talking about the accuracy of this particular one right so uh like in terms of our distances or maybe classification based or maybe clustering based but yeah so these are the things which is possible like mathematically and you will be able to do it yeah okay so uh fine this is one of the like a possible answer right now you can think of like a cluster based and classification waste but classification best answer that you were trying to give right so that is not uh correct okay uh coming to okay coming to when you're good at something never do it for free right you understand that okay so uh we'll come to the feedback uh i like your amount of knowledge that you have uh still your basics needs to be much more strengthened uh because uh whatever things you have written in a resume it will be revolving around that because whatever questions we asked you was based on your resume right so when you are writing something with bold character make sure that that surrounding topics needs to be covered well again uh it's okay uh sometimes what happens is that yes in interview they may ask you some more additional questions but the thing was that you did good well you tried to handle some of the things yes we forced you to take out the answer from your brain and uh come out you know that all things should be in the tips of the lift right so yes the interview was good uh i'd suggest but still your basics needs to get strengthened and whatever things that you do only write that thing which you're good at okay which you are good at try to write that thing over there okay and whenever you are saying something about some machine learning algorithm make sure that you learn about that machine learning algorithm properly because when when i told that logistic regression is used for multi-class classification or not have you heard of ovr one versus rest so these all kind of examples are some of the common examples to test your mind the interview will try to test you and since you're a statistician related to statistics more concepts will be asked to you okay yeah so think of like a solving uh see like everyone learns to uh like uh solve a problem of machine learning based on some kind of i'll say a toy data set right which is widely available across the internet right okay many people are asking about the pin code okay so guys pin code you have various techniques like target mean encoding you know mean encoding so many people had written the right answer before just check my future engineering playlist everything has been explained yes sir so that's good yeah so basically see what happens is like uh everyone right so whoever try to practice machine learning and they try to start a machine learning so generally like they start with the toy data set that is fine that is a good start i'll say but unless until you are not able to relate those things to a real world right so i'll say like uh your learning is not complete and you will not be able to get experience always and again so many people used to say that i don't have a data set right if they can start thinking about a real world problem right you will be able to create a data set even by yourself you don't need someone to give you a real-time data set in many situations i'm not talking about the core domain right i'm not talking about the core banking domain or some some supply chain domain or maybe a retail domain right but there are many problems that you can think of in even in your real time and for which you can generate even a real-time data and then you can try to utilize it so i think that is our best way to learn to design a solution right so think about again that there will be like there can be a multiple number of solutions so always like a try to learn in that way uh anyhow like uh yeah you're doing very good right so hoping for the best and uh yeah so for sure like uh like uh you you will be able to do better in your like a near future yeah so yeah this is it from my side and uh thank you so much for coming to this platform and i'll see you again yeah yeah thank you so uh one last thing uh just the second one last thing uh keep on going on with this you're still in your final year you have a lot of chances make your basics strong and uh yes we can have one more interview after one month just drop me a mail again you know uh you have my number also so you can message me at any point of time and then what we'll do is that will try to conduct after one month again interview at that time whatever things you write in your resume based on that will revolve around the question and let's see the difference okay now for every interview that we take first we'll take the first round and probably the next round in after one month then you try to see that how much improvement is there in you and that will also make you satisfied okay so thank you thank you science thank you sudanshu our wonderful host uh guys please uh sudanshu opie should be there in message uh everybody sudanshu opie sudanshu op sudanshu opi again a wonderful last time last time chris told me about the meaning of op yes now i know like what is the full form so thank you thank you guys thank you sudanshu and

Original Description

Sahil LinkedIn:linkedin.com/in/sahil-saini-3a420217a ⭐ Kite is a free AI-powered coding assistant that will help you code faster and smarter. The Kite plugin integrates with all the top editors and IDEs to give you smart completions and documentation while you’re typing. I've been using Kite for a few months and I love it! https://www.kite.com/get-kite/?utm_medium=referral&utm_source=youtube&utm_campaign=krishnaik&utm_content=description-only All Playlist In My channel Interview Playlist: https://www.youtube.com/playlist?list=PLZoTAELRMXVM0zN0cgJrfT6TK2ypCpQdY Complete DL Playlist: https://www.youtube.com/watch?v=9jA0KjS7V_c&list=PLZoTAELRMXVPGU70ZGsckrMdr0FteeRUi Julia Playlist: https://www.youtube.com/watch?v=Bxp1YFA6M4s&list=PLZoTAELRMXVPJwtjTo2Y6LkuuYK0FT4Q- Complete ML Playlist :https://www.youtube.com/playlist?list=PLZoTAELRMXVPBTrWtJkn3wWQxZkmTXGwe Complete NLP Playlist:https://www.youtube.com/playlist?list=PLZoTAELRMXVMdJ5sqbCK2LiM0HhQVWNzm Docker End To End Implementation: https://www.youtube.com/playlist?list=PLZoTAELRMXVNKtpy0U_Mx9N26w8n0hIbs Live stream Playlist: https://www.youtube.com/playlist?list=PLZoTAELRMXVNxYFq_9MuiUdn2YnlFqmMK Machine Learning Pipelines: https://www.youtube.com/playlist?list=PLZoTAELRMXVNKtpy0U_Mx9N26w8n0hIbs Pytorch Playlist: https://www.youtube.com/playlist?list=PLZoTAELRMXVNxYFq_9MuiUdn2YnlFqmMK Feature Engineering :https://www.youtube.com/playlist?list=PLZoTAELRMXVPwYGE2PXD3x0bfKnR0cJjN Live Projects :https://www.youtube.com/playlist?list=PLZoTAELRMXVOFnfSwkB_uyr4FT-327noK Kaggle competition :https://www.youtube.com/playlist?list=PLZoTAELRMXVPiKOxbwaniXjHJ02bdkLWy Mongodb with Python :https://www.youtube.com/playlist?list=PLZoTAELRMXVN_8zzsevm1bm6G-plsiO1I MySQL With Python :https://www.youtube.com/playlist?list=PLZoTAELRMXVMd3RF7p-u7ezEysGaG9JmO Deployment Architectures:https://www.youtube.com/playlist?list=PLZoTAELRMXVOPzVJiSJAn9Ly27Fi1-8ac Amazon sagemaker :https://www.youtube.com/playlist?list=PLZoTAELR
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Krish Naik · Krish Naik · 0 of 60

← Previous Next →
1 Natural Language Processing|Stemming
Natural Language Processing|Stemming
Krish Naik
2 Natural Language Processing|BagofWords
Natural Language Processing|BagofWords
Krish Naik
3 Gaussian distribution or Normal Distribution in statisctics
Gaussian distribution or Normal Distribution in statisctics
Krish Naik
4 Natural Language Processing|TF-IDF for Machine Learning| Text Prerocessing
Natural Language Processing|TF-IDF for Machine Learning| Text Prerocessing
Krish Naik
5 Log Normal Distribution in Statistics
Log Normal Distribution in Statistics
Krish Naik
6 Covariance in Statistics
Covariance in Statistics
Krish Naik
7 Confusion matrix, Precision, Recall| Data Science Interview questions
Confusion matrix, Precision, Recall| Data Science Interview questions
Krish Naik
8 Tutorial 44-Balanced vs Imbalanced Dataset and how to handle Imbalanced Dataset
Tutorial 44-Balanced vs Imbalanced Dataset and how to handle Imbalanced Dataset
Krish Naik
9 Implementing a Spam classifier in python| Natural Language Processing
Implementing a Spam classifier in python| Natural Language Processing
Krish Naik
10 Tutorial 11-Exploratory Data Analysis(EDA) of Titanic dataset
Tutorial 11-Exploratory Data Analysis(EDA) of Titanic dataset
Krish Naik
11 Face Recognition using open CV and VGG 16 Transfer Learning
Face Recognition using open CV and VGG 16 Transfer Learning
Krish Naik
12 Pedestrian Detection using OpenCV from Videos
Pedestrian Detection using OpenCV from Videos
Krish Naik
13 Face and Eye Detection from Videos using HAAR Cascade Classifier
Face and Eye Detection from Videos using HAAR Cascade Classifier
Krish Naik
14 Reading, Writing and Displaying images with Opencv| OpenCV Tutorial
Reading, Writing and Displaying images with Opencv| OpenCV Tutorial
Krish Naik
15 OpenCV Installation | OpenCV tutorial
OpenCV Installation | OpenCV tutorial
Krish Naik
16 Face and Eye Detection from Images using HAAR Cascade Classifier
Face and Eye Detection from Images using HAAR Cascade Classifier
Krish Naik
17 Car Detection using HAAR Cascade and Opencv from Videos.
Car Detection using HAAR Cascade and Opencv from Videos.
Krish Naik
18 Using OpenFace for Face recognition in Keras
Using OpenFace for Face recognition in Keras
Krish Naik
19 OpenPose Tutorial with Tensorflow
OpenPose Tutorial with Tensorflow
Krish Naik
20 Multiple Linear Regression using python and sklearn
Multiple Linear Regression using python and sklearn
Krish Naik
21 Dimensional Reduction| Principal Component Analysis
Dimensional Reduction| Principal Component Analysis
Krish Naik
22 Movie Recommender System using Python
Movie Recommender System using Python
Krish Naik
23 TPR,FPR,FNR,TNR, Confusion Matrix
TPR,FPR,FNR,TNR, Confusion Matrix
Krish Naik
24 Precision, Recall and F1-Score
Precision, Recall and F1-Score
Krish Naik
25 Artificial Neural Network for Customer's Exit Prediction from Bank
Artificial Neural Network for Customer's Exit Prediction from Bank
Krish Naik
26 GridSearchCV- Select the best hyperparameter for any Classification Model
GridSearchCV- Select the best hyperparameter for any Classification Model
Krish Naik
27 RandomizedSearchCV- Select the best hyperparameter for any Classification Model
RandomizedSearchCV- Select the best hyperparameter for any Classification Model
Krish Naik
28 K Nearest Neighbor classification with Intuition and practical solution
K Nearest Neighbor classification with Intuition and practical solution
Krish Naik
29 K Means Clustering Intuition
K Means Clustering Intuition
Krish Naik
30 Create custom Alexa Skill- Lambda function- Part2
Create custom Alexa Skill- Lambda function- Part2
Krish Naik
31 Hierarchical Clustering intuition
Hierarchical Clustering intuition
Krish Naik
32 Implement Transfer Learning with a generic Code Template
Implement Transfer Learning with a generic Code Template
Krish Naik
33 Gender Classifier and Age Estimator using Resnet Convolution Neural Network
Gender Classifier and Age Estimator using Resnet Convolution Neural Network
Krish Naik
34 Unlock Your Application With Your Face using OpenCV
Unlock Your Application With Your Face using OpenCV
Krish Naik
35 Draw rectangle from webcam and sketch process it on a live feed
Draw rectangle from webcam and sketch process it on a live feed
Krish Naik
36 Complete Life Cycle of a Data Science Project
Complete Life Cycle of a Data Science Project
Krish Naik
37 How we can apply Machine Learning in Finance
How we can apply Machine Learning in Finance
Krish Naik
38 Deep Learning in Medical Science
Deep Learning in Medical Science
Krish Naik
39 How to switch your career to Data Science.
How to switch your career to Data Science.
Krish Naik
40 Linear Regression Mathematical Intuition
Linear Regression Mathematical Intuition
Krish Naik
41 Handle Categorical features using Python
Handle Categorical features using Python
Krish Naik
42 Machine Learning Algorithm- Which one to choose for your Problem?
Machine Learning Algorithm- Which one to choose for your Problem?
Krish Naik
43 DBSCAN Clustering Easily Explained with Implementation
DBSCAN Clustering Easily Explained with Implementation
Krish Naik
44 Curse of Dimensionality Easily explained| Machine Learning
Curse of Dimensionality Easily explained| Machine Learning
Krish Naik
45 Feature Selection Techniques Easily Explained | Machine Learning
Feature Selection Techniques Easily Explained | Machine Learning
Krish Naik
46 Tutorial 29-R square and Adjusted R square Clearly Explained| Machine Learning
Tutorial 29-R square and Adjusted R square Clearly Explained| Machine Learning
Krish Naik
47 Cross Validation using sklearn and python | Machine Learning
Cross Validation using sklearn and python | Machine Learning
Krish Naik
48 Handling Missing Data Easily Explained| Machine Learning
Handling Missing Data Easily Explained| Machine Learning
Krish Naik
49 Deploy Machine Learning Model using Flask
Deploy Machine Learning Model using Flask
Krish Naik
50 Deployment of Deep Learning Model using Flask
Deployment of Deep Learning Model using Flask
Krish Naik
51 How to Visualize Multiple Linear Regression in python
How to Visualize Multiple Linear Regression in python
Krish Naik
52 K Nearest Neighbour Easily Explained with Implementation
K Nearest Neighbour Easily Explained with Implementation
Krish Naik
53 Predicting Heart Disease using Machine Learning
Predicting Heart Disease using Machine Learning
Krish Naik
54 Predicting Lungs Disease using Deep Learning
Predicting Lungs Disease using Deep Learning
Krish Naik
55 Stock Sentiment Analysis using News Headlines
Stock Sentiment Analysis using News Headlines
Krish Naik
56 Random Forest(Bootstrap Aggregation) Easily Explained
Random Forest(Bootstrap Aggregation) Easily Explained
Krish Naik
57 Voting Classifier(Hard Voting and Soft Voting Classifier)
Voting Classifier(Hard Voting and Soft Voting Classifier)
Krish Naik
58 Credit Card Fraud Detection using Machine Learning from Kaggle
Credit Card Fraud Detection using Machine Learning from Kaggle
Krish Naik
59 Hyperparameter Optimization for Xgboost
Hyperparameter Optimization for Xgboost
Krish Naik
60 Tutorial 45-Handling imbalanced Dataset  using python- Part 1
Tutorial 45-Handling imbalanced Dataset using python- Part 1
Krish Naik

Related Reads

📰
Call conventions : ce qui se passe vraiment quand tu appelles une fonction (et pourquoi C et C++ se…
Learn how call conventions work in C and C++ and why they sometimes conflict
Medium · Programming
📰
SvelteKit 2 Complete Guide: From Zero to Production (2026)
Learn how to build a full-stack application with SvelteKit 2 and Svelte 5's rune system, from setup to production, and boost your productivity
Dev.to · Carlos Oliva Pascual
📰
Implementing a Game Character Evolution System Backend: A 4-Stage System Design
Learn to design a 4-stage game character evolution system backend to integrate with point shops and inventory systems
Dev.to · 박준희
📰
High Level System Design Day-16 — Websockets (Part 4)
Learn to design scalable systems using WebSockets for real-time communication, a crucial skill for backend engineers
Medium · Programming
Up next
How To Install iOS 27 Beta on iPhone for FREE! (Step-by-Step)
Ksk Royal
Watch →