Distributed Storage with Hadoop | Big Data Engineering
Key Takeaways
Distributed Storage with Hadoop, Big Data Engineering with Azure and AWS Cloud, led by expert mentors Krish Naik and Mayank Aggarwal
Full Transcript
[Music] [Music] hello everyone uh welcome back let me see if we able to just stream it so for anyone who have joined can you please just confirm me if my voice is Audible and if I you are able to like see me fine that would be helpful so people who have joined please help me to confirm Mons [Music] [Music] great okay so people who have joined can you just confirm me let me hi lakman I hope you're able to hear me as well hello everyone yeah so I see the audience is increasing surely you can help me just quickly confirm if we are if I'm Audible and visible so [Music] yeah thanks a lot so [Music] cool I think it's 8 so if you have attended any of the previous session I hope you know the drill just a Kat here so today uh if you find my energy low or voice little low I just having some cold but I will make sure that you know that doesn't affect the session so yeah please if anytime anyone has any problem with the voice or anything let me know in the chat so great everyone let me just also put up my screen so today uh I'm just take a couple of steps back because many students were having few doubts and many people were there who were having this thing that they are coming from a totally unrelated background with with like like majorly with respect to your uh this thing one minute just a second everyone yeah should be fine okay so yeah many students were also having the doubt that if they're coming from a totally unrelated background maybe let's say they're in testing or they are from non- Tech so I wanted to make sure that the session is also uh give them an idea that how exactly are we going to go about things and yes that should give them a taste if big data is there for them or not not okay so today again I will be just talking a little bit more about Hadoop and distributed storage and this will give you an idea that what exactly do we do overall in data engineering and Big Data uh engineering right cool so I see people have joined so let me just maybe go to the page course page so everyone I hope till now let me share this link as well so we are going to start with this particular course this Saturday now so it's going to start from Saturday right December 21st and uh this most probably would be the second last session tomorrow session will be with chrish sir so chrish sir will also be joining tomorrow and there we will be answering some more of your doubts right everyone so yeah do ask away if anyone has any doubt meanwhile let me just quickly tell about all these things yeah I am a Java developer can we do it in Java or python required chandan uh we will be doing it in Python okay things can be done in Java if I be honest I also start Ed the first project of my big data Journey like harop and everything in Java but now honestly python is a lot preferred so it should not be a problem you can quickly Learn Python like it's lot easy than py Java and majorly most things which you are going to study let's say if you want to like deep dive into AI or data science ml python of course is the language of choice now so with that in mind we will be doing things in Python only but provided the overall I would say AI today like right now I was just checking that chat GPT has also opened up canvas so you can quickly change the code but we are going to focus on py spark and python right so on second let me have the banner so everyone please do ask your question so Chan I hope that is clear so python we will like we can actually watch the series on my YouTube channel or Chris Channel as well both are good to start with nothing more than that required okay so yeah this is the case everyone and we have few seats like less seats left I think CHR and he mentioned there was around one3 now if I remember I think less than 50 seats are left so Chris 10 is the code okay which you can still use it's giving you a 10% discount everyone through which you can join at this amount now again uh basically we have to make sure that if you have any doubts you can reach out to the number which is on screen as well as here I can quickly I'm going to quickly run through all the things uh which are basically helpful to you you will always see them uh like in your in the description of this particular video on YouTube okay so the session is going to start from this Saturday 8 to 12 we are saying proficiency level should be professionals I 2 plus years if it is less than that you can join just make sure that you are having some experience with coding if no experience is there then I will suggest to First do some coding and get that experience okay that again is very much required okay I won't say that if you have not done coding then big data is the course for you that thing you have to understand okay so let us see any other doubts yeah will the meeting invite be sent yes tar uh we have a good platform through which you will be getting that along with that if you miss the session then you will get the recordings right so so no worries on that I recent graduate BBA yes an so yes you can just call the number once uh again we have made sure that we are not doing any wrong sales with this uh like normally we will be guiding you properly so you can just call on the number the counseling team can help you understand your situation better and if you still have any doubt just let me know and you can connect with me as well okay so let me just share the uh share the ways you can connect with me so everyone this is my YouTube Channel and this is the LinkedIn through which you can connect with me so YouTube channel will have your email ID as well which you can just have okay uh KAG yes you can do that if you have a good exposure to data science surely this is going to helpful to you so I don't see and if you have no issues with coding then it's going to be a very good uh I would say learning experience because data engineering as such is a very good course right honestly it's great course with with the mind that you will be learning exactly major thing which happens in companies because every company is working on data building those pipelines so yeah that is going to be a very good exposure cool everyone so do keep your questions coming do have your Q&A so we are just taking questions as of now let me meanwhile go to slus as well and let me share this link with you all as well so you can access the slab yeah uh we will also cover details on how to make store data in a local database like post your SQL or my SQL and connect it to us your yes ab ab we are going to do that so I think I showed so tomorrow or day after we are going to I'm going to actually show you one of the video so it's a full project uh which will be available there you will get a lot of idea so actually we will be doing the database also on cloud so I found a good platform and connecting that it will be very close to how one of the projects which I did was at a very good level okay not an easy at all so yeah everything we'll be doing we will be making sure that our projects are complex enough and you will get an exposure once you see that project so you will be able to understand what exactly are we going to do how we are going to do the project as well hello my call said to make a data science project and they said the data is collected by you means we can take it from Kel and it should be big uh get said you can go to Kel and download the data like many datas are available just to show you okay because no worries let me help with that as well Ecom gigle uh let's go here this is one is small I remember one was there with was in GBS as well okay let me see multic category store I think this was the one if I remember right yeah see I don't think uh what is get set your college will expect your data to be greater than this and honestly you will not be able to even handle this uh in your normal laptop or somewhere you will have to work by distributed computing only right because this data is in size of GBS so yeah a very big CSV file right uh hello am I audible everyone can you just confirm once if I'm audible or not hello okay cool no worries great thanks for for the confirmation uh so anytime you face any issue with the either the voice screen or the uh video let me know in the chat okay we'll quickly solve that thank you everyone so yeah let me take the questions I hope uh get said you get your answer can you compare career in Big Data as compared to Java fullstack in terms of growth and progression uh uh see okay if I have to discuss about the career in big data and if you are talking about in terms of a proper Big Data developer right so opportunities are certainly more the reason is that many companies and this again I can tell from my experience only Java developer seats a very uh fault proof career honestly because many big companies are still working on Java like in gold for example I was working in Java and you will find projects which are in Java for any company which youed which was 10 founded 10 years earlier because at that time Java was of course a very good language and everything now what happened is that what I have personally seen again I am not a very Pro that okay one language is going to uh basically let's say I believe that the concepts and Basics should be clear okay now coming to Java developer part right of course there are opportunities but they are majorly with big companies and you can get a very specialized skill set and your demand is going to be a lot okay it may happen that if you are let's say just a Java developer with no idea about system design stuff or I would say let's say big data and exposure to these fields your career growth might also be hampered as well I have seen many people who are working in the service based industry they have great Java skill set but they have not been exposed to problems which are let's say difficult to solve they are majorly working on on creating apis okay again it also depends a lot on what type of problem are you trying to solve with Java okay language always just remember is a tool so talking about service based companies which I have been exposed to Via like mentees which I have connected with they are the problem was not that they working in Java the problem was that they were not getting exposed to different uh I would say complexities of things which they were solving it was majorly like they have to work on a code which was in Java create some API okay and of course like I have done the same but thankfully I was exposed to problems which are complex than that okay so using your uh spring Boot and everything have done that created those apis and everything but you should make sure that you are getting exposed to different different kind of problems like for example in today's world uh any companies which is new majorly they prefer javal L now again I not saying that they don't do it at all but let's say if any compan is getting created right now they will be using Ai and stuff for which Python and other languages are preferred right similarly for front end we have all together different stack so on that note if I have to choose one I will say that yes okay you are good with Java no problem but please make sure that you see how can you solve different problems using Java like for example I did my first project in big data using Java so there now language became a tool and I got exposed to Big Data Technologies spark Hadoop everything so that is something which you should understand okay Jan sir I'm mainly targeting data scientist profile do you think this course would give me an edge uh kushagara it will be helpful to you but again along with this you should be very good uh where I can see you fitting if you complete this course and let's say if you are good with data scientists is that small companies or small startups who want you to handle each and everything like see for them they will not have departments like let's say one team is there which is handling all this thing one thing is one team is there which is just for the data science part okay there they will be having and now if you have seen the project which we did right in the pipeline only you can apply machine learning models okay spark sorry this Azor which we did there also many things are there then AWS Sage maker is there so all those things can now be done on the go so if you have this information and you can build end to end pipeline starting from data sourcing data inje all the way to your uh let's say applying of ML and data science that is going to be a lot helpful to you okay okay let me see sir I'm many tar data science have done that hi I'm 7 experience business analyst having no experience in data will I be able to yes sham so this course should be a lot good for you because if you are from business analyst you will understand that how we work on getting the data and stuff just one thing if you have not exposed to coding right if you have been working in Excel and stuff I will suggest that you just start that right now because you will be using coding and SQL a lot SQL I think you will be using a lot but doing things via coding is going to be helpful and required okay will we also see how to work in a notebook in the cloud or other ways of coding in the cloud AB uh we will be doing everything so I showed in one of the project session which I took that how we are coding on the data breakes notebook right we will be coding on the terminal as well so again each and everything we will be doing as in as in the way it is required plus we will see how we can make a project end to end in software only let's say using vs code or your py ch okay so yeah just honestly I have done this all in the production and grade environment so that is something which I'm going to share with you and the main focus will be to make sure that your problem and understanding gets clear so if tomorrow someone is asking you to code in a terminal or a notebook you are good with both so that is going to be my focus again I'm not that kind of a person who feel a lot uh I would say impressed or something if someone is coding in a vs code and not or terminal I just want to make sure that problem gets solved and you have that optimistic and logical mindset how to understand the problem and how to code towards the solution okay great uh can I join if I have a little bit knowledge about Big Data yes Kip that is going to be helpful because again this will make sure that you have more knowledge on Big Data then okay so just five more minutes everyone I'm going to take question then we are going to start with today's session the way we have been doing so my project is on prediction and Sir said you also need to make a URL means HTTP how can I uh get set your uh this is a lot unclear honestly so I will suggest if you can somehow basically tell me that what exactly are the requirements that is a lot helpful okay okay I think sham yes thanks even I'm interested in distributed system so even backend developer would help me CH yes as a backend developer to it would help you because see in backend development also we are majorly working in writing those scripts Where We Are getting the data from one place or another and honestly you are without realizing doing a lots of data engineering stuff right so big data if you know that's just a cherry on the cake because then you will be getting exposed to multiple more things right multiple more Technologies and stacks which is going to be helpful to you okay can I go with cloud computing as AWS because it's booming right now I'm confused uh Kip we are going to use the cloud computing on your AWS Azure and gcp so you will get exposed to each and every one of them so don't worry uh we are not we might be doing very less thing on local maybe just the spark tutorial I might show you on local or maybe that also on Google collab so not anything on this right pratik are we going to cover Microsoft fabric pratique it is not there as of now okay uh actually I tried and Microsoft fabric you can actually not use with your normal account plus we will be using synapse and other things so you will get the taste fabric is just an all-in-one plan platform and again yes it is having lake house and everything so yeah uh you will get an idea how to use in fabric as well that is not going to be a problem okay uh bio I think uh Chris said in the basically that the you can just refer to the YouTube videos I will just update the link today don't worry okay prere requisit pratik uh Kang are these one so Python and SQL uh we have this on YouTube as well so you can just do that right uh KAG we will be doing projects in between as well and majorly towards the end but yes you will get the exposure that how we have to use these Technologies Okay cool so I think we are around 20 minutes map now let me quickly go through this cabus because many students I see are joining for the first time because they are asking on the prerequisits so yeah let us start with that before that first let me share this link again feel free to check this out at your end this is the laus which I'm referring to so big data boot camp with AWS and as your okay gcp also we will be using like just that majorly the projects we will be doing on AWS and Azor the reason for that is that they are majorly used in the industry plus if you know onecloud surely you will be able to work on other clouds as well right great so uh comprehensive yeah so these are all the things everyone I think I've read this multiple times uh this is just the basic idea that what all we will be doing okay along with that let's go to to module Z Now module Z is cross prerequisits python SQL and database uh Basics so what exactly is the database how to write SQL queries okay if you are given a let's say rdbms database what exactly it is uh I'm going to show you that you can create a database on the cloud then work on that that all is fine but I was uh just basically have I will like that if you know that okay yeah this is how we write python you can write basic functions in Python you can do some complex coding on python because that is a re it okay if required you can actually do that because the first uh around 1 month we will be focusing on lots of theory that is where I will be telling you about Big Data Hadoop spark uh sorry map reduce all these things Yan and then we will jump to your spark so you are all good if you are just doing this thing okay great so let us move forward everyone then first is going to be brief overview of Big Data Concepts and Foundation people who are just uh joining from let's say some tot unrelated domain let me tell you the way I will be teaching and you will see a little bit idea today as well and in the previous sessions I will be teaching everything from Basics okay so it's not that I will say that okay you know what big data is let me move forward no that is not going to be the case we are going to learn each and everything from Basics so that everything is clear okay great uh okay te dive okay so next is Hadoop architecture now I will be telling you each and everything about Hadoop because it is important to understand you see the thing is again you might be using S3 or blob storage in Azor or let's say whatever gcp uses but the main point or main idea is that you should know how the distributed storage and things work right then Foundation of Apache spark very important spark is very very important thing I will first teach you spark and then we will be jumping to data bricks if anyone feels that they can Master data brakes without spark extremely sorry that is going to be a very difficult part to do okay so these are the things which we are going to cover in spark let me share this again link so that you can download this at your end right then data frames and structure data processing so spark basically have rdds data frames SQL all these things we are going to cover okay so Advanced data processing and optimization in terms of joins your caching all these things are also there which we are going to cover okay let us move forward performance tuning and optimization again of course we have to make sure that we our code which we write is a L tuned with performance it should be running quickly we should see where the bottom leg might arise so all that we have to solve then again we have to I will be teaching you about no SQL databases okay and it's comparison with SQL database SQL you will be doing as a prerequisite and no SQL we will be covering two database mongodb and cassendra okay so these are a lot used in industry I have actually used both of them in the industry only so not like we are studying something which is not getting used okay and plus SQL knowledge is still required I hope all of you understand that SQL is still required a lot okay then let's come to Hive architecture Hive is like a data warehouse which is inbuilt in like which you can install on top of Ado so we are going to study about Hive as well then Advanced hi features and everything okay okay then yeah so introduction to Kafka now let's talk about the real time scenarios so if you want to develop a chatbot or let's say you want to make a co-pilot kind of a thing so spark steaming and Kafka the real time publisher subscriber model these things are going to be helpful they are also going to be a lot helpful in your software development Journey then Kafka producer consumer we will set it up online and on the cloud as well as on local as well then spark structure streaming again okay then we are going to understand about Apache airflow how we can make awesome pipelines data pipelines using the same we will also have projects here okay so you'll get an idea this thing again is a lot helpful when you do data science projects as well okay then cloud computing and overview of aor so this again I think is a lot clear I will be explaining about Cloud but again one thing more everyone because we will be starting and I want to make sure that anyone who doesn't have an idea though we will use cloud here and there I will be explaining properly that what cloud exactly is and how we use it okay then as your data storage services so many things about a we are going to cover a your data breaks okay a your data Factory then Advan data Factory transformation monitoring and error handling so these things are going to be covered then we are going to jump to AWS AWS EMR okay then let us move forward AWS S3 again for storage then Athena and glue for Server squaring and ETL so making those pipelines which you see and then we are going to do projects so at least five projects we are going to do the information is there you can see that we are using things which we have learned so big data ATL pipelines ofo P then we are do to going to use spark and Kafka steaming we are going to make sure that we create an end to endend proper uh I would say uh pipeline which is handling industry level data and you're working and finally serving it as well and using AWS also we are going to do the project Okay cool so I hope that all is fine now let me jump to today's session uh this is little bit about me for people who are not aware I have spent enough time in the starting so I am a graduate of nsit I have previously worked at Goldman Sachs myle oo rooms and - level startups like I have a fair share of exposure so that is how I will be teaching I am a very strong proponent that if your teacher he has the hands-on experience he's he will be able to guide you in a lot better way so I want to make sure that all the students you are in good hands and yeah that I can assure you okay cool then everyone let us start with today's topic so today's topic everyone is distributed storage with Hadoop right okay let me just quickly see if we have got any doubts I think isan has answered the one how much practice required apart from regular classes rajes I would expect that you are spending close to 10 hours a week to make sure that you are getting the exposed to other things many things will be there which you can of course see on your own so yeah that my practice I will expect that you do okay hi sir hope fine and doing well I have been confused about one topic since one year time series belong to AI specialization or not belong to AI it doesn't as such belongs to AI Ahmed time series is something separate than a but then again we are trying to predict that how with time things are going to change so majorly it will come as a artificial intelligence only but yes as a topic it is not that you need any major things in AI to do that okay uh okay I think I uh entered yeah I have asked that so should be fine great everyone so let us start in this uh I'm going to also explain about uh Hadoop as well okay Hardo is going to be a lot important everyone so let us begin uh now if you have any doubts you can just uh basically enter that uh in the chat and I will take that when we pause or normally okay so for the next half an hour 45 minutes we'll be teaching just before I start uh can anyone please help me to just tell if audio video and Screen everything is fine and maybe I can just change to this particular uh way so one minute yeah oh I think both works fine only should be fine so can anyone confirm me all these three three things are fine that is going to be a lot helpful maybe you can write yes yes yes but cool thank you great so now let us actually start and see first thing I would like to clear about Hadoop uh how many of you think Hadoop is a new technology relatively new technology or something like that like Hadoop is something which has just come to the market uh Hadoop is has to majorly do with I would say after like it is launched in 2017 18 how many of you think that it's a complet comparatively new technology is it is it new or old like what is your mindset behind that is it new or old so yeah Haro is actually been there there for a lot of time okay it's or launched in 2005 so Google actually wrote a paper on distributed computing and map use like processing okay just that overall the thing is that because it was so good and it gave idea to a lot lots of new things that the same idea right let's say the distributed storage idea that is getting used in different different Technologies now that is what makes understanding Hadoop a lot lot more important okay everyone so many non-technical people out people who have not used it they feel that it has just come into picture or like this big data something no it has been there for quite a long time okay so in Hadoop we have three components which we are going to discuss as well today first is hdfs that is for storage second is map reduce Now map reduce is for uh processing third is Yar now these are the major components like of course you have Hive Pig all these things but they can be installed afterwards right but Hive is kind of like let's say data warehouse provided by Hado again there are some use cases some not so we have to clear that idea then we can use some other data house as well right just like we know let's say if you know Java you can you understand that okay there is a loop that Loop concept Remains the Same in python as well okay so let's talk about Hadoop a little bit everyone okay it was made by uh duck cutting if I remember the name right anyone can just correct me if the name is wrong or anything okay so let's talk about how it came into picture like Duck uh basically invented or worked on hero and then he made hero open source okay so it came under Apache and came to be known as Apache Hadoop anyone who knows or understand about open source you will see that yes it is available for everyone free to use not a worry but when you have something as open source there's a problem as well like everyone is basically coding on that and because it's open source no one in a way owns that it they are going to help you with your support request right so for that you will find many paid implementation of aoop or let's say companies have their S3 and other things but yeah paid implementations of her doop in a very easy way if I have to just tell you or you have to understand and majorly this will be a lot uh this will help you to clear lots of things think of how exactly Android and Apple Works Android OS and Apple OS now see Android is open source right again created by Google made open source it is being used by many companies Mi Samsung uh nothing OnePlus everyone is using this right they are basically making sure that they can provide some extra features on top of that and everything and it's normally a lot buggy right because again everyone is trying to make their own versions and everything and the core Android major very less companies use they are trying to make sure that they are developing that based on their requirement whereas if I talk about Apple OS a single company is handling that so they know they have the version specified okay so for example let's say if some issue is there in Samsung Android version it might not be there in mi okay so they are not comparable even but Apple we know that a single version is there right and here if you see okay it's a lot stable comparatively lot stable than Android or anything right everyone so again this thing is very important to understand that what happened is that Apache Hadoop yes it is there you can use you can download that you can go to Apache right now like I can actually just show you that as well I can go to Google Apache Hadoop and why am I clearing all these things because this is also something which students have doubt right see you can download Hadoop 3.4.1 is the latest version second version is also available okay Source download check some signature like things are available you can download it you can set it up yourself as well it is not going to provide you with support for that the very first company which made this thing possible okay which provides support and everything as well was Cloud era okay they made sure that they are commercializing it they are picking up this Hadoop providing you with the service and also making sure that we provide the support so everything they are handling so that you can just use it and any issue which you face they are going to help you with that okay so I hope this is clear everyone anyone has any doubt in just a basic history of Ado let me know in the chat if anyone has any doubt I hope everything is fil 10 now anyone has any doubt let me know quickly great now let us talk about a little bit about distributed computing distributed computing see in a very easy way distributed computing is when you're using multiple nodes or multiple desktops or computers right now multiple things come into picture here as well which I want to make sure that you understand from very Basics right uh for many people you are you might have a little bit idea basic idea but I want to make sure that you have a clear clear cut idea on this let's say this is your one machine everyone okay this is your one machine this is your single node part and if I just uh let's say make something like this a cluster like this this is your multi note part okay this is your multi note part so on the left hand you have a single node on the right side you have a multi- node okay whenever you are using more than one machines in doing something we can say that okay you you are using distributed computing okay but few things which I want all of you to understand okay is that uh see maybe it can happen that you have four PCS on your home right now you have one your brother has one maybe your father has a work one who will write that code okay and this is very important to understand who is going to write that code which connects these things with one another that is where Hadoop as a framework comes into picture like someone has done that effort that hey install me into these machines okay tell me how where they are kept like tell me the network location and I will handle the rest of the part okay so it's not that uh because see you will never be setting this up so you will never get this kind of a thinking as well so when Google wrote that paper right this was the main thing which they soled that to make a software through which this storing and and your handling and processing of data can be done in a distributed way okay and this was not an easy thing to solve like you can yourself think of that how many programs you might have to write how much long code you might have to write so this was something which was solved by Google and then taken over as an inspiration for hu right and this is the very basic about distributed computing like you can think that it might be a lot easy for you to think that yes I have four computers I will just use them no someone might have to write to make sure that you can use them in a production setup yes you can maybe use them for a basic downloading something but how will they sync right how are they going to syn let's say if you're trying to just download a single file how will they get divided this was the tough nut to crack which Google did and which is now used in hadu okay so that is where I want to make sure that you understand exactly this is the distributed computing now to explain you in a very easy manner let's say you have this puzzle to solve let's say we have this puzzle divided into four parts and it's a very big puzzle right is not a small puzzle what you can do is you can say that hey it is having four parts 1 two three and four you can give this to one of your friend this to another friend this to another and maybe you are working on this yourself now what they can do is all of them they can take this 1 2 3 4 part at their home okay let's say this is a puzzle let's have multip pieces so they can take this up and store it in their home and they can work on that on their home locally and this is what this thing right to make sure that computers can do that this is what Hadoop makes possible or any distributed architecture for that and Hadoop is an architecture it's not a software it's not something which you uh will like let's say take like a chrome or something no okay and it can be installed on any machines and you can form a cluster for how much machines or notes you want that is the basic idea behind distributed computing and where Hadoop comes into picture now why whenever we talk about Big Data distributed computing comes into picture the reason for that again is that because data is becoming so big that it is not possible for a single machine to process yes you can still store I can still say or give you that thing that you can have some Network assisted drives okay or solution or space okay it is known as any commonly and you can store that data not a problem the problem happens when you have to process on that and that to in a real time in a quick way okay so for that basically because Big Data as we say the data is increasing a lot in size typically it is very difficult for a single system to handle and that's why we first went to the distributed uh architecture and then we come up with something known as Hadoop which can connect these machines to work as a single unit to solve your problem okay so what Hadoop and map ruce will do is Hadoop makes it possible for each and every one of them to take it to their home basically divide this puzzle in a proper way and map reduce ensures that you are able to or sorry these all are able to solve this independently like this first person is able to solve it independently at their home so we can do local processing right so this computer this computer this computer all of them can do local processing and once their solution is done we can combine the results to get your total results okay and all these this all is handled by a Yan which is a resource manager which makes sure that you know this person he is able to solve he has the resources like it is not that the light power cut is there at its home no Yan will make sure that everyone is having their resources if someone's get free then he can take up some other part all these things are taken up by Yan so a very easy implement uh instructions and idea anyone has any doubt in what all I have explained till now I hope you like the way I'm teaching just basic thing nothing no definition nothing so I don't like definition honestly I might give you them but I will make sure that I explain you using examples only so for even a person who has no exposure to uh who are like let's say even working in testing or has no exposure I hope you get this Basics right that what exactly are we trying to achieve and where Hadoop and everything kicks in because this thing is is not a lot clear to many people like many people just have like we will have a group cluster and I will work on that no someone is written the major part of your code which you don't have to worry about so when you say to this machine that hey uh upload my file Hadoop handles this thing right not like uh you don't have to write the code but yes Hadoop has done that heavy lifting okay uh sham uh torrent downloading is peer-to-peer so see okay that's a nice question and let me explain that as well see what happen is that hadu follows a master worker architecture or Master Slave architecture where one machine is basically handling as the admin or the major and it is kind of telling other machines to work on the problem that is why my uh this thing is like uh the arrows are like this coming from a single machine okay in a peer to-peer that is something on which if I remember right cassendra is also based they are connected like this so on a very very big note what happens in a torrent is that this person has the file and he is known as a seedar okay he is seeding this is a leecher he is leeching that file from this person and it's a peer-to-peer connection peer-to-peer connection is also used in many places uh if you are anywhere where or maybe have used the to browser it's also peer-to-peer connection so architecture in the way you are using distributed computing that is different in torrent there is no one single machine from which everyone is downloading no it is like uh these people are connected via torrent maybe these TOS have the file downloaded and let's say 50 people are trying to download from them so two seeders 50 leers normally it is suggested and a very good habit to seed your torrent right now it no one I think now less people are downloading that uh I used to seed the torrent till 150% like let's say once I download them if I'm downloading a 100 MB file 150 MB I made sure to see it just being a good internet citizen but yes torrent works on that principle it can go a lot uh maybe complexity is there but yeah that is what I learned way back in my school days because I was a lot interested in these things but yeah torrent Works peerto beer Okay cool so anyone has any other doubt before I move forward I hope this thing is clear to everyone including anyone who has a no exp no background in Tech or anything I hope this is clear can I get a quick yes see everyone be a lot active not a movie okay we have to make sure that we are lot active if we are spending the time after office to learn something please make sure if you have any doubts you just clear it out great so let us move forward and do a few things more okay now as I said harop as an architecture works on a master worker type so this is one let's say node we can say it as a node machine something like that this is the first this is the second and this is the third one okay now what happen is that when you install Hadoop Hadoop asked that hey Master node is y no s you cannot say that Master node is not y uh sh Master Slave architecture is the same only no like this is what Master Slave architecture is soan I did cs50 back in my college uh any assignment or anything which you will be given in this uh course you will be able to do them properly with the things which I will be teaching maybe something extra right projects and everything it will be complex but coming from what I will be teaching you okay I hope it is clear great uh sham I hope you get the idea that this is Master worker Master Slave this is the architecture now what happened is that we have some uh demon or a program running on these machines okay so this is master this is let's say worker worker one worker two worker three let's say all of them is having three TB of space now you will have total of nine TBS some little less because something is used for uh like preserve for OS and everything so we will be installing operating systems on this if any of you remember in uh gcp data proc I showed you that how we can create it easily so you will not be stalling them in today's world for sure because any company they are they either have admins or uh I would say operational guys to handle that or you do that quickly via Cloud but again the architecture needs to be clear so hero as I said is an architecture okay we have to now make sure that see if uh these are the machines right if these are the machines and if I ask my this particular architecture that hey I want to store this 5 TB of file okay now coming to distributed storage this file is going to get divided into blocks okay so it's going to get divided into multiple blocks in the latest version of Haro the default block size is 128 MB so you can do the math with how many blocks are going to get created we will have this information so this master node it maintains a metadata it maintains a metadata table okay it's going to send basically going to make sure that these blocks get saved into these things now again we are going to like this is a lot complex not that straightforward but yes just to give you an idea this is how it's going to happen it's going to say that hey let's say this first block is in worker one second block is in worker two okay there are some Demon which is running okay demon is a program which runs on machines so so that you can again don't want to make it a little lot complex as of now right but there is a in master node we have a separate demon worker we have a separate demon okay these are the cases so we have to make sure that this the weight is happening again because these machines which you are taking they are not very heavy machine like they are not Apple M1 or something they are commodity Hardware which are very cheap so now many people will say that okay it is pretty easy right but what if tomorrow this machine dies this machine due to some reason the hard drive just fails okay Hado makes sure that it is also doing that fault tolerance it is doing making copies so that if one of the machine dies your data or your overall work is still not uh I would say failed it is it will still be able to get you the files so that is where replication comes into picture how it I can explain you in an easy term let's say if you have a very important file you store it on your laptop on your phone and on one of the hard drive or pen drive so that if tomorrow your phone gets stolen okay or it just becomes dead you know that you have your file right uh example can be I think many of you might have your IDs okay your different different IDs on phone maybe on some WhatsApp chat and on the laptop as well so that you know that okay if by chance you don't have any one of them or any one of them gets bad then you can get it from other places so this is in a very easy way data replication to achieve fall tolerance okay uh let me quickly see the questions uh C let me know what kind of data set we will be working on I would really appreciate that some more substantial data set uh suan uh we will be uploading a project so you will get an idea okay again the data set we will be taking will be close we might take some data from Kel we can make up some data so it will give you an idea that okay what kind of things you will be solving okay so don't worry on that how much time should dedicate outside of class to keep up with the material uh I answered this already suan 10 hours majorly I will uh suggest because so that you can learn about more things and do many things on your own as well so at least 8 to 10 hours you will have to spend uh because see we will have three three hours class uh I believe that it will take you around that much time only to revise the class and on top of that if you have some doubts to clear that out to discuss with the community you will have to spend that much time is the class really four hours Longs uh don't worry on that suan I will be I can teach for 3 hours okay it won't be exhausting I have taken classes for more than eight hours also in a day have been doing this teaching thing for a long time okay so you don't have to worry on that okay and you don't have to uh face any problem with that I normally I don't lose the energy while teaching so yeah you will not face any that problem I hope that you are ready because yes for sitting in a three hours I will be giving breaks in between but yes that is required okay any doubt anyone any other doubt which I can help you clear see the thing is that if you want to actually transform your career and if I am teaching a professional batch like data engineering so I have done that again previously as well Big Data data engineering data science ml AI gen we have to spend that time and 3 hours is a normal time I I think that and plus it will be all learning and everything right so don't have to worry on that any doubt anyone I'm just uh taking a minute to ask let you all ask your doubt so that we can then move forward please ask the same if everything is clear just then also write a yes or if you like the way I'm explaining just let me know so that I know if I have to change something or if I can go forward with this only and if anyone is feeling that uh this is very basic stuff well uh we have took sessions earlier where we have done an inin project where we have discussed about uh data breaks Spark all these things so we have done the fair share of difficult topics introduction and everything as well okay just that this I want to this particular topic I wanted to discuss so that students who have uh who are coming from separate Fields because many of them reach out to me on LinkedIn they also get an idea that how we are to discuss things and go about a particular topic okay so sort of after partition into 12 128 MB these partitions get uh this uh these partition these blocks okay get stored into different different machines okay now again there are lots of complexity with increasing the partition decreasing the number of partition increasing this partition size like for example you can take it 1 GB as well but yeah there will be some things which will not be right then you can take it 2 MB as well then also we will see what happens but yeah by default it is 128 MB uh yes praque for sure we'll have to cover all these things okay because people will be joining right so I will have to make sure that if someone misses who have joined right and doubt anyone let me know so see what happen is that we normally call this particular thing as name node once you send the file name node basically has this full information that okay where are all the machines kept okay where uh let's say which machine has s space all these things are maintained these machines right they are in constant connection of your master so they are in constantly sending their pulse so Master know that they are alive again as I said the major heavy lifting has been done to you when you install hadu it is not that this machine will be kept and out of the blue it will be used no they are in constant connection with each other if for 10 seconds this machine doesn't send any communication it might be assumed to be dead we will assume that it is dead right these machines can be even kept anywhere normally we try that they are of course close to each other but yes they can be kept anywhere as well uh Kalani the prerequisites for the course are let me go back to the first this thing are SQL and database Basics along with python so python SQ and database Basics very normal things nothing major right soan if you are a TPM right now and you want to shift to data engineering uh it's going to be honestly a little difficult and it also depends that what kind of projects you have managed when you are TPM like I know majorly you will be working on I think J if you can just give a little idea that what exactly you have done at TPM you will have to put in that extra efforts to show and also align and say that yes I was managing these kinds of project and then I got that overall uh interest into these feeds and thus I did that it's not that you will directly be having uh no no that is fine but it also depends a lot on what kind of work you're doing right everyone so that thingone have to uh handle anyone any other question great so let me just quickly complete the class I meant to take for an hour only today majorly the questions so yeah let us now discuss a little bit about map reduce as well okay map reduce yes it is not used like map reduce honestly not used a lot now but I feel that it is still required map redu has two words map plus reduce so what we do here let's say if again if you have this overall one machine one of the workers and this is a worker only second worker third [Music] worker uh how many of you are aware about the map function map function in Python map function in Python uh s you can go to krishna's YouTube channel under that in live last two sessions are all the things which we have done okay in a very easy sense if you think about map function what it does is it takes let's say a list or something and a function and apply that to every element in the list just a normal analogy again not this map is not very straightforward or equivalent to this but here similar things happen let's say if I'm giving this three words okay let's say uh first second and third these are the words I'm giving and I am passing all of them the instruction that they have to run the same program what is that program the program is Count letters right so all of them are going to apply this and get their individual so they are going to do local processing right so this is going to process this machine is going to process what is kept here only it doesn't care about any other machine right it just have its connection with the master and it will be doing this local processing right so it can happen that this is your one text file one big text file okay and you are just divided this text file and then working on this right so based on let's say the block size your file is divided and you are working on this now it's going to get the answer which is let me delete this remove this as well it's going to get the answer which is five six and five this is the map part okay this is the map part overall and then we reduce this answer so we are going to now reduce the answer by adding them it's going to be 16 now again very simplified explanation lots of things happen behind the the back I'm not saying that this is what but yes the very basic fun is this only that this is map reduce now there are challenges limitations like for example what if you have to do filter you have to make sure that you write your program in such a way that it fits that map thing plus you cannot now write something more on top of that okay you cannot write more operations on top of that so you can just do reduce once only so again we will discuss that but the basic idea is that this is how we do the processing in our big data framework okay now why is it not used uh the programming in map Ru is a lot difficult we are going to write a Java program just to show you you don't have to write that but the basic idea is that I will be teaching that why map ruce what map ruce is the very first thing second why is it not used now that again is very important to understand okay third why have we then jumped to spark and what benefit spark does give us all these things will be discussed so you don't have to worry but yes this is a very basic idea about how the processing happens in your uh distributed architecture especially in Haru okay and behind the scenes this like resource management and everything happens via the help of your uh yan yan also has an architecture again everyone has this architecture just to show you uh like see all these things are complex and I don't want see I want to make sure that you understand these things in depth okay let's say do write architecture so you will see how complex this is let's say if I go to image you will see many components see this is the full architecture which Hadoop Yaks right takes into picture you will see multiple steps so again going to discuss each and everything of that but I hope you get the basic idea right because before that I personally make sure that student has idea about everything okay so that is how I like to teach so that if he is standing somewhere he is doing any project he can say with confidence that yes I know things right because that is how my journey has accelerated not by just learning things from above no okay everyone so yeah if anyone has any doubt let me know this was majorly the things which I have to discuss wanted to keep it because these are very basic things wanted to keep it a lot uh easy only but yeah I hope you get the idea so let me just drop the screen share and yes everyone so if you have any doubts just let me know and let me just actually add the screen once and show you once again yeah so for people uh who have to check out the course please do it here we have limited streats left I think isan just messaged me that we have around 35 some seats left uh we have not kept this course with the mindset ke let's say thousand people can join so please uh if you're planning do just quickly enroll and if you have any doubts whatsoever please reach out to us and we will be more than happy to help you not sell you okay and yeah I think that is the major thing again just quickly repeating this is the full cabus which you can download the link I have shared from where you can enroll proficiency level all these things are given and dashboard access you will have for one and a half years you will get all the recordings okay the duration of the course is 5 to 7 months though if uh many people are able to understand things we might try to complete it in four months as well because again I want to make sure that I you are very quickly ready and giving interviews then resume discussion mock interviews job referrals all the things are going to happen you will also get a community chat forum for discussion so with your uh basically peers you will be able to discuss each and everything so that you do the learning with each and everyone so if someone is getting a job maybe they can refer to you as well all these things we will be doing okay everyone if you have any doubt reach out to this number if you want to connect with me I'm giving my LinkedIn ID and YouTube idea YouTube uh Channel as well feel free to connect with me on any Forum okay you haveo and its one code terminology map ruce elgo created by Google data move query move par to every node yes take answers from every node and join yes exactly is awsn aure mandatory uh yes sham exactly see again my aim is also not to make sure that uh I am just teaching you things so that you know things but you are not able to get the job AWS is used in the industry but I have said it multiple times uh you might have to just have two credit cards which I don't think is difficult ult in today's time uh then you have to make sure that you quickly work and learn on the same if you follow all the instructions highly like hardly you will have to like normally if let's say you're not slacking off then from the trial only you will be able to do each and everything if not then maybe it can take you maximum 500 something but then again that is when you will be slacking off so but we will be doing it on cloud and unfortunately I tried finding uh that level of exposure which these platform these clouds are providing no one else is providing so I cannot say that for as your data Factory no let us do it on some other tool no not not there okay plus because they will be used in the industry it is a lot required right let me show you one of the projects so actually that has been just done I have to miss just do that thing so this is the architecture of the project which I'm creating a lot related to how I have done projects in the industry we have multiple data sources HTTP SQL table data inje then we are using Azor Technologies here lots of them then visualization so yeah you will get an idea that what level of projects we will be doing honestly if you do this level of projects very nicely and you know all things you will be able to crack jobs easily I feel that reason is that many Technologies are discussed there and again Basics will also be clear right cool everyone so uh five more minutes for the question then we are going to meet tomorrow so there's a session tomorrow as well CH is also going to join that there we will be taking some commonly asked question like for people specifically from I would say this non-tech field okay so there we will have K understanding as well let me just tell what exactly this is who is [Music] this uh one minute let me just tell you the topics so we got some common doubts okay like how big data is connected with web development data science other domain career transition from non technical to Big Data and how to make a career in Big Data if you have a career Gap so all these things we are going to discuss okay we tayor resum to show ourselves as data engineer for data engineering job uh yes sham like we will try like I will tell you so see this is also based on lots of my exposure when I so I have taken actually uh many interviews I like to actually take interviews okay to be very close to hiring process we will try to add those subtle things okay to make sure that it as you may like interview and know that yes this person is trying to go towards uh this big data and we will add projects we will make sure that we add those Technologies technical skills which are a lot sort of so yeah we will be doing that of course that is required I will be showing you examples and way so yeah J with big data no salum we are not introducing gen here okay plus like J I think can help you to write the code and not sure how it will help you to make that pipeline or something so J is not there in the course okay I know that you can now include gen in everything like J with Excel and all these things but yeah J can help you to write code and stuff help you to do stuff easily but the core things again that generative I will not be able to do okay so it's the same like gen and maths yes if you know maths gen can help you with stuff but basic math should be still needed okay anyone any other question before we drop let me know I'm just just going to wait till 9:10 can I be able to crack mang with this course uh S as someone who have cleared interviews of these companies and someone who have actually friends in all of these companies and can also uh know about the process okay uh normally I try to use the term tier one companies you will have to be very good in other things as well okay and I'm being very honest here you have to be very good in DSA because all of these companies are going to ask you DSA first then they can ask you on system design things as well right on your previous experience and then let's say if you are specifically going for data engineering then I can be sure that yes if you do this course properly you will have that understanding you will not be lacking behind in anything if you do this properly now it may happen that they are asking you some tool let's say uh you are doing HD insights they're asking you for EMR or you did it in Azor they are asking some other tool you will be able to make sense and do that but again for this tier one companies like I can give you an example for Goldman I give 10 inter they were on lots of things specifically a lot in depth in DSA things right then in the next company I was asked for system design DSA lowlevel design uh hld all these things along with the projects and CV so these companies understand the overall character of a interview right anyone who says that do this course and you can crack pr1 companies uh he might be lying or TR to sell you but yeah being honest here if you do this properly data engineering side your thing will be done but if you come tomorrow to me sir they ask DSA and I have not done DSA it's going to be difficult but you can do one thing I have a course on udmi as well for DSA so if you go on udmi and because of good overall things so let's say if I search DSA python uh yeah you will see it on top premium bestseller okay if I search I think DSA only you might get this on few top so it has been going pretty well right lots of good reviews 4.6 rating uh you can take this that's a very good start I will be enhancing this course but yes this is all based on my experience but yes you will have to put in the time uh celum okay please cover each and everything related to Big Data also the new Big Data text tags in such a way that enough to crack remote and T1 companies uh we'll already be doing that salum so yeah again see Sal are uh a lot focused get your Concepts clear you will be able to crack those companies it will also depend a lot on other things that how good you are with other things and I'm being really honest here right I could have sell this easily that yes do this course at M no do DSA do things make your holistic profile a lot good so that you they know that you're a good engineer okay I'm a bit worried as I am s years of experience as a business anal SQL how will ible to justify my experience as a data engineer uh sham we will connect and talk on that we will have to discuss see I cannot give generic Gan I will have to understand your overall history of work experience how where you are coming from and then normally I'm pretty good in this thing so as I told like I have taken more than 500 plus mock interviews these sessions and everything as a part of scaler coding ninjas function up and other platforms we will be able to come up with something so that at least you are able to give interviews and then cook up some story to how to show that smartly I cannot say that that no it will be handled no I will have to understand what you have worked on and everything okay sir is DSA important for data analyst or data scientist data analyst might not be data scientist uh basically depends very less companies but top tier companies are going to ask you that uh DSA because they it is normally used as a filtering criteria plus what I feel is that if you spent enough time on DSA you can understand difficult of the code as well right same Mas like [Music] shubam suan and shubam see the thing there is that again then what options do you have will you just sit back and be like okay I spent seven years let me know not do that I think it's still better now than to try then to basically not do anything so I personally believe that you can change your Fields if you put in enough efforts and companies are maybe your seven years of experience is helpful to you if you showcase that in a proper way and this is the way I'm going to teach so I hope anyone coming from a different background he doesn't face those problems right uh sarak if you are in fourth year I would expect that either you are very good in coding and you understand few things like database SQL and everything and you have done some software development uh project in a way that you have an idea how the code flows that level of understanding I would expect if you're in fourth year okay your benefit for you will be that if you give companies and you are able to explain these things uh companies will be a lot uh impressed provided you know the basics first right like you cannot be like okay I don't know coding I don't know I don't know DSA I don't know uh anything else I just know big data then it's going to be a problem uh s of again uh totally different have to know your background cannot just suggest like that I I don't give generic Gan I try to understand student personally where he is coming from what he has done College skills and then I am basically prefer because see I have done mentoring for a long time as I'm saying like since College I have been doing that because my seniors mentored me very nicely I have seen one thing gend G doesn't work I can say that do DSA do this this this but where are you coming from are you currently in service based company or product based companies if you're in a product based company that's say in payments maybe I can say that hey just Target another company in payment from PTM to phone pay is easier so yeah but as a again yeah DSA has to be done properly I will be making content on that which I feel is important but again DSA DSA should be done at that level that you will be having uh great uh okay let me tell you like this way okay I cannot again will be but you should be able to do at least 250 questions properly not like just seeing the solution you know that yes you do this properly that is what is to be done huh that should be fine when can we expect gcp data engineering sha we are going to use D gcp here so again what again gcp data engineering is just a person who use Google Cloud platform I have made this point clear that in my classes I will focus on Basics so that tomorrow if you have to do the same thing in aor or in gcp you are able to do that so in this course you will get that exposure and and and and and and yeah it should be more than good for that uh Sal this course is starting on Saturday it's mentioned here see start date this Saturday now December 21st okay okay time to brush up my DSA with the course will two hours of DSA be enough sir no I think you were in fourth year no I just to give you an idea I used to do DSA for 10 to 14 hours every day in my fourth year after placement just to give you an idea okay not showing off for anything I used to do that do that because if you're in college then to DSA is like you are something which you're doing in your free time two hours per day might not cut it that was my experience again if you're are smart enough you can maybe do that I was not able to do it two hours per day uh okay so I hope that was all the question everyone uh we are already 9:15 so maely we take this session for an hour or something because I see that then students keep on uh they basically get that exposure and I don't want to make it boring like teaching for three hours uh if required let me know in the comments and I can take their classes as well but yeah so everything is there everyone uh you can just reach out you can just check out on this link uh in the description of this video you also have this information okay or uh for people who I said that I might need to know you better please connect with me Depending on time I will be more than happy to connect I have been doing that on uh LinkedIn with people LinkedIn and YouTube also people reach out to me on mail because as I said generan I cannot give that might that will surely not work for you meting me this is one thing which I have learned uh if I just tell you H do this two 50 questions do this do this it's better if you understand where you are coming from and then I guide you my experience that has been my experience of like I have seen students genuinely take the leave when I do like that cool then uh I don't see any other questions so just a reminder everyone it's going to start Saturday tomorrow we are going to have the last session before the batch okay I might take other sessions maybe once we start okay so on that note yes let us meet on Saturday for people who have joined for other people feel free to connect with me on LinkedIn I'm sharing my uh account once again sharing my YouTube channel as well uh have few videos on Big Data there as well okay so yeah you can just see that there as well the videos we'll be uploading many videos in the coming time and yeah that I think should be all if anyone has any doubt feel free to reach out to this number or connect with me uh most of the things are I think uh done okay yeah cool great so thank you everyone uh thanks a lot for your patience thanks a lot for being with me and we will meet tomorrow please make sure that you doubt get your doubts clear and yeah you should be please just put in the effort so that we can have a very good transition stories and we are able to enjoy our career okay good everyone so bye-bye uh let's meet tomorrow
Original Description
🚀 Big Data Bootcamp with Azure & AWS Cloud | Starts Dec 21, 2024 🗓️
Ready to become a Big Data Engineer with expertise in Azure and AWS Cloud? Join our Big Data Bootcamp led by expert mentors Krish Naik and Mayank Aggarwal, starting December 21, 2024. This program is perfect for professionals with 2+ years of experience who are eager to enhance their skills in cloud-based data engineering!
Why Join?
Master Big Data concepts & tools like Apache Spark, Kafka, and more.
Learn how to integrate Big Data solutions with Azure and AWS Cloud.
Hands-on projects & real-world use cases to build your expertise.
Get support with doubt clearing sessions, hackathons, resume reviews, and mock interviews.
Course Details:
Start Date: December 21, 2024
Timing: 8:00 AM - 12:00 PM IST (Sat & Sun)
Proficiency Level: Professionals with 2+ years of experience
Mentors: Mayank Aggarwal & Krish Naik
What's Included:
1.5-year access to the course dashboard
Community forum for discussions
Live doubt-clearing sessions
Job referrals (if available)
Resume discussions and mock interviews
Enroll Now: https://learn.krishnaikacademy.com/we...
For any questions, feel free to contact our counseling team at 📞 +919111533440. We’re here to help you succeed!
👉 Download the full syllabus: https://bit.ly/3OqqN78
Don't miss out on this incredible opportunity to accelerate your career in Big Data Engineering with the power of Azure and AWS Cloud! 🌐
#BigData #Azure #AWS #CloudComputing #DataEngineer #DataEngineering #BigDataBootcamp #CloudBootcamp #DataScience #TechCareer #KrishNaik #MayankAggarwal
Playlist
Uploads from Krish Naik · Krish Naik · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Natural Language Processing|Stemming
Krish Naik
Natural Language Processing|BagofWords
Krish Naik
Gaussian distribution or Normal Distribution in statisctics
Krish Naik
Natural Language Processing|TF-IDF for Machine Learning| Text Prerocessing
Krish Naik
Log Normal Distribution in Statistics
Krish Naik
Covariance in Statistics
Krish Naik
Confusion matrix, Precision, Recall| Data Science Interview questions
Krish Naik
Tutorial 44-Balanced vs Imbalanced Dataset and how to handle Imbalanced Dataset
Krish Naik
Implementing a Spam classifier in python| Natural Language Processing
Krish Naik
Tutorial 11-Exploratory Data Analysis(EDA) of Titanic dataset
Krish Naik
Face Recognition using open CV and VGG 16 Transfer Learning
Krish Naik
Pedestrian Detection using OpenCV from Videos
Krish Naik
Face and Eye Detection from Videos using HAAR Cascade Classifier
Krish Naik
Reading, Writing and Displaying images with Opencv| OpenCV Tutorial
Krish Naik
OpenCV Installation | OpenCV tutorial
Krish Naik
Face and Eye Detection from Images using HAAR Cascade Classifier
Krish Naik
Car Detection using HAAR Cascade and Opencv from Videos.
Krish Naik
Using OpenFace for Face recognition in Keras
Krish Naik
OpenPose Tutorial with Tensorflow
Krish Naik
Multiple Linear Regression using python and sklearn
Krish Naik
Dimensional Reduction| Principal Component Analysis
Krish Naik
Movie Recommender System using Python
Krish Naik
TPR,FPR,FNR,TNR, Confusion Matrix
Krish Naik
Precision, Recall and F1-Score
Krish Naik
Artificial Neural Network for Customer's Exit Prediction from Bank
Krish Naik
GridSearchCV- Select the best hyperparameter for any Classification Model
Krish Naik
RandomizedSearchCV- Select the best hyperparameter for any Classification Model
Krish Naik
K Nearest Neighbor classification with Intuition and practical solution
Krish Naik
K Means Clustering Intuition
Krish Naik
Create custom Alexa Skill- Lambda function- Part2
Krish Naik
Hierarchical Clustering intuition
Krish Naik
Implement Transfer Learning with a generic Code Template
Krish Naik
Gender Classifier and Age Estimator using Resnet Convolution Neural Network
Krish Naik
Unlock Your Application With Your Face using OpenCV
Krish Naik
Draw rectangle from webcam and sketch process it on a live feed
Krish Naik
Complete Life Cycle of a Data Science Project
Krish Naik
How we can apply Machine Learning in Finance
Krish Naik
Deep Learning in Medical Science
Krish Naik
How to switch your career to Data Science.
Krish Naik
Linear Regression Mathematical Intuition
Krish Naik
Handle Categorical features using Python
Krish Naik
Machine Learning Algorithm- Which one to choose for your Problem?
Krish Naik
DBSCAN Clustering Easily Explained with Implementation
Krish Naik
Curse of Dimensionality Easily explained| Machine Learning
Krish Naik
Feature Selection Techniques Easily Explained | Machine Learning
Krish Naik
Tutorial 29-R square and Adjusted R square Clearly Explained| Machine Learning
Krish Naik
Cross Validation using sklearn and python | Machine Learning
Krish Naik
Handling Missing Data Easily Explained| Machine Learning
Krish Naik
Deploy Machine Learning Model using Flask
Krish Naik
Deployment of Deep Learning Model using Flask
Krish Naik
How to Visualize Multiple Linear Regression in python
Krish Naik
K Nearest Neighbour Easily Explained with Implementation
Krish Naik
Predicting Heart Disease using Machine Learning
Krish Naik
Predicting Lungs Disease using Deep Learning
Krish Naik
Stock Sentiment Analysis using News Headlines
Krish Naik
Random Forest(Bootstrap Aggregation) Easily Explained
Krish Naik
Voting Classifier(Hard Voting and Soft Voting Classifier)
Krish Naik
Credit Card Fraud Detection using Machine Learning from Kaggle
Krish Naik
Hyperparameter Optimization for Xgboost
Krish Naik
Tutorial 45-Handling imbalanced Dataset using python- Part 1
Krish Naik
Related Reads
📰
📰
📰
📰
Excavating Legacy ETL: The AI Never Asserts a Fact It Could Look Up
Dev.to AI
Announcing Orchestra and n8n | The ultimate way to automate workflows
Medium · Data Science
ELT is moving back to best-of-breed and Orchestration is the missing piece
Medium · Data Science
Azure Data Engineer Course in Telugu: Build a Successful Data Engineering Career
Medium · DevOps
🎓
Tutor Explanation
DeepCamp AI