Data Engineer Full Course in 10 Hours | Data Engineer Course For Beginners | Edureka

edureka! · Beginner ·🔄 Data Engineering ·1y ago

Key Takeaways

This video teaches data engineering skills for beginners, covering the full course in 10 hours

Full Transcript

[Music] hello everyone and welcome to this video on data engineer full course by edu Rea data engineering involves designing building and maintaining the infrastructure for data collection storage and processing data Engineers create robust pipelines to extract transform and load data ensuring its quality and accessibility they work with databases Big Data techn Technologies like Hardo and Spark and data warehousing Solutions such as Amazon red shift and Google big query by streamlining data processes and working with teams data Engineers help organizations use data for insights and better decisions with that said now let's outline today's agenda for this data engineer full course we will Begin by exploring data engineering and the steps to become a data engineer next we will provide an overview of Big Data followed by Guidance of becoming Big Data engineer and insights into the associated salary as we move forward we will outline the pathway of becoming an Azure data engineer then we will cover Azure topics such as Azure data Factory Azure database Services Azure SQL database and Azure data link advancing further we will explore Advanced Data modeling using powerbi and Azure including Azure data break as we progress to Advanced topics we will explore an introduction to Hadoop and cover the Hadoop ecosystem which consist of various tools such as hdfs Yan map reduce spark Pig Hive and hbas next we will focus on understanding KFA streams a key Concept in data engineering to conclude this course we will discuss essential Hadoop interview questions and answers to help you advance in your data engineering career before we begin please consider subscribing to our YouTube channel and hit the Bell icon to stay updated on the latest tech content from edureka also visit the Ed a website for Microsoft certified as your data engineer associate dp23 the link to which is in the description box below without any further Ado now let's dive into our first topic what is data engineering hey have I told you about the incredible data engineering journey of a multinational e-commerce company no I don't think so what happened Alex well let me share this story with you Bob this company had lots of valuable data but struggled with SC systems and outdated databases oh that's a tough situation what did they do about it Alex they embarked on a data engineering initiative BB it was quite a journey I can imagine so what changes did they make Alex they integrated their data into a centralized system and automated data cleaning and transformation it made a huge difference Bob that sounds so promising did it impact their operations Alex definitely Bob they gained faster assess to accurate data and generated realtime analytical reports impressive did they address data governance as well Alex yes they implemented data governance practices for data quality privacy and compilance that's commendable it's a great example of how data insuring can transform a business what an inspiring story Alex yeah I thought you would find it fascinating Bob data insing chly has the power to drive success and Innovation together yeah thanks for sharing the story Alex it reinforces the importance of data Engineering in today's data driven World hello everyone this is Saia from edua and in this video we will be diving into the fascinating world of data engineering so data engineering is all about harnessing the power of data to drive meaningful insights and today's rapidly evolving digital landscape organizations are grappling with massive amounts of data and that's where data and sharing comes into play now let's take a quick look at the agenda for this video we will start by the introduction of what data insuring is and why it is important for so many businesses outside then we will delve into the key components of data engineering such as data inje data integration data transformation and data storage next we will discuss some of the responsibilities of data engineers and the comparison between data engineer data analyst and data scientist after that we will delve into the installation process of popular tools and Technologies which is used in data engineering finally we will will wrap up with a glimpse of what data pipelines are so let's understand what is data engineering First Data engineering involves the design development and maintenance of system and infrastructure to handle large volumes of data effectively it focuses on the extraction transformation loading and storage of data which ensure its quality and scalability of analysis now let's understand what are the key components of data engineering data engineering helps to collect data from the disparate sources and integrate into a unified format which allows organization to have a comprehensive view of their data it also involves processing and transforming data into a suitable format which ensures data quality consistency and applications it includes the development of data pipelines and processing system that enable the efficient processing and Analysis of data this includes batch processing for large scale data transformation and realtime processing for immediate insights and actions data engineering plays a crucial role in enabling data driven World by providing clean integrated and accessible data data insuring empowers organizations to make a informed decision based on accurate insights and Analysis overall it is used to handle the complexity of data processing storage and integration which ensures that the organization can leverage the full potential of their data assets for strategic operations and Innovations now after the thorough understanding of what what data engineering is and its key features also we will delve into the importance of data engineering so here are some of the key reasons why data engineering should be important or why data engineering is important so our first reason is it focuses on building scalable and efficient data processing systems by optimizing data pipelines leveraging distributed computing Technologies and employing the performance tuning techniques data Engineers ensure did the organization can handle large volumes of data and Achieve faster processing times it provides the foundation for advanced analytics and machine learning initiatives by structuring and preparing data in a switchable format data Engineers enable data scientists and analyst to extract valuable insights build predictive models and develop machine learning algorithms data insuring enables organization to process and analyze data as it arrives this is crucial in scenarios such as fraud detection recommendation systems iot applications and monitoring systems that require immediate insights and actions as we already know the importance of data engineering we will cover the real word applications based on this so our first application is e-commerce sites so data insuring is used to collect and process large volumes of customer data transaction data and product data this enables e-commerce companies to predict recommend recommendations optimize pricing strategies and improve Inventory management our second example is social media sites like data insuring plays a crucial role in collecting processing and analyzing social media data which allows companies to monitor brand sentiment track user interactions and derive insights for targeted marketing campaigns so our third example is in finance and banking sector data insuring is used to handle financial data including transaction records customer information and Market data it enables fraud detection risk assessment algorithmic trading and personalized Financial Services also okay so now the fourth application is in the healthcare sector data insuring is employed to manage and analyze patient records Medical Imaging data and clinical child data it enables Healthcare Providers and researchers to gain insights improve patient outcomes and develop predictive models so this examples highlight diverse application and the importance of data Engineering in various Industries demonstrating how it enables organization to Leverage The Power of data for operation and Innovation now the data insuring process involves several key steps including data injection transformation storage processing and integration let's discuss each step in more detail so our first step is data inje data inje refers to the process of collecting and importing data from various sources process into a data system or data pipeline this can involve extracting data from databases files apis streaming platforms or the other resources the goal is to GA relevant data and make it available for the further processing and Analysis Second Step would be the data transformation once the data is ingested it often needs to be transformed into a suitable format for analysis or storage this data transformation involves cleaning validating and restructuring the dat data to ensure consistency and usability this step may include tasks such as data filtering aggregation normalization data type conversion or the application of business rules so our third step would be data storage so after the data is transformed it needs to be stored in a structured manner right so this typically involves using databases or data storage system that provide efficient storage and retrial capabilities popular choices for data storage include relational databases data where warehouses data laks or distributed file systems the selection depends on the specific requirements of the project such as data volume SS patterns and the analytical records so our next step would be data processing data processing involves performing computations and analysis on the stored data this step can include tasks such as data aggregation data Improvement data summarization statistical calculation or machine learning algorithms data processing can be done through various tools such such as dql queries data processing engines or the custom scripts Next Step would be the data integration so in this step data integration involves combining data from multiple sources to create a unified and comprehensive view this step is crucial when dealing with heterogeneous data sources or when different style produced data this needs to be Consolidated data integration can be achieved through data consolidation or by using extract transform load process to combine and merge data from various sources next our last step would be the data governance so data governance is nothing but a framework in a set of process which ensures the effective management and utilization of data assets within an organization it involves establishing policies procedures and guidelines for data management data quality and data usage all the steps and data Engineering Process are typically iterative and may require continuous monitoring optimization and maintenance to ensure data quality reliability and performance data Engineers play a crucial role in designing in implementing efficient and scalable data pipelines to support datadriven applications and analytics Now we move ahead with the key responsibilities of data Engineers which typically includes designing and developing scalable and efficient data pipelines which extract transform and load data from various sources to a Target systems next would be the managing and optimizing data storage infrastructure for performance scalability and reliability selecting and configuring appropriate data storage technology such as relational databases data warehouses or no SQL databases to meet data storage and retrival requirements data Engineers are also responsible for applying data transformation techniques to clean filter agregate and structure data for analysis or consumption by Downstream systems there also Al responsible for implementing data processing task using programming languages or data processing Frameworks to manipulate and transform data efficiently now after analyzing the key responsibilities of data insuring we will proceed with our next topic and understand the key differences between these three roles that is data engineer data analyst and data scientist these three are the distinct role within the field of data science each with its own set of responsibilities and skill requirements so our first profile is of data engineer so data engineer are responsible for Designing building and maintaining the infrastructure and system that enable data storage processing and retrial as we all know they focus on creating and managing the data pipelines and architecture necessary for efficient data collection transformation and storage data Engineers also work closely with software engineers and databas administrator to ensure data is accessible reliable and scalable they typically work with tools like Hardo spark SQL ETL Frameworks and cloud-based platforms for data processing and storage then our next profile would be the data analyst so data analysts are focused on analyzing and interpreting data to derive meaningful insights they work with structured and unstructured data to identify patterns strengths and correlations data analysts are proficient in statistical analysis data visualization tools and data quaring techniques as well they often use tools like SQL Excel table or powerbi to analyze data and create reports and dashboards after that our next profile would be the data scientist data scientist possess a blend of skills from mathematics statistics programming and domain knowledge they leverage their expertise to develop and Implement complex algorithms and models to solve integrate data patterns or extract insights from large data sets data scientists employ techniques like machine learning predictive modeling and statistical analysis to build predictive models uncover patterns and make predictions also they also collaborate with stakeholders to Define business problems and design experiments to together their data so to summarize this data Engineers focus on the infrastructure and data pipelines data analysts work on analyzing and Reporting data while data scientists concentrate on Advanced modeling and extracting insights so to make all this happen data Engineers rely on powerful tool tools and Technologies platforms like Apache spark Apache Kafka SQL and no SQL databases and cloud services such as AWS and gcp provides the building blocks for efficient data engineering this tools help process vast amounts of data facilate realtime data streaming and ensure secure and scalable data storage now let's understand what are the data pipelines are so here is the glimpse of what data pipelines are and why we use data Pipelines so data pipelines are the series of steps that extract transform and load data from source to destination itself okay so they enable the efficient and automated flow of data through different stages ensuring data quality consistency and avability let's understand each step one by one so the first step is extraction the extraction phase of ETL involves retrieving data from various sources such as databases files apis or streaming platforms the goal is to extract the relevant data needed for further processing and Analysis this process typically includes establishing connection to the data sources performing data queries or using data extraction tools to pull the required data into the data pipeline Next Step would be the transformation the transformation phase of ETL focuses on cleaning validating and reshaping the exctic data to ensure its quality and consistency this step involves applying various oper operations and rules to the data such as data filtering data type conversions data aggregation data enrichment or data normalization the transformation process aims to make the data suitable for analysis storage or integration into the destination system Next Step would be the loading the loading pH of ETL involves storing the transformed data into the targeted system such as databases data warehouses or data LS then the data is loaded in a structured format that aligns with the schema or format of the destination system this phase may include tasks such as data mapping schema matching data passing or indexing to optimize data storage and retrial after that our next step would be the or maybe we call this our last step would be the monitoring and handling monitoring and handling refer to the ongoing monitoring management and the maintenance of data pipelines on okay so now let's discuss about batch processing and realtime streaming Pipelines so in a batch processing pipeline there is a delay between time data is collected and when it is processed this delay can range from minutes to hours or even days depending on the schedule intervals on the other hand realtime streaming pipeline aim to process data as it arrives which results in a minimal latency data is processed analyzed in near real time or within a very low delay the second point is batch processing pipelines are designed to handle large volumes of data efficient ly they can process and analyze massive amounts of historical data in a batch mode on the other hand realtime streaming pipelines focus on processing data as it arrives making them more suitable for handling data streams with continuous High Velocity data updates now the third point is batch processing pipelines typically utilize batch processing Frameworks like Apache spark or Hadoop map reduce this Frameworks process data in chunks or batches which allows for parallel process and optimize resource utilizations also while realtime streaming pipelines often use streaming Frameworks like Apache Kafka Apache fling or Apache stom this Frameworks enable continuous processing of data streams supporting low latency operations and realtime analytics now the fourth point is batch processing pipelines are commonly used for tasks that involve historical analysis generating periodic reports or data preparation for machine learning models they are well suited for scenar SC arios where processing time is not critical but analyzing large volumes of data is essential on the other hand realtime streaming pipelines are ideal for application that require real-time monitoring immediate response or instant insights based on live data use cases include like fraud detection realtime recommendation systems network monitoring or iot sensor data analysis so our last point is batch processing pipelines often require significant Computing resources during the processing phase as Z process large volumes of data in a batch mode whereas realtime streaming pipelines also requires Computing resources but are more focus on low latency processing and continuous data streams requiring efficient resources allocation and management so it's important to note that there can be the overlap between batch processing and realtime streaming pipelines and hybrid architectures combining both approaches are common so that wraps up our Deep dive into the world of data engineering we have covered the essential aspects of data pipelines governance and security giving you a comprehensive understanding of how data is ingested transformed stored and processed here is the question for you guys which of the following task is not typically performed by a data engineer and the options are a data cleaning and transformation B data storage and retrieval C data visualization and presentation or D Building data pipelines if you know the correct answer please comment down below in our previous module we have just learned about the essential aspects of data engineering moving forward in how to become a data engineer you'll discover the skills education and career path necessary to embark on a successful data engineering Journey now how to become a data engineer the road map to become a data engineer can go like from being proficient in programming language to learning Automation and scripting then understanding your databases and mastering data processing techniques to studying cloud computing and internalizing infrastructure now we'll see how to get started with becoming a data engineer to get started with the learning process you can look into our Eda YouTube channel to start with even with no prior knowledge of data engineering one can just go through the videos and get an understanding of the whole subject we even have edureka blogs to help you with detailed information on the topics and can give you a clearer picture apart from these we even have premium courses which can help you understand the topic at ease with the personal trainer and 24 hours access to Lifetime content you can learn here at your own pace with a live trainer who is extremely efficient and knowledgeable and experienced in the particular field these courses can even get you certificates which you can add in your CV for better opportunity in your job market so that's it for today I hope this video helped you and will help you decide how to become a data engineer having covered the basics of becoming a data engineer our next topic is Introduction to Big Data here Learners will learn about characteristics sources and significance of Big Data data in various Industries now I feel Sor of it's the best time to tell the story about how data evolved and how big data came fine RMA so we'll move forward so sort of what can you notice here RMA I see how technology has evolved earlier we had landline phones but now we have smartphones we have Android we have IOS that are making our lives smarter as well as our phone smarter apart from that we were also using bulky desktops for processing MBS of data now if you can remember we were using floppies and you know how much data it can store right then came hard dis for storing TBS of data and now we can store data on cloud as well and similarly nowadays even self-driving cars have come up I know you must be thinking why are we telling that now if you notice due to this enhancement of Technology we're generating a lot of data so let's take the example of your phones have you ever noticed how much data is generated due to your fancy smartphones your every action even one video that you send through WhatsApp or any other messenger app that generates data now this is just an example you have no idea how much data you're generating because of every action you do now the deal is this data is not in a format that our relational database can handle and apart from that even the volume of data has also increased exponentially now I was talking about self-driving cars so basically these cars have sensors that records every minutu details like the size of the op obstacle the distance from the obstacle and many more and then it decides how to react now you can imagine how much data is generated for each kilometer that you drive on that car I completely agree with you RMA so let's move forward and focus on various other factors behind the evolution of data I think you guys must have heard about iot if you can recall in the previous slide we were discussing about self-driving cars it is nothing but an example of iot let me tell you what exactly it is iot connects your physical device with internet and makes the device smarter so nowadays if you have noticed we have Smart ACS TVs Etc so we'll take the example of smart air conditioners so this device actually monitors your body temperature and the outside temperature and accordingly decides what should be the temperature of the room now in order to do this it has to First accumulate data from where it can accumulate data from internet through sensors that are monitoring your body temperature and the surroundings so basically from various sources that you might not even know about it is actually fetching that data and accordingly it decides what should be the temperature of your room now we can actually see that because of iot we are generating huge amount of data now there's one startat also that is there in front of your screen so if you notice by 2020 will have 50 billion iot devices so I don't think so I need to explain much that how iot is generating huge amount of data so we'll move forward and focus on one more factor that is social media now when we talk about social media I think RMA can explain this better right RMA yeah sort of but I'm pretty sure that even you use it so let me tell you that social media is actually one of the most important factor in the evolution of big data so nowadays everyone is using Facebook Instagram YouTube and a lot of other social media websites so these social media sites have so much data for example it will have your personal details like your name age and apart from that even each picture that you like or react to also generates data and even the Facebook pages that you go around liking that is also generating data and nowadays you can see that most people are sharing videos on Facebook so that is also generating a huge amount of data and the most challenging part here is that the data is not present in a structured Manner and at the same time it is huge in size isn't that right Sor can't agree more the point you made about the form of data is actually one of the biggest factor for the evolution of big data so due to all these reasons that we have discussed have not only increased the amount of data but it has also shown us that data is actually getting generated in various formats for example data is generated with videos that is actually unstructured same goes for images as well so there are numerous or you can say millions of ways in which data is getting generated nowadays absolutely and these are just few examples that we have given you there are many other driving factors for the evolution of data so these are few more examples because of which data is evolving and converting to Big Data we'll discuss about the retail part I'm pretty sure that all of you must have visited websites like Amazon flip cart Etc and rishma I know you visited a lot of times yeah I do and suppose RMA wants to buy shoes so she won't just directly go buy shoes she'll search for a lot of shoes so somewhere her search history will be stored and I know for sure that this won't be the first time that she's buying something so there will be her purchase history as well along with her personal details and there are numerous ways in which she might not even know that she's generating data and obviously Amazon was not present earlier so at that time there is no way that such huge amount of data was generated similarly the data has evolved due to other reasons as well like Banking and finance media and entertainment etc etc so now the deal is what exactly is Big Data how do we consider a data as big data so let's move forward and understand what exactly it is okay now let us look at the proper definition of Big Data even though we've put forward our own definitions already so sort of why don't you take us through it yes RMA sure so big data is a term for collection of data sets so large and complex that it becomes difficult to process using onhand database system tools or traditional data processing applications okay so what I understand from this is that our traditional systems are a problem because they're too oldfashioned to process the this data or something no RMA the real problem is there is too much data to process when the traditional systems were invented in the beginning we never anticipated that we would have to deal with such enormous amount of the data it's like a disease infected on you you don't change your body orientation when you get infected with a disease right RMA you cure it with medicines couldn't agree more sort of now the question is how do we consider some data as Big Data how do we classify some data as big data how do we know which kind of data is going to be hard for us to process well Sor of we have the 5 vs to tell us that so let's take a closer look at what are those so starting with the first V it's the volume of data it's tremendously large so if you look at the stats here you can see the volume of data is rising exponentially so now we're dealing with just 4.4 zettabytes of data and by 2020 just in 3 years it is expected that the data will rise up to 44 zettabytes which is like equal to 44 trillion gigabytes so that's really really huge it is because all these humongous all this humongous data is coming from multiple sources and that is the second V which is nothing but variety we deal with so many different kinds of files at all once there are MP3 files videos Json CSV tsv and many more now these are all structured unstructured and semi-structured all together now let me explain you the this with the diagram that is there on your screen so over here we have audio we have video files we have PNG files we have Json log files emails various formats of data now this data is classified into three forms one is structured format now in structured format you have a proper schema for your data so you know what all columns will be there and basically you know the schema about your data so it is structured it is in a structured format or you can say in a tabular format now when we talk about semi-structured files these are nothing but Json XML and CS files where schema is not defined properly now when I go to unstructured format we have log files here audio files videos and images so these are all considered as unstructured files and Sor it is also because of the speed of accumulation of all this variety of data Al together which brings us to our third V which is velocity so if you look here earlier we were using Mainframe systems huge computers but less data because there were less people working with computers at that time but as computers evolved and we came to the client server model the time came for the web applications and the internet boomed and as it grew among the masses the web applications got increased over the internet and everyone started using all these applications and not only from their computers and also for mobile devices so more users more appliances more apps and hence a lot of data and when you talk about people generating data or Internet RMA the one kind of application that strikes first in my mind is social media so you tell me how much data you generate alone with your Instagram post and stories uh it will be quite a boast if I only talk about myself here so let's talk including every social media user so if you see the stats in front of your screen you can see that for every 60 seconds there are 100,000 tweets actually more than 100,000 tweets generated in Twitter every minute similarly there are 695,000 status updates on Facebook when you talk about messaging there are 11 million messages generated every minute and similarly there are 698,000 45 Google searches 168 million emails and that equals to almost 1,820 terabytes of data and obviously the number of mobile users are also increasing every minute and there are 2117 plus new mobile users every 60 Seconds gez that's a lot of data I don't even want to go ahead and calculate the total it would actually scare me yeah that's a lot now the bigger problem is how to extract the useful data from here and that's when we come to our next we that is value so over here what happens first you need to mine the useful content from your data basically you need to make sure that you have only useful fields in your data set after that you perform certain analytics or you say you you perform certain analysis on that data that you have cleaned and you need to make sure that whatever analysis you have done it is of some value that is it will help you in your business to grow it can basically find out certain insights which were not possible earlier so you need to make sure that whatever big data that has been generated or whatever data that has been generated it makes sense it will actually help your business to grow and it has some value to it now getting the value out of this data is one big challenge let me tell you why and that brings us to our next V which is veracity now this big data has a lot of inconsistencies obviously when you're dumping such huge amount of data some data packets are bound to lose in the process now what we need to do we need to fill up these missing data and then start mining again and then process it and then come up with a good Insight if possible so if you can notice there's a diagram in front of your screen so over here we have this field which is not defined similarly this field and if you can notice here when we talk about this minimum value you see the other minimum values and when you talk about this it is it is way more than the other fields present in this particular column similarly goes for this particular element as well okay so obviously processing data like this is one problematic thing and now I get it why big data is a problem statement well we have only five vs now but maybe later on we'll have more so there are good chances that big data might be even more big okay so there are a lot of problems in dealing with big data but there are always different ways to look at at some things so let us get some positivity in the environment now and let us understand how can we use Big Data as an opportunity yes RMA and I would say the situation is similar to the proverb when life throws you lemons make lemonade yes so let us go through the fields where we can use Big Data as a boon and there are certain unknown problems solved only because we started dealing with big data and the boond that you're talking about rishma is big data analytics first thing with big data we figured out how to store our data cost effectively we were spending too much money on storage before until Big Data came into the picture we never thought of using commodity Hardware to store and manage a data which is both reliable and feasible as compared to the costly servers now let me give you a few examples in order to show you how important big data analytics is nowadays so when you go to a website like Amazon or YouTube or Pandora Netflix any other website so they'll actually provide you certain Fe in which they'll recommend some products or some videos or some movies or some songs for you right so how do you think they do that so basically whatever data that you are generating on these kind of websites they make sure that they analyze it properly and let me tell you guys that data is not small it is actually big data now they analyze that big data and they make sure that whatever you like or whatever your preferences are accordingly they'll generate recommendations for you and when I go to YouTube I don't know if you guys have noticed it but I'm pretty sure you must have done that so when I go to YouTube YouTube knows what song or what video that I want to watch next similarly Netflix knows what kind of movies are like and when I go to Amazon it actually shows me what all products that I would prefer to buy right so how do you think it happens it happens only because of big data analytics okay so there is one more example that just popped into my mind I'll share with you guys so there was this time when the Hurricane Sandy was about to hit on New Jersey in United States so what happened then the Walmart used big data analytics to profit from it now I'll tell you how they did it so what Walmart did is that they studied the purchase patterns of different customers when a hurricane is about to strike or any kind of natural Calamity is about to strike on a particular area and when they made an analysis of it so they found out that people tend to buy emergency stuff like flashlight life jackets and a little bit of other stuff and interestingly people also buy a lot of strawberry poptart strawberry poptarts are you serious yeah now I didn't do that analysis so I Walmart did that and apparently it is true so what they did is so they stuffed all their stores with a lot of strawberry Pop-Tarts and emergency stuff and obviously it was sold out and they earned a lot of money during that time but my question here RMA is people want to die eating strawberry poptarts like what was the idea idea behind strawberry poptarts I'm pretty unsure about it but yeah since you have given us a very interesting example and Walmart did that analysis we didn't do it so yeah so it is a very good example in order to understand how big data analytics can help your business to grow and find better insights from the data that you have yeah and also if you want to know why strawberry poptarts maybe later on we can start making an analysis by gathering some more data also yeah that can be possible okay so now let's move ahead and take a look at had a case study by IBM how they have used big data analytics to profit their company so if you have noticed that earlier the data that was collected from The Meters that you have in your home that measures the electricity consumed it is actually sending data after 1 month but nowadays what IBM did they came up with this thing called smart meter and that smart meter used to collect data after every 15 minutes so whatever energy that you have consumed after every 15 minutes it will send that data and because of it Big Data was generated so we have some stats here which says that we have 96 million reads per day for every million meters which is pretty huge this data the amount of data that is generated is pretty huge now IBM actually realized the data that they're generating it is very important for them to gain something from that data so for that what they need to for that what they need to do they need to make sure that they analyze this data so they realize that big data analytics can solve a lot of problems and they can get better business inside through that so let us move forward and see what type of analysis they did on that data so before analyzing that data they came to know that energy utilization and billing was only increasing now after analyzing Big Data they came to know that during Peak load the users require more energy and during off peak times that users require less energy so what advantage they must have got from this analysis one thing that I can think of right now is they can tell the industries to use their Machinery only during the off weak times so that the load will be pretty much balanced and you can even say that time of use pricing encourages cost savy retail like industrial heavy machines to be used off peak time so yeah they can save money as well because off peak times pricing will be less than the peak time prices right so this is just one analysis now let us move forward and see the IBM Suite that they developed so over here what happens you first dump all your data that you get in this data warehouse after that it is very important to make sure that your user data is secure then what happens you need to clean that data as I've told you earlier as well there might be many feeds that you don't require so you need to make sure that you have only useful material or useful data in your data set and then you perform certain analysis and in order to use this Suite that IBM offered you efficiently you have to take care of a few things the first thing is that you have to be able to manage the smart meter data now there is a lot of data coming from all this million Smart Meters so you have to be able to manage that large volume of data and also be able to retain it because maybe later on you might need it for some kind of regulatory requirements or something and next thing you should keep in mind is to monitor the distribution grid so that you can improve and optimize the overall grid reliability so that you can identify the abnormal conditions which are causing any kind of problem and then you also have to take care of optimizing the unit commitment so by optimizing the unit commitment the companies can satisfy their customers even more they can reduce the power outages that is they can reduce the power outages so that their customers don't get angry more identify problems and then reduce it obviously and then you have also to optimize the energy trading so it means that you can advise your customers when they should use their appliances in order to maintain that balance in the power load and then you also have to forecast and schedule loads so companies must be able to predict when they can profitably sell the Excess power and when they need to hedge the supply and continuing from this now let's talk about how Encore have made use of the IBM solution so Encore is an electric delivery company and it is the largest electrical distribution and transmission company in Texas and it is one of the sixth largest in the United States they have more than 3 million customers and their service area covers all almost 117,000 square miles and they began the advanced speeder program in 2008 and they have deployed almost 3.25 million meters serving customers of North and Central Texas so when they were implementing it they kept three things in mind the first thing was that it should be instrumented so this solution utilizes the smart electricity meters so that they can accurately measure the electricity usage of a household in every 15 minutes because like we discussed that the smart meters were sending out data every 15 minutes and it provided data inputs that is essential for consumption insights next thing is that it should be interconnected so now the customers have access to the detailed information about the electricity they are consuming and it creates a very Enterprise wide view of all the meter assets and it helped them to improve the service delivery the next thing is to make make your customers intelligent now since it is getting monitored already about how each of the household or each customer is consuming the power so now they're able to advise the customers about maybe to tell them to wash their clothes at night because they're using a lot of appliances during the daytime so maybe they could divide it up so that they can use some appliances at off peak hours so that they can even save more money and this is beneficial for both of them for both the customers and the company as well and they have gained a lot of benefits by using the IBM solution so what are the benefits they got is that it enabled onore to identify and fix outages before the customers get inconvenience that means they were able to identify the problem before it even occurred and it also improved the emergency response on events of severe weather events and views of outages and it also provides the customers the data needed to become a active participant in the power consumption management and it enabled every individual household to reduce their electrical consumption by almost 5 to 10% and this is how onore used the IBM solution and made huge benefits out of it just by using big data analytics that IBM performed but let me just interrupt right now so since RMA told us in the beginning as well that there are no free lunches in life right so this is an opportunity but there are many problems to encase this opportunity right so let's focus on those problems one by one so the first problem is storing colossal amount of data so let's discuss few starts that are there in front of your screen so data generated in past 2 years is more than the previous history in total so guys what are we doing stop generating so much amount of data and it said that by 2020 total Digital Data will grow to 44 Zab bytes approximately and there's one more start that amazes me is about 1.7 MB of new information will be created every second for every person by 2020 so storing this huge data in traditional system is not possible the reason is obvious the storage will be limited for one system for example you have a server with a storage limit of 10 tabt but your company is growing really fast and data is exponentially increasing now what you'll do now at one point you'll exhaust all the storage so investing in huge servers is definitely not a cost effective solution so what do you think what can be the solution to this problem uh according to me a distributed file system will be a better way to store this huge data because with this we'll be uh saving a lot of money let me tell you how because due to this distributed system you can actually store your data in commodity Hardware instead of spending money on high-end servers don't you agree sov completely now we know storing is a problem but let me tell you guys it is just one part of the problem let's see if few more okay so since we saw that the data is not only huge but it is present in various formats as well like unstructured semi-structured and structured so you not only need to store this huge data but you also need to make sure that a system is present to store this varieties of data generated from various sources and now let's focus on the next problem now let's focus on the diagram so over here you can notice that the hard disk capacity is in increasing but the disc transfer performance or speed is not increasing at that rate let me explain you this with an example if you have only 100 MVPs input output Channel and you are processing say 1 terabytes of data now how much time will it take maybe calculate it'll be somewhere around 2.91 hours right so it'll be somewhere around 2.91 hours and I've have taken an example of 1 terabytes what if you're processing some Zeta bytes of data so you can imagine how much time will it take now what if you have four input output channels for the same amount of data then it'll take approximately 72 hours or convert it to minutes so it'll be around 43 minutes approximately right and now imagine instead of one TB you have Zab bytes of data for me more than storage accessing and processing speed for huge data is a bigger problem okay so RMA has a very good example to discuss yeah so since you were talking about accessing the data and you told us already about how Amazon and different websites and YouTube they make those recommendations so if there was no solution for it if it would take so much time to access the data the recommendation system wouldn't work at all and they make a lot of money just by recommendation system because a lot of people go there and click over there and buy that product right so let's consider that that it is taking like hours or maybe years of time in order to process my that big amount of data data so let's say that at one time I purchased an iPhone 5s from Amazon and after 2 years I'm again browsing onto Amazon and since it took so much time to access the data and I've already switched over to a new iPhone and they are recommending me the old iPhone case for 5S so obviously that won't work I won't go there and click it because I've already changed my phone right so that will be a huge problem for Amazon the recommendation system work anymore and I know that RMA changes her phone every year so if she has bought a phone and people are recommending if she has bought a phone now and someone's recommending the case for that phone after two years doesn't make sense to me at all yeah only it'll work if I have both the two phones at the same time but yeah I don't want to waste money on purchasing new iPhone case for my old phone so basically it won't be fair if we don't discuss the solution to these problems RMA we can't leave our viewers with just the problems right it won't be fair what is the solution Hadoop Hadoop is a solution so let's introduce Hadoop now okay so now what is Hadoop so Hadoop is a framework that allows you to first store big data in a distributed environment so that you can process it parallell there are basically two parts one is hdfs that is Hadoop distributed file system for storage it allow allows you to store data of various formats across a cluster and the second part is map reduce now it is nothing but a processing unit of Hadoop it allows parallel processing of data that is stored across the hdfs now let us dig deep in hdfs and understand it better yeah so hdfs creates an abstraction of resources um let me simplify it for you so similar to virtualization you can see hdfs logically as a single unit for storing big data but actually you're storing your data across multiple systems or you can say in a distributed fashion so here you have a Master Slave architecture in which the name node is a master node and the data nodes are slaves and the name node contains the metadata about the data that is stored in the data nodes like which data block is stored in which data node where are the replications of the data block kept and etc etc so the actual data is stored in the data nodes and I also want to add that we actually replicate the data blocks that is present in the data nodes and by default the replication factor is three so it means that there are three copies of each file so s going to tell us why do we need that replication sure RMA since we are using commodity Hardwares right and we know failure rate of these Hardwares are pretty high so if one of the data nodes fail I won't have that data block and that's the reason we need to replicate the data block BL now this replication Factor depends on your requirements right now let us understand how actually Hadoop provided the solution to the big data problems that we have discussed so RMA can you remember what was the first problem yeah it was storing the big data so how hdfs solved it let's discuss it so hdfs provides a distributed way to store Big Data we've already told you that so your data is stored in blocks in data noes and you then specify the size of each each block so basically if you have a 512 MB of data and you have configured hdfs such that it will create 128 megabytes of data block so hdfs will so hdfs will divide the data in four blocks because 5 in2 divide by 128 is four and it will store it across different data nodes and it will also replicate the data blogs on the different data notes so now we are using commodity hardware and storing is not a challenge so what are your thoughts on it s I will also add one thing RMA it also solves the scaling problem it focuses on horizontal scaling instead of vertical now you can always add some extra data nodes to your hdfs cluster as and when required instead of scaling the resources of your data nodes so you're not actually increasing the resources of your data nodes you're just adding few more data nodes when you require let me summarize it for you so basically for storing one TB of data I don't need a one TB system I can instead do it on multip multiple 128 GB systems or even less now RMA what was the second challenge with big data so the next problem was storing variety of data and that problem was also addressed by hdfs so with hdfs you can store all kinds of data whether it's structured semi-structured or unstructured it is because in hdfs there is no pre- dumping scheme of validation so you can just dump all the kinds of data that you have in one place and it also follows a right one and read many model and due to this you can just write the data once and you can read it multiple times for finding out insights and if you can recall the third challenge was accessing the data faster and this is one of the major challenge with big data and in order to solve it we're moving processing to data and not data to processing so what it means s of just go ahead and explain it yes RMA I will so over here let me explain you what do you mean by actually moving process to data so consider this as our master and these are our slaves so the data is stored in these slaves so what happens one way of processing this data is what I can do is I can send this data to my master node and I can process it over here but what will happen if all of my slaves will send the data to my master node it'll cause Network congestion plus input output Channel congestion and at the same time my master node will take a lot of time in order to process this huge amount of data so what I can do I can send this process to data that means I can send the logic to all these slaves which actually contain the data and perform processing in the slaves itself so after that what will happen the small chunks of the result that will come out will be sent to our name node so in that way there won't be any network congestion or input output congestion and it will take comparatively very less time so this is what actually means sending process to data in the previous module we introduced the concept of big data now in how to become a big data engineer you will gain insights into qualifications skills and the steps required to specialize in handling and processing large data sets so who's a big data engineer now every datadriven business needs to have a framework in place for the data science and data analytics Pipeline and a data engineer is the one who's responsible for building and maintaining this framework now these Engineers must Ensure that there is an uninterrupted flow of data between servers and applications so in simple words a data engineer builds tests maintains data structures and architectures for data ingestion processing and deployment of large-scale data intensive applications now data Engineers work in tandem with data Architects data analysts and data scientists so they must all share these insights to other stakeholders in the company through data visualization and storytelling but what does a big data engineer do exactly now the most crucial part of a big data engineer is to design develop construct install test and maintain the complete data management and processing systems they are basically the ones who handle the complete endtoend infrastructure for data management and processing they build a pipeline for data collection and storage and funnel the data to data analysts and scientists so basically what they do is they create the framework work to make data consumable for data scientists and analysts so they can use the data to derive insights from it note that the data Engineers are the Builders of data systems and not those who mine for insights so the data engineer Works more behind the scenes and must be comfortable with other members of the team producing Business Solutions from this data now all their responsibilities revolve around this they need to take care of a lot of things while performing these activities hence one of the most sought-after skills in data engineering is the ability to design and build data warehouses this is where all the raw data is collected stored and retrieved from without data warehouses all the tasks that a data scientist does will become obsolete it is either going to get too expensive or very very large to scale now data Engineers should always keep in mind that the system which he or she builds needs to be scalable robust and fault tolerant so that the system can be scaled up without increasing the number of data sources and can handle a huge amount of heterogeneous data without any failure now imagine a situation wherein the source of data is doubled or tripled but the system cannot scale up will it not cost a lot more time and resources to build the same system again which is suitable for this kind of intake exactly this is why the Big Data Engineers have a role here next he or she is the one that handles the extract transform and load process which is basically the blueprint for how The Collector raw data is processed and transformed into Data ready for analysis now you're going to acquire a lot of data from different sources how do you bring them together to one platform ETL is your answer apart from all this a data engineer should always aim at deriving insights by acquiring data from new sources some of the responsibilities of a data engineer also include improving data foundational procedures integrating new data management Technologies and the software into existing systems and building data collection pipelines and finally one of the major roles of a data engineer is to include performance tuning and make the whole system way more efficient which is pretty self-explanatory if you ask me now most of us have some idea about who a big data engineer is but there's still some confusion about their responsibilities now this ambiguity further increases when we gain more information about the role now let me help you debunk all your queries about it so let's talk about some big data engineer responsibilities first up we have data ingestion now this is associated with the task of getting data out of the source systems and ingesting it into a data Lake now a data engineer would need to know how to efficiently extract the data from a source including multiple approaches for both batch and real-time extraction as well as needing to know about the incremental data loading fitting within small Source windows and parallelization of data loading as well now another small subtask of data ingestion is data synchronization but because it's such a big issue in the Big Data world we are going to talk about it now since Hadoop and other big data platforms don't support incremental loading of data a data engineer would need to know how to deal with detecting changes in the data source merge and sync change data from sources into the Big Data environment next we have data transformation this is basically the T in the extract transform and load that we had discussed earlier it is basically focused on integration and transformation of data for a specific use case now a major skill set here is the knowledge of SQL as it turns out not much has changed in terms of the type of data Transformations that people are doing now compared to purely relational environments now imagine all this data that you've acquired from various sources what would you have to do to make them all palatable in the same platform you need to transform that data and this is what a data engineer does here and finally we have performance optimization which is one of the tougher areas because anyone can build a slow performing system the challenge is to build data Pipelines that are both scalable and efficient so the ability and understanding of how to optimize the performance of an individual data Pipeline and the overall systems are a higher level of data engineering skill now for example Big Data platforms continue to be challenging with regard to query performance and have added complexity to a data engineer's job in order to optimize performance of queries and creation of reports the data engineer needs to know how to denormalize partition and index data models he also needs to understand tools and Concepts regarding in-memory models and olap cubes now let's quickly move ahead and look at the required skills to fulfill these responsibilities now we'll be going through these skills in a clockwise order so starting with big data Frameworks now with the rise of big data in the early 21st century a new framework was born and that is Hadoop all thanks to Doug cutting for introducing this framework it not only stores big data in a distributed manner but also processes the data parallel there are several tools in the Hadoop ecosystem which cater differently for different purposes and Professionals for a big data engineer mastering Big Data tools is a must some of the tools which you will need to Master first of all you have hdfs which is the storage part of Hadoop being the foundation of Hadoop knowledge of hdfs is a must to start working with Hadoop framework next we have yarn which performs resource management by allocating resources to different applications and scheduling jobs Now map ruce is a parallel processing Paradigm which allows data to be processed parallely on top of the hdfs next we have pig and Hive now Hive is a data warehousing tool on top of hdfs which caters to professional from an SQL background to perform analytics on top of hdfs whereas Apache pig is a high level platform which is used for data transformation on top of Hado now hi is generally used by data analyst for creating reports whereas pig is used by researchers for programming both are pretty easy to learn if you already familiar with SQL next we have Flume and scoop Flume is a tool which is used to import unstructured data to hdfs and scoop is used to Import and Export structured data from our dbms now next we have zookeeper which acts as a coordinator among the distributed Services running in a Hadoop environment it basically helps to configure management and synchronize services and finally we have Uzi which is basically a scheduler which binds multiple logical jobs together and helps in accomplishing a complete task next up we have realtime processing Frameworks now real-time processing with quick actions is the need of R either it is a credit card fraud detection system or a recommendation system now imagine if you wanted a red dress today and Amazon decides to suggest Ed to you a month later now wouldn't that be completely useless for you in this case you need realtime processing it is very important for a data engineer to have knowledge of realtime processing Frameworks now Apachi spark is one of the distributed realtime processing Frameworks which is used in the industry rigorously it can be easily integrated with Hadoop leveraging hdfs as well next we have dbms now a database management system stores organizes and and manages a large amount of information within a single software application now data Engineers need to understand the database management system to manage data efficiently and allow users to perform multiple tasks with ease this will help data engineers in improve data sharing data security data access and better data integration with minimized data inconsistencies these are the fundamentals that data Engineers should know prior to building a scalable robust and fall tolerance system next we have SQL based Technologies now there are various relational databases that are used in the industry such as Oracle DB Microsoft SQL Server Etc now data Engineers must have at least the knowledge of one such database now knowing SQL is also a must this structured query language as SQL is also known as used to structure manipulate and manage data stored in relational databases as data Engineers work closely with rdbms's they need to have a strong command on SQL now next we have no SQL Technologies as the requirements of organizations have grown Beyond structured data no SQL databases have been introduced into this environment it can store large volumes of structured semi-structured or unstructured data with quick iteration and agile structure as per application requirements some of the most prominently used databases are hbase Sandra and mongodb now hbas is a column oriented nosql database on top of hdfs which is great for scalable and distributed Big Data stores it is also great for applications with optimized read and range based scan and it provides consistency and partitioning out of capap now Cassandra is a highly scalable database with incremental scalability and the best part about Cassandra is the minimal Administration and no single point of failure it's good for applications with fast and random read and writs it provides available and partitioning out of c and finally we have mongodb which is basically a document oriented nosql database which is a schema free database it gives full index support for high performance and replication for fall tolerance it has a Master Slave sort of architecture and provides CP out of capap it is rigorously used by web app apption and semi-structured data handling next we're going to discuss programming and scripting languages so various programming languages can serve for the same purpose so knowledge of one programming language is enough I'm saying this because the flavor of language may change but the logic Remains the Same if you're a beginner you can go ahead with python as it is an easy language to learn due to its syntax and good Community Support whereas R has a steep learning curve which is developed by statisticians and and it is mostly used by analysts and data scientists the next skill we're going to discuss is an important one it is ETL or data warehousing now data warehousing is very important when it comes to managing a huge amount of data coming in from heterogenous sources where you need to apply extract transform and load now data warehousing is used for analytics and Reporting and is a very very crucial part of every business intelligence solution because this is the part which is going going to take you most time now it is very important for a big data engineer to Master One data warehousing or ETL tool after mastering one it becomes pretty easy to learn new tools and as the fundamentals remain the same now Informatica click View and talent are very well-known tools used in the industry Informatica and talent Open studio are data integration tools with ETL architecture the major benefit of talent is its support from the Big Data Frameworks if you're new to data warehousing and ETL tools I would definitely recommend you start with talent because after learning this any data warehousing tools will become a piece of cake and finally we have our operating systems now intimate knowledge of Unix Linux and Solaris is very helpful as many mathematical tools are going to be based off of these systems due to their unique demands for root access to hardware and operating system functionality above and beyond that of Microsoft's windows or Mac OS now some level of understanding of how to act upon this data is also very valuable for data Engineers for this reason some knowledge of statistical analysis and the basics of data modeling are also hugely valuable knowledge of machine learning in Cloud also will serve as a big plus while machine learning is technically something relegated to a data scientist knowledge in this area is helpful to construct Solutions usable by your cohorts now this knowledge has the added benefit of making you extremely marketable in this space as being able to put on both hats in which case makes you a really formidable tool we have just explored the pathway of becoming a big data engineer next we will look at Big Data engineer salary where you will learn about the earning potential and factors influencing the salaries of Big Data professionals here you can see the job distribution per salary range in India for a data engineer as we you can see people who get paid more than 5 lakhs inom are about 33% 730 perom are 26% 870,000 about 20% Which are very high salary brackets apart from that the average salary for a data engineer is almost 8 lakhs in anim and for a senior data engineer is almost 16 lakhs in anim if we look at the same numbers in the US there are 32% of professionals who make more than $90,000 a year and 27% professionals who make $105,000 a year the average salary in the US for a data engineer is way more than $90,000 and for a senior data engineer it is $124,000 per anom now as we have discussed the salary of a big data engineer let's look at a few factors in the form of skills and technology that they know on which their salary depends here we've carefully curated a table which lists out the skills and the average salary which can be encashed through them you can see services such as AWS data analysis data mining warehousing machine learning and even programming languages like Java and R apart from that you can see bi tools and statistical tools like Tableau database architecture ETL and structured query languages now another influence on the salary is experience because experience is also a very important factor in deciding the Big Data engineer salary the distribution of salary is like so an entry-level data engineer makes about $85,000 a year people who have 5 to8 years of work experience bag nearly $103,000 a year and people who are experi exped I'm talking like 10 years of Industry experience get over $118,000 a year now this salary must be coming from somewhere presenting to you the companies that hire in this job role as you can see there are some very big names like Amazon Google Bosch Microsoft and IBM who hire Big Data Engineers now companies that hire Big Data professionals are companies that are invested in in the future and the worldwide big data market revenues for software and services are projected to increase from $42 billion in 2018 to $103 billion in 2027 attaining a compound annual growth rate of about 10.5% that sort of a growth needs some kind of work after going through multiple job descriptions we found that the Big Data engineer salary has many variables having discussed Big Data engineer salaries we will now move to how to become an Azure data engineer in this session Learners will learn about the skills and certifications needed to specialize in Microsoft's aure platform so let's go ahead and also understand the average salary of a data engineer so in the US the average salary of a data engineer is $133,000 but remember guys this is basically an average salary which basically means that there are people who are above this salary or can even be a person who is below this salary okay but later in this session I'm also going to talk about some of the job description which can tell you the kind of salary that you can earn once you start applying for a data engineer profile in India the salary is around 6.5 lakhs perom or 7 lakh perom and the same goes for an Indian jobs as well that the salary can be higher than this average package or it can be lower than the package as well well now that you have seen the basic or an average scale of a aor data engineer let's go ahead and see the job description of an Azor data engineer all right guys so thinking again about what a data engineer does an Azor data engineer is responsible for Designing implementing and maintaining data management and data processing system on the Microsoft aor cloud platform now they work with large and complex data sets and are responsible for ensuring that data is stored process and secure efficiently and effectively now when you basically Define a job role like that you have many jobs description which are floating in the market right now how can you identify which job description you have to apply to now let's go ahead and understand that so talking about the job description which basically exists in the market guys you will see a job description which is an entry-level job description and then on the next slide I will show a job description which is a mid or a senior level job description all right so heading back to the entry level so if you see the first job description which is an aor data engineer the salary is anywhere around 6 lakh perom to 14 lakhs perom and these are the skills that you require now as I mentioned before these are the expected skills required for an aor data engineer now what are the skills which are expected now you are expected to know no SQL or a cosmo database skill now they are expected to know data Lake data factory data warehouse and strong experience in building pipelines in Azor data or in Azor data Lake then you should be able to analyze and understand complex data you should be able to understand business requirement and actively provides input from data perspective now at the same time if you look at these skill sets these skill sets are all the scales at which you will basically know after you study for or clear the Azor data engineer certification so once you're done with the certification once you are done with the skill sets which is just there in the certification you can easily go ahead and apply for a job that lies in the salary range which is for an aor data engineer now talking about a mid senior level data engineer profile the list is quite long as you can see now you can see that over here apart from all these skill sets a lot of other things are also mentioned here as well for example you should have some four or five plus years of experience in implementing or designing solution using Azor Big Data Technologies then you should have an experience with an Hands-On in Azor data Factory Azor devops aor data Lake storage Etc now you should have a knowledge of Big Data pipeline then design and build modern data pipelines and maintain the data warehouse schematics layouts architecture and relational or non-relational database for data access and advanced analytics now you should also know Java jQuery SQL or Scala or any preferred programming language right so over here what we recommend to our Learners is you should go ahead and Learn Python because although they have mentioned only these programming language but companies are very much flexible if the target profile has any programming experience but a strong one in any of the programming languages right next thing that they expect you to know is Advanced skill using one or most common language for example like python batch Etc so this python will basically serve as a dual purpose that is all scripting language and a programming language as well the next thing that they expect you to know is the ETL process using big data Technologies such as spark Kafka doop and others now the next thing that they expect you to know is the ETL process using big data technology such as spark Kafka hups and others and then you need to understand the data visualization experience using python python here is a plus guys so you need to understand this and then you actually need to learn table or powerbi any one of the tools or technology is a plus point so you either have to have a skill on tblo or a powerbi so this again is an important skill to have and this again coincide with what a data analyst does right because he is also responsible for data visualization to some extent then you have a solid understanding and experience implementing cloud data platform in Microsoft Azor devops so if you're a guy who wants to start off you can start off with the the fresher profile and after having experience in the fresher profile and learn all the skill sets like big data aor these are the skill sets that if you gain you can actually apply to the senior or mid level in particular okay so Guys these are the few job description which are related to the data engineer profile now that we are clear with why who and what are the career opportunities and salary of an aor data engineer we shall see the path towards becoming an aor data engineer so first of all you will have to talk about the different aor storage which are out there now you will have to learn about these like the blob storage table storage file storage and the Q storage now you will have to learn about the relational database options as well which are there in the market such as the SQL database SQL DB Warehouse analystic Services now you will also have to learn about no SQL you will have to learn about big data services in aor like data Lake analytics data Lake storage now you also have to learn about the data factories Azor function stream analytics iot hubs even hubs Etc now apart from that you will also need to learn redis cach and aor search so these are the services that are basically required for you to understand in order to clear the certification and go ahead and become an aor data engineer now these also conside with the job description that we had a look earlier right now all the services which are mentioned there is basically a part of what you have learned in order to correct the certification now apart from this we also recommend going through the open source Services of our dup such as spark hi and the Hadoop itself right now let us go step by step to reach our goal in becoming an aor data engineer so in order to become an aor data engineer you will also need to have a strong foundation in data engineering and cloud computing here are some steps you can take to develop the skills and knowledge needed for a career as an aor data engineer so the first one is learn the basics of data engineering now in order to become an aor data engineer you should first develop a strong foundation in data engineering Concepts such as data modeling data pipelines data processing and data storage now you can learn these Concepts through online courses or books or by working on practical projects now the second is a learning of programming language now as an AER data engineer you will need to proficient in at least one programming languages as I've already mentioned earlier now python is a very popular choice for data engineering but you can also consider learning languages such as Java C or Scala now the next step is you need to learn an SQL now as a data engineer you will be working with a large amount of data if already known that since the name itself as an aor data engineer so you will need to be a proficient in SQL to extract and transform data now the next thing you need to learn is Azor data storage options now as you know aor aor offers a range of data storage options such as Azor blob storage Azor data L Storage Azor Cosmos DB Azor SQL database and Azor signups analytics formally SQL database warehouse now you should familiarize yourself with the features and capabilities of these storage options now the next thing you need to learn is Azor data processing Technologies now Azor provides several Technologies for processing data such as Azor stream analytics Azor data bricks Azor data fact and aor HD insights you should learn how to use these Technologies to build data pipelines for ingestion transforming and processing data so the next thing you need to learn is aor data management and security so you as an aor data engineer should learn about the Azor tools for managing and securing data such as Azor data catalog Azor data Lake security and Azor private link so what is the next step the next step is get certified as you know getting certified is very crucial and it's very important in today's generation because all the organization as I mentioned earlier if I just have to go back in my slide I will show you the job description where they have actually mentioned that you need to actually pass the certification all right this is the certification you required that is a DP 200 and a DP 2011 but as for now you don't need dp2 200 or dp21 you only have to give one exam which I'll be first further talking about it that is the dp23 all right guys so moving ahead again so you need to earn an aort data engineer associate certification by passing the dp23 exam right now what is the next thing you need to do is you need to gain practical experience now the best way to learn aor data engineering is by working on practical projects now we all know this right now you can find data engineering projects on online platform such as keigle or you can work on projects in your own organization or you can enroll with edura and they provide you a tons and tons of Life projects which are developed by the instructor who have already worked as a data engineer in their organization now the next thing you need to learn is continue learning now as a data engineer you will need to keep your skills and knowledge up to date as Technologies and best practices evolve right now making sure to stay current by learning about new aor data engineering features and participating in professionals development activities all right guys so now let us understand the things that you need to know about this certification that I've just talked about that is your dp23 so guys if you're applying for an aor data engineer associate there used to be two exam that you have to give which I've just mentioned now on the screen in the previous slide right one is implementing an Azor data solution and the next exam is design an aor data solution now after clearing both these exam it will give you the aor data engineering associate certificate if ation and these two exam basically have the code like I've mentioned dp2 200 and DP 2011 right but now this exam have been retired on 23rd February 2021 now you guys just have to give one exam to clear the data engineer certification and this exam is dp23 now what they have done is they have Club both these exam and they have now included its syllabus in just one examine they're asking questions from it right so earlier what you have to do is you have to pay for two exams then prepare for two exams and give them and then only you could get the certification but now just by passing one exam you can clear out the data engineer certification right now as part of the new exams now the skill sets have been updated as you can see on the screen so basically these are the distribution of topics and they'll be covered in this exam now first of all you will be asked most of the questions on design and Implement data storage now this is going to have 40 to 45% of weightage okay now I'm going to tell you what are the topics that comes under the design and Implement data storage so make sure you write it down okay guys so the first is design a data storage structure then the second is design a partitioning strategy then the next question is design the serving layer then you have the Implement physical data storage structures then Implement logical data structure and lastly implement the serving layer now after this you will have have design and develop data processing now this is of 25 to 30% weightage now there are four points coming under this design and develop data processing now the first is you need to learn about the ingest and transform data second is design and develop a batch processing solution the third is design and develop a stream processing solution and lastly we have the manage batches and pipelines all right so coming to the third is design and Implement data security which is of 10 to 15% so there are two main topics here that is design security for data policies and standards and the second is Implement data security Now talking about the last is you have the Monitor and optimize data storage and data processing which is again of 10 to 15% now even here also we have two main topics that you need to be covered that is the monitor data storage and data processing and the second is optimize and troubleshoot data storage and data processing now if you are with me at this point you shall have understood what are all the things that you have to learn in order to clear the exams and become an aor data engineer right since we have already discussed what are the path and what are the skills to learn to prepare ourself so like I said earlier you now just have to give this dp23 exam to get the Microsoft certification associate data engineering certification okay so just one exam and now you will get the certification for it all right so now let's move on guys and now let's talk about how you guys can get started in clearing this exam and developing those skills and go ahead and apply for the job and become a successful zor data engineer so we have mentioned a lot of things that you have to learn but how exactly you should go forward and start learning these things let's go ahead and clear that out so first of all guys what we can do for you is you can basically refer to a lot of blocks that we have written on or we frequently update videos on YouTube as well such as this video which has info about the engineering certification or how to become an aor data engineer now we frequently put more videos for such topics now you can go through them and basically get a jump start into how you can prepare now my recommendation to you will be to plan out your working plan hours after or before working shift of yours now you should at least spend 3 to 4 hours every day for the next two or 3 months in order to clear this certification exam otherwise if you don't invest this much amount of time guys it is going to be very difficult to crack the exam because there are a lot of things to learn especially if you not from a data engineer domain now it is going to be a little difficult because you will have to read documentation you'll have to do Hands-On right now at some point you will get stuck and you'll have some issues you have to figure out what went wrong or what's wrong with the handson I know this because I've also been there guys but getting stuck is the most beautiful part of learning anything because that is where you start the actual research and about how things work right so spend at least 2 to three hours every day either after your workshift or before your workshift to basically learn these Technologies right and now for those people who feel like they do not have the time or they do not want to invest time in researching and they want someone to help them out in getting this exam cleared and become a successful aort data engineer so guys we are a director also of a course of Microsoft aor certification training of the Azor data engineer associate certification course and we also have a master program as well all right which will basically help you in clearing the aor data engineer associate certification right now if you need a helping hand and you need someone or you need to be taught by someone who's already cleared this examination and is already working as a data engineer in the industry then this is the right course for you guys in our last module we have covered the steps of becoming an data engineer next up is azure data Factory where you will explore Microsoft's cloud-based data integration services and learn how to create and manage data pipelines so why as your factory because again we know that we have been generating data at an exponential rate especially since last 5 years and in 2015 we were generating data at the rate of almost 3.4 exabytes per month and now we are generating at the rate of almost more than 44 exabytes of data per month but that's almost 15 times increase in the amount of of data in just a span of 5 years and it is going to be increased by almost 30% more because of the current lockdown phas by all the entire globe and that's why the entire consumption of data is again has been increased prominantly right that is that exactly is what we have the that's why we need to have the most optimized solution for data driven Solutions out there especially on the cloud computing platforms right and mod and modern data handling requires us to move from on premise to database to Cloud database services and that to quickly and then here we have to make sure again and that's why the data needs processing and goes through a series of steps making the that process TDS because again when we are trying to load the data if we are trying to migrate other data from our database servers from our on premise store Services then that has to go through a series of steps and that makes the entire process much more complicated it makes it much more slower as well and plus since it requires a good amount of investment both in times of both in terms of time and money it becomes we can see not feasible at all and data Factory simply help us in automating this entire process and does serve the cost that exactly is why we have data Factory now what exactly is data Factory don't what exactly is data Factory here so data Factory is basically a CL cloudbased integration service through which we can Define the entire we can design the data the entire workflow of data in cloud and making sure that we can use for orchestration and for the automation purposes so for example suppose if we have now if you want we can create a complete pipeline we can Define what is a source and how the pipeline should be structured how the data is coming in where to where it should be stored and that to on a regular manner so we can Define the Pipeline and we can automate the entire data movement that means if we make any changes to any particular file or any objects in the op promise that will be automatically replicated or I can say moved into the cloud services as well by the help of data Factory right and using data Factory here we can create and schedule the entire data workflows called as pipelines that can inest data from disparate s data source for example if you have multiple data source defined here we can simply inchest data from multiple sources and then we can process by simple a single pipeline it is basically used for processing large amount of large volume of data and these are done by using by integrating it with services like we have SEO SG inside Hadoop spark SEO data L analytics and ISO machine learning so basically if we are looking to make sure that we are we have the optimized data available and optimized stream of data available for analytics and for machine learning then we have to do that by using data Factory to maintain that consistency and if we talk about the entire workflow here first of all the pipelines are all datadriven workflows in N data Factory typically performs the following four steps it simply help us in connecting and collecting the entire data set so first of all if we have multiple sources if we have wide variety of data sources here we have to make sure we are able to set the source for each and every each and every connector we have to make sure that we are able to connect you to multiple services and then we can and through multiple apis or if we have multiple sources of data says for example if we have data from our own CRM coming in if we have data from our own Erp tool from multiple social media analytic PL dashboard here we have to connect and collect all the different data sets then we have to transform and enage because do data always consists of multiple anomalities right they may be multiple missing values they will be some encourag data formats available they the the data may be incomplete or it may be maybe multiple mistakes in that data correct so you have to make sure that we take care of entire data transformation that means if you want to convert the data format from one to the other like we have the ETL so again we can take care of the transformation and the enrichment of data that means data pre-processing and making sure it is much clear for and it is usable by the antical platforms for which we want to use it then we have to publish it to to make it usable and then we have toly monitor it so that the entire process can be monitored in case there have been some breakdowns in any of the Clusters from which we are collecting data then that can be monitored and if something can be done it can be it can be processed as well we can do that now let's understand each and every concepts for data Factory here so if you talk about the entire Concepts first of all we do have pipeline we do have something as pipeline so pipeline is a logical grouping of activities for example Suppose there are 10 different sequence like we discussed first of all we have to connect then we can transform then we have to process it then we have to make it available for the analytical platforms so there if there are multiple sequences that needs to be followed again that then that is something that we Define as a part of pipeline alog together and then we have data sets so obviously without data sets the entire data pipeline is of is of no use so data set simply represent the data structure within the available data stor so there have multiple data stores available what exactly those are we are going to discuss step by step and then we have activities for activities simply represents a processing step in a pipeline for example there are 10 different steps here for example we have to first of all connect to the source then we have to work on transforming data Cel then we have to work on pre-processing then we have to work on setting of The Connection so each now this entire process itself is called as pipeline maybe it consists of multiple sequence steps and each and every individual step here each and every individual step these are termed as activities so a pipeline is what we can say pipeline is simply a collection of different activities so we can so that activity is focused on completing one different task and pipeline simply defines the structure or we can say sequence of those STS and it simply makes sure that that sequence is followed whenever the entire pipeline is being implemented so let's erase this up and next we have Link services so it simply the information needed to connect to the external sources like for example we have apis if you have third party vendors and we have third party sources then again we do need an active apis for that and that's why these are all termed as link sources from which we can Source the entire dat ass sets here these are all additional links available and as we know data Factory is basically used for on premise itself if we are looking to to connect this if we are looking to connect this to multiple on premise data sets here then that exactly is what we use it for let's understand what exactly is data L service so data L as we know data l so data L as you know is simply an Enterprise wide hyperscale repository so here okay it is simply an Enterprise wi hypers scale repository for big data analytics workload and now it simply holds now it has a capability of paby so it can hold data for any size it will allow us to do again using that we can we can do multiple operational and exploratory analytics as well for example if we have multiple sources of data so for example if we have data sources like from on premise we have sensors data like for example if we when we talking about aviation industry then we have multiple we have tons of data sets coming in from different sensors especially for and same way for M for any manufacturing sectors as well same way if you have any data collected for any websites for example we are talking about any so any uh any stream of data coming in for analytics for any special websites for example suppose we have any big e-commerce solution for Amazon flip cart Airbnb so they have tons of data available on the platforms if we had data coming in from different devices from be it can also be a part of the I complete iot networks they can be data in the format of videos multiple social social media streams coming in for example we have streams of Facebook on Twitter RIT on redit on multiple platforms so if you have multiple social media streams coming in in the M of post or if you or let's say if you want to understand the real time I can say if you use case is to work on studying the user sentiments right then that in that case we have to make sure we are we are pitching in the multiple social media streams coming in if we have data in the format of images on application then these all Concepts have to be these all sources needs to be connected as a part of data Factory right and that's why here we get in now once we have these data sources available then we can connect it to ad analytics we can connect this to HG insight to R spark and machine learning purposes now if you have been aware if we have been aware of the fact such as now we can say data warehousing we can say data Lake works like a data warehous for example if you're working on the realtime analytics right and if you have data sources from multiple platforms so instead of connecting each and every Source manually or we can say one by one to spark we can store the entire data and a Sy centralized location and then we can we only need to connect a centrer location with a single connection to spark that's it so it simply improvises the entire entire performance as well just like we have red shift available in awss same way here we have data leag now let's also understand multiple data Le Concepts let's understand multiple data Le Concepts here so data le as we know again has multiple components inside it like we have analytics now in Daily components we have anal we have components for analytics such as a Insight we have AO data Lake available we can use data store for the complete storage purposes or for analytics we can indicate this with a inside and then we have hard inside and then we have as your data available and then when we are trying again in terms of analytics we have to make sure we are we do remember these three main key points analytics we can do on data of any size there is no Limited ation users all users are productive on day one and then we have to make sure that ready it is all scal up exactly as per our Enterprise requirement we have to make sure of that part in terms of type of data stored here it supports all data types we have structured semi structured and unstructured data types supported we have CSC files we have XML files right we have the emails we have Jason's files right so again these all are a part of structure data when we don't have a direct format available on which we can start and perform the entire sorting and that is example for sem structured data set and then we have unstructured data set like we have images videos audio clips these all are part of semi structured data sets and then if you compare data League to Data Warehouse if you compare data League to Data Warehouse here again data le as you know is simply complimentary to the data warehouse where as data whereas if you talk about the data housing service it may be soured to data leag data leak is basically used for detail data whereas we can see data warehouse is basically used for filter summarize and for refinement of data all together data L offers schema on read whereas data house offers schema on write and data L has one language to process data of any format we as in data bhouse we can process it using the SQL complying it all together now let's move into the handson here and see how exactly it this is implemented on top of Edo data Factory portal here data warehouse and data L let's understand this by simple use case here and for doing that let's open up our notepad here for example let's say we have multiple data sources for example we have data sources available from our CRF correct for example our main use cases we want to perform the analysis for sales report for sales for any company correct we are here to for analysis or multiple St data here right or suppose if we have data coming from from CR and we have data from Erp tools as well right and then we have data from the own data set available for example we have the own CC file also stored here so if you have multiple sources of data coming in then in in if you want to Club it if you want to use this on top of any antical platform so there are two ways of doing that either we have to connect each and every source with the antical platform and then only we can start working on it correct using the concept of data warehouse what we can do we can connect each and every Source we can store the in the data from each of source to a centralized location as in called as a data warehouse right and once the entire data is available in data house then we can connect that data house directly to us through a single collection to our antical platform so that whatever anst we are try to perform the entire process can be streamlined as a part of data warehousing service let's erase this up we getting started on data Factory here we can open we can come back to a portal so here we can come back to a portal from in case you have not signed up on asure we can open this entire portal here where we can start working on top of data Factory one one by one now for getting started here what we can do is we can simply come back we can simply come back and first of all for having the for having the entire database that we can install and can interpret locally we can use this we can use entire platform called as ssms in case you don't have the entire setup done at your end then we can simply go ahead and download the SQL Server management studio in case you don't have the setup done now once we have configured these now what we can do we can come back to our Su portal now in the S portal what we can do we can simply first of all once we are into dashboard here we have to work around with first of all setting up the entire data warehousing service so what we can do is here we can set up the entire dat warehousing service and for setting it up we can come back to our entire dashboard now for creating a new resource we here we have to click on this option which says create a resource and then we can choose resource type here so from databases here we can choose now here these are the most commonly used resources here for example if you want to go for web application for functional applications for SQL databases we can choose accordingly and if you want to start with SQL database we can open up SQL database and then we can set up the entire database platform but again in here our main goal is not to create a secq database but we set up the entire data with housing first and then we can saish in data services and then get started on top of it right so here what we can do here either we can go ahead and search for the services using these categories or we can use it directly we can use the service bar to search for the services directly for example here we can use data l as you can see currently we are planning to use data l so here we can choose data L here now if you want to create it now here we can simply click on create here we can Define the name of the current data storage service that we are going to create we can Define the entire name for example let's say we we call it as AA app itself here we can choose subscriptions So currently here either we can go for two type of sub description here we can choose pay as you go or here we can choose the the fre file if in case we have fre file available then we can choose that then here we can choose Resource Group in case we don't have the resource Group created we can choose a new Resource Group here let's enable the entire Services there so that we can so we can enable the entire subscription model here so let's enable the in subscription model here we can verify the code in case we you haven't set it up we can simply set it up you by simply specifying our entire details set it up so that we can use it for creation of our entire data data league and data warehousing platform now even though once you if we are signing this up for the first time here once we and once we sign it up we will be charge a nominal a ninal fee for making sure the entire account is well authenticated that we can use as a part of signing up for as your platform and once it is done this will take us back to the D dashboard now once we have entered the account here we can simply click on create resource and here we can choose data L for example let's say we are starting to we are plan to start with by setting up entire data warehouse first right so ear it was called as data warehouse which was now changed to data Le itself now we can open up Dela data Le account all all together so here we can go for data Le generation here we can click on create here we can Define the entire resource name let's say we name it ASA here we can choose the subscription model that we want plan to go for here we can choose Resource Group Resource Group are simply like when we are planning for creation of now when we are creating 10 different resources in aure and at the end we want to Simply segregate them based on a certain project for example we want to we want to know which all services have been deployed for which particular project and if you want to see a consolidated building for that particular project then we can go for Resource Group for example let say here we are going to create a resource Group for app one we can choose any available Resource Group name that you want to go ahead with all right and here we can go for p ASO model we can choose the encryption if you want to enable encryption we can choose the keys or we can simply choose the default keys from ad as suppose let's say we choose a key from the key generation itself we can configure the up it may take a couple of second for this entire data to be created and as you can see it says deployment in in progressor so it may take a couple of of minutes for this entire data to be processed here step by step and once it is processed we will be able to see this entire end live action and in the meantime in case you want to go ahead and set up the entire ssms that we have discussed we have to have another SMS for that we can connect our local SQL databases and then we can import directly into the a portal and then we also have to set up entire data Factory for setting up the entire data Factory here we can open up the data Factory servers available Ino platform here we can click on ADD and here we have to define the data Factory name for example let say we want to call us as ATA TF as in data Factory here we can choose a current version so so basically earlier we have been using V1 now but now since last year we have been using V2 so here we can choose the subscription model that we have subscribed for our for we can choose Resource Group that we have already created by the name of Eda app so that if we have 10 different Services if we have 40 different resources being deployed then we can and again all of those resources are mapped for a single application then we can map it through a single application Al together we can see that and once it is done we can simply choose a in which we are trying to launch this particular data Factory for again we have to choose a region which is obviously closer to our end users because at the end if we are not using a location which is not closer to end users then that will create a huge amount of lency so for example if our users base are based our users are based in Singapore so here we can choose region for Singapore Al together if we know that you our uses are based not from Singapore suppose from from Central India Australia East Africa north from Europe from USC so here we can choose our regions accordingly for example's suppose here we want to start with North Europe so here we can choose North Europe and then if you want to specify a get URL that will be used as a as a main data source here for for this data Factory for example let's say we use our own GitHub URL for this one let's log into our GitHub and for example suppose here we may have any particular depository here we have any that we want to connect here we can simply connect that let's say here we can enter our GitHub URL here if we have multiple repositories here for example here we have repository as Eda 1 we can choose Ina 1 here if we have Branch as we want to go for the for let's suppose for the development Branch so here we have a branch name for development we can choose a branch name so here we have the entire get URL here we have to define the repository name and then we have to define the branch name and then we can choose a root folder for which we are trying to connect to the connect to here so for example if we have Ro folder by the name of index we can Define index and now we can click on create So currently this is getting initialized as you can see our deployment is done for data L and currently the deployment is being in progress for data Factory that we currently deployed so far now currently this has been deployed here now if you want to download the entire deployment details we it's a good practice to download the entire employment details which will contain the entire list of all the piece of of information and then for getting the entire operations detail here we can simply choose the operations detail where we can Define the entire operations name duration and now if you want to configure this we can open up the resource where and this resource here we first of all here if you want to quick if you want to see the entire activity that means what exactly it has been and again what exactly has been the IPS the storage usages and the processor us we can go ahead and see the entire monitoring part and if we are Conn if we are planning to connect to to this particular inance we have to make sure we are defining the current IM user rules and their policies because again if we are using this through some some particular account here then we can end as we Define the rules for them they will not be able to make any changes to this particular server out there that is something that we have to take care of and along with that let's go ahead and create one storage account as well so again for doing that we can use a service bar available on top so here we have to open up our storage account let's open this up now currently as you can see currently by default when start up they won't be any storage account created and deployed so here we have to go ahead and create one click on create storage account here we can choose a subscri deson model here we can choose a resource Group for the same application resource Group that we have currently created here we have to define the entire storage account name let's say we name it ASA itself and then we can choose the location now remember this okay we already have configured that so here one two so here we have to choose the same location in which we had deployed the other data L and our data house our data Factory itself so we had deployed this for northern Europe so here we can choose North Europe and based on where our locations are where users are basically located then we can choose a performance to be standard and premium so again in terms of account kind here we can choose storage or we can choose if you're going for general purpose version one or version two so general purpose one again they depending upon requirement we can choose accordingly if you want a higher performance that means if we are looking to to have a higher workload then we can use Gen 2 it was released earli last year and then if you want to replicate the entire data we want to go for Zone redundant we want to go for locally redundant as if we want to maintain local copy or we want to maintain a read access stor that means again based on multiple regions it will be copied but again the copy will have only the read only access that means this will be used as a primary account for storing data for writing data and the replications will be used for just for the read purposes so that the entire situation of bottleneck is also not created and then we can choose assets here to be hot or to be cool depending upon the requirement we can choose it to be hot or cool here just like we in case we have been Avail if we have been familiar with the concept of multiple storage classes in AWS same way we we have standard and then we have infrequent access so here we can choose school if you want to move to some to something like infrequent access which is not used frequently or we can keep it to hard for those standard we can most frequently used files here so currently we can keep it to host then we have to define the networking point if we are trying to deploy this in our own isolated Network then we can choose public then we can choose a private or public endpoint depending upon the requirement here so if we have if we have a virtual Network created then we can go ahead and use selected networks if you want to deploy this on public endpoints we can go for public end points or if okay again in case you want to deploy this on our private then first of all we have to configure our prior endpoint and then only we can get started then we can Define the production type here whether we want to go for stop or desktop delete or if sof as in if you want to retrive it we can simply do that and then we have the Gen 2 hierarchal should be disabled as you don't need it as a part of our Handel currently so we can keep it to disable then we can Define the tags here tags are simply used for sorting and filtration purposes if you want to sort this out later on if there are multiples service that be deployed and now we want an easier way to host it easily when we can easily do that and in here we can define t if you don't want to use it we can review it we have to wait for this one to be reviewed here once we are done we can click on create let's and let's wait for this one to be created if you are looking to move files here from one from one part to the other again we can easily do that is it access to free to create database for creating databases here here we here we can choose a engine for database for creation of multiple databases here we have different services for for doing that for example suppose for creation of databases here we can click on create resource we can choose resource type as databases and here we can choose which part database engine we are going to deploy here just like we have RDS available in adus same way here we can change it up and if we looking to create a complete data pipeline then that's why we use data Factory so for example if we are trying to to transfer one data from the other account here we can simply share the across multiple accounts if you're trying to start to transfer data from on premise to SEO or from SE to on premise again for if you looking to transfer from a to on promise then we can from on a to on promise that means locally then we can use another service called as storage Explorer as a part of AO platform we can use that if we looking to transfer from AO platform to on premise we can do that we can take the help of storage Explorer our service and then we can easily use the AO import export to Simply transfer the service from our AO platform to optimise just like we have a physical device offered by as a part of snowball in AWS in case we have been working with AWS and we there we have a service called as snowball where it is a simple physical device for if we looking to get connected we can use that we can take the help of storage gate phase in that so just like storage gate phas here we can take the help of storage Explorer back into Aur for transfering data from in and out of AO platforms moving on to our next question which of the following is a key component of aure data Factory for creating data driven workflows and the options are a pipelines B buckets C containers or D modules if you know the correct answer please comment down below we have just learned about Azure data Factory moving forward our next topic is azure database Services where you will dive into the various database services offered by Azure and how they support data storage and management now let's have an introduction of a database first let's understand why do we need a database so the various reasons a database is important are first of all it manages large amounts of data a database stores and manages a large amount of data on a daily basis this would not only be possible using any other tool such as a spreadsheet as they would simply not work second is its accuracy so a database is pretty accurate as it has all sorts of building constraints checks Etc this means that the information available in database is guaranteed to be correct in most cases it's easy to update data in a database so in a database it is easy to update data using like various data manipulation languages available one of these languages SQL fourth is security of data so databases have various methods to ensure security of data there are user logins required before accessing a database and various access specifiers these allow only authorized users to access the database fifth is data Integrity this is ensured in databases by using various constraints for data data Integrity in databases makes sure that the data is accurate and consistent in a database the last is easy to research data it is very easy to access and research data in a database this is done using data query language which allow searching of any data in the database and performing computations on it now that you have understood the need of a database let's briefly understand what actually it is so a database is an organized collection of structure information or data typically stored electronically in a computer system a database is usually controlled by database management system together the data and the database management system along with applications that are associated with them are referred to as a database system often shortened to just a database so data within the most common types of database in operation today is typically modeled in rows and columns in a series of tables to make processing and data quaring efficient the data can then be easily accessed managed modified updated controlled and organized most databases are structured query language for writing and querying data databases are used to support internal operations of organizations and to underpin online interactions with customers and suppliers databases are used to hold administrative information and more specialized data such as engineering data or economic models example includes computerized Library System flight reservation system computerized past inventory system and many content Management systems that store websites as collection of web pages in a database now that you have an understanding of Microsoft Azure as well as of a database let's now take a look at different types of databases in Azure first is relational database a relational database is a type of database that stores and provide access to data points that are related to one another relational databases are based on the relational model and intuitive straightforward way of representing data in tables in a relational database each row in the table is a record with a unique ID called the key The Columns of the table hold attributes of the data and each record usually has a value for each attribute making it easy to establish the relationships among data points in a relational database all data is stored and accessed by relations so relations that store data are called base relations and in implementations are called tables other relations do not store data but are computed by applying relational operations to those relations these relations are sometimes called derived relations in implementations these are called views or queries derived relations are convenient in that they act as a single relation even though they may grab information from several relations each relation or table has a primary key this being a consequence of a relation being a set a primary key uniquely specifies a tuple within a table while natural attributes are sometimes good primaries so this is all about relational database then second we have is non- relational database also known as no SQL databases so no SQL database or non- relational database provides a mechanism for storage and retrival of data that is modeled in means other than the tabular relations used in relational databases so non- relational databases are increasingly used in big data and realtime web applications and non- relational databases are also like sometimes called not only SQL to emphasize that they may support SQL like query languages or sit alongside SQL databases so the prominent non relational databases provided by aour is Cosmos database that I will EXP explain you further in this video and third is inmemory database so an inmemory database also like in short form we say it as IMDb also like a main memory database system or mmdb or memory resident database these are all the names of it so an in memory database is a database management system that primarily relies on Main memory for computer data storage it is contrasted with database management system that employ a disk storage mechanism in memory databases are like faster than dis optimized databases because this access is slower than memory access the internal optimization algorithms are simpler and execute fewer CPU instructions accessing data in memory eliminates seek time when querying the data which provides faster and more predictable performance than disk a potential technical hurdle with inmemory data storage is the volatility of ram specifically in the event of a power loss intentional or otherwise data stored in volatile Ram is lost with the introduction of nonvolatile Random Access Memory technology in memory database will be able to run at full speed and maintain data in the event of power failure let's Now understand the architecture of database Services provided by so you can see the it looks like a complex architecture but I will explain you quite easily so the basic fundamental building block that is available in azour is the SQL database so Microsoft offers this SQL server and SQL database on azour in many ways we can deploy a single database or we can deploy multiple databases as part of a shared elastic pool you can see the elastic pools and single database okay Microsoft introduced a managed instance that is targeted towards on premises customers so if we have some SQL databases within our on premises Data Center and we want to migrate the database into Azure without any complex configuration or ambiguity then we can use a managed instance because this is mainly targeted towards on premises customers who want to lift and share their on premises database into Azure with the least effort and optimized cost we can also take advantage of Licensing we have within our on premises data center Microsoft will be responsible for maintenance patching and related services but in case if we want to go for the infrastructure as a service for the SQL Server then we can deploy SQL server on the Azure virtual machine if the data has a dependency on the underlying platform and we want to log into the SQL server in that case we can use the SQL server on a virtual machine we can deploy a SQL Data Warehouse on the cloud aor offers many other database services for different types of databases such as my Marb and also post SQL once we deploy our database into hro we need to migrate the data into it or replicate the data into it okay then we have is azure database services for data migration so services that are available in Azure which we can use to migrate the data from our on premises SQL Server into Azure so in that the first one is azure data migration service so it is used to migrate the data from our existing SQL server and database within the on premises data center into the Azure then we have azure SQL data synchronization if we want to replicate the data from our on premises database into azour then we can use azour SQL data sync then we have SQL stretch database so it is used to migrate cold data into Azure SQL stretch database is a bit different from other database offerings it works as a hybrid database because it divides the data into different types like hot and cold so hot data will be kept in the ones Data Center and cod data in the Azure then we have is data Factory so Azure data Factory is used for ETL means transformation extraction and loading so using the data Factory we can even extract the data from our on premises data center we can do some conversion and load into the Azure SQL database data Factory is an ETL tool that is offered on the cloud which we can use to connect to different databases and like extract the data or transform it and load into a destination then there's azour security Center so all the databases that exist in our need to be secured and also we need to accept connections from known Origins for this purpose all these database Services comes with firewall rules where we can configure from which particular IP address we want to allow connection we can define those firewall rules to limit the number of connections and also reduce the service attack area so now let's talk about the cosmos DB so Cosmos DB is nothing but a SQL data store that is available in aour and it is designed to be globally scalable and also very highly available with extremely low latency Microsoft guarantees latency in terms of reading and writing wres with Cosmos DB for example if we have any application such as iot or gaming where we get a lot of data from different users spread across globally then we will go for cosos DB because Cosmos DB is designed to be globally scalable and highly availablee to which users will like experience low latency finally there are two things and one is we need to secure all the services for that purpose we can integrate all these services with azour active directory and manage the users from Azure active directory also to monitor all these Services we can use the security Center so there is an individual monitoring tool too but AZ security Center will keep on monitoring all these services and provide recommendations if something is wrong I hope the architecture of Azo database Services is now clear to you now let's uh move forward to briefly understand the database Services provided by Azure so the first is azure SQL database so SQL database is the flagship product for Microsoft in the database area it is a general purpose relational database that supports structures like relation data Json spatial and SML the Azure platform fully manages every Azure SQL database and guarantees no data loss and a high percentage of data availability Azure automatically handles patching backups replication failure detection underlying potential Hardware software or network failure deploying bug fixes failovers and like database upgrades and other maintenance tasks so there are three ways we can Implement our SQL database so first is managed instance this is premar targeted towards on premises customers in case if we really have a SQL Server instance in our on premises Data Center and you want to migrate that into Azure with minimum changes to a application and the maximum compatibility then we will go for manage instance second is single database so we can deploy a single database on AJ its own set of resources managed via L logical server okay then we have this elastic pool we can deploy a pool of databases with a shared set of resources managed via logical server we can like deploy the SQL database as an infrastructure as a service that means we want to use the SQL server on azure virtual machine but in the case we are responsible for managing the SQL server on that but in that case we are responsible for managing the SQL server on that particular AO virtual machine so then we have is the purchasing model so there are two ways we can purchase the SQL server on AO so first is Vore purchasing model also as virtual core purchasing model so the Vore purchasing model enables us to independently scale compute and storage resources match on premises performance and optimize price it also allows us to choose a generation Hardware it also allows us to use Azure hybrid benefit for SQL Server to gain cost savings best for the customers who value flexibility control and transparency so second is DD model it is based on a bundle measure or compute storage and inut output resources so sizes of the compute are expressed in terms of database transaction units means ddus for single databases and elastic database transaction units for elastic pools this model is best for customers who want on simple pre-configured resource options in the second database service is azure Cosmos database so Azure Cosmos database is a no SQL data store it is different from the traditional relational database where we have a table and the table will have a fixed number of columns and each row in the table should adhere to the scheme of the table in the no SQL database you don't Define any schema at all for the table and each item or row within the table can have different values or different schema itself so so now let's understand the cosmos database structure first one in the structure is database so we can create one or more Azure Cosmos database under our account a database is analogous to a name space and it is the unit of management for a set of azure Cosmos containers so the second is Cosmos account so the Azure Cosmos account is the basic unit of global distribution and high availability for globally Distributing our data and throughput across multiple Azure regions we can add or remove Azure regions from our azure Cosmos at any time I mean Azure Cosmos account at any time so the third is a container so an Azure Cosmos container is the unit of scalability for both provision throughput and storage of items a container is horizontally partitioned and then replicated across multiple regions then let's understand the types of consistency under Cosmos DB so Azure Cosmos database approaches the data consistency as a spectrum of choices instead of two extremes so strong compatibility and eventual consistency are at the ends but but these are many consistency choices along along the Spectrum so the consistency levels are region agnostic the consistency level of our Azure Cosmos account is guaranteed for all read operations regardless of the region from which the reads and rs are served the number of areas associated with the Azure Cosmos account or whether our account is configured with a single or multiple right regions then there is request units so we pay for the throughput we provision and the storage we consume on an hourly basis with Azure Cosmos DB remember this DB means database so then there are request units in Cosmos DB means Cosmos database so we pay for the throughput we provision and the storage we consume on an hourly basis with Azure Cosmos GB the cost of all the database operations is normalized by Azure cosos DB and is expressed in terms of request units the price to read a 1 KB item is a one request unit all other database operations are similarly assigned with a cost in terms of research units the number of research consumed will depend on the type of operations item size data consistency query patterns etc for the management and planning of capacity aure CMOS database ens shows that the number of research units for a given database operations over a given data set is deterministic and the third database Services Azure data Factory so Azure data Factory is a data integration service based on the cloud that allows us to create data driven workflows in the cloud for orchestrating and automating data movement and data transformation data Factory is a perfect ETL tool on cloud data Factory is designed to deliver extraction transformation and loading process within the cloud the ETL process generally involves four steps so the first one is collecting collect we can use the copy activity in a data pipeline to move data from both on premises and Cloud secure data stores so the second is a transform so once the data is present in a centralized data store in the cloud process or transform the collected data by using computer services such as HD inside Hadoop spark data leak analytics and machine learning third is published so after the raw data is refined into a business ready consumable form it loads the data into an azour data warehouse azour SQL database and Azure Cosmos database Etc so fourth is Monitor so Azure data Factory has built in support for pipeline monitoring via azour monitor API Powers shell log analytics and health panels on the azour portal so then there are components of data Factory so data Factory is composed of six key elements all these components work together to provide the data form on which you can form a data workflow with the structure to move and transform the data so first one is pipeline a data Factory can have one or more pipelines it is a logical grouping of activities that perform a unit of work the activities in a pipeline perform the task Al together for example a pipeline can contain a group of activities that inest data from a Azure blob and then runs a hi query and an HD Insight cluster to partition the data so second is activity it represents a processing step in a pipeline for example we might use use a copy activity to copy data from one data store to another data store then we have a data sets so it represents data structure within the data stores which point to or reference the data or we want to use in our activities as input or output then there are Link services so it is like connection strings which Define the connection information needed for data Factory to connect to external resources a link service can be a data store and computer sources linked service can be a link to data store or a compute resource also then we have a triggers so it represents the unit of processing that determines when a pipeline execution needs to be disabled we can also schedule these activities to be performed at some point in time and we can use the trigger to disable an activity then the last one is control flow so it is an orchestration of pipeline activities that include chaining activities in a sequence branching defining parameters for the pipeline label and passing arguments while invoking the pipeline on demand or from a tiger we can use control flow to sequence certain activities and also Define what parameters need to be passed for each of these activities I hope you have now understood the major Services of azure databases so now let's have a look at some of the use cases for Azure database Services first let's see the use cases for SQL database so the first one is developer or test environment an important use case for replicating or migrating data to SQL hosted on Azure is for developer or test environments before deploying to the production environment it is perent that the data is tested against developer and test environments so Azure SQL database can act as a target for such environments the life production environment can be replicated to the developer or test environment using a database copy so the second is business continuity one of the most important use cases for SQL on azour is using it as a Dr Target to maintain business continuity Azure SQL databases can provide an SLA of up to 99.99% by maintaining several copies of the data this provides business continuity as allows you to restore GE redundant copies of the data or use active Geo redundant copies as failover points in use of outages at data centers or in regions besides SQL databases you can also use availability groups to fulfill business continuity demands not only can you use availability groups in Azure SQL virtual machines but also use Azure SQL virtual machine instances as a target for high availability and disaster recovery and the third one is scaling out readon workloads apart from Prov providing PC or Dr capabilities active GE replication can also be used to offload readon workload such as reporting jobs to secondary copies you can also extend on premises SQL Server instance using readable always on replicas and the fourth one is backup and G so Azure SQL database are backed up automatically on a regular basis and there are no storage costs for to 200% of the maximum provision database storage you can restore backups to any point in time going back to a pretended period which which is determined by the Azure SQL Service Tire in use on premises SQL Server databases and transaction locks can also be bagged up directly to Azure using the backup to URL feature and stored in Azure storage so Azure SQL databases can also be stored on local storage by exporting them to backpack files means backup and the St files so fifth one is Advanced analytics so another important reason for hosting SQL in azour is to make use of azure's advanced gentics platforms such as azour storage blob and azour data leak store a common scenario with Advanced analytics is when users reference data from various data sources use Azure data leak store as the staging area or perform transformation activities using higho spark and finally load the data into Azure data vhouse for bi and Reporting bi means business intelligence now let's see the use cases for Cosmos database first of all they are used in iot and telematics so iot use cases commonly share some patterns in how they ingest process and store data first these systems need to ingest burst of data from device sensors of various locals next these systems process and analyze streaming data to derive real time sites the data is then archived tool Co storage for batch analytics Microsoft Azure offers Rich services that can be applied for iot use cases including Azure Cosmos database Azure event hubs Azure stream analytics Azure notification Hub Azure machine learning Azure HD insight and powerbi burst of data can be ingested by Azure event hubs as it offers High throughput data ingestion with low latency data ingested that needs to be processed for realtime Insight can be funneled to aure stream analytics for realtime analytics data can be loaded into Azure Cosmos database for an ad hoc query once the data is loaded into azour Cosmos database the data is ready to be queried in addition new data and changes to existing data can be read on changed feed so Chang feed is a persistent append Only log that stores changes to Cosmos containers in sequential order then all data or just changes to data in Azure Cosmos database can be used as reference data as part of a realtime analytics in addition data can further be refined and processed by connecting Azure Cosmos database data to HD insight for pig hiive or map reduce jobs defined data is then for a sample of iot solution using Azure Cosmos database event hubs and storm see the HD Insight storm examples repository on GitHub okay then we have is retail and marketing so Azure Cosmo database is used extensively in Microsoft's own e-commerce platform that runs the Windows store and Xbox Live it is also used in the retail industry for storing catalog data and for event sourcing in order to process pipelines so catalog data storage scenarios involve storage and querying a set of attributes for entities such as people places and products some examples of catalog data are user accounts product catalog iot devices Registries and build of material systems attributes for this this data may vary and can change over time to fit application requirements consider an example of a product catalog of an automative part supplier every part may have its own attributes in addition to the common attributes that all parts share furthermore attributes for a specific part can change the following year when a new model is released Azure Cosmos database supports flexible schemas and hierarchical data and thus it is well suited for storing product catalog data Azure Cosmos database is often used for Event Source to power event driven architectures using its change feed functionality the change feed provides Downstream microservices the ability to reliability and incrementally read inserts and updates made to an Azure Cosmos database this functionality can be leveraged to provide persistent event store as a message broker for State changing events and drive order processing workflow between any microservices in addition data store in Azure Cosmos database can be integrated with HD insight for big data analytics via Apache spark jobs so the third one is gaming the database tire is a crucial component of gaming applications modern gaming app perform graphical processing on mobile or console clients but rely on the cloud to deliver customized and personalized content like in-game stats social media integration and high school leaderboards games often require single millisecond latencies for leads and rights to provide an engaging in-game experience a game database needs to be fast and be able to handle massive Spice in request rates during new game launches and feature updates so Azure Cosmo database is used by games like The Walking Dead No Man's Land by next games and hello five guardians so Azure Cosmos database provides the number of benefits to game developers like Azure Cosmos DB allows performance to be scaled up or down elastically this allows games to handle updating profiles and stats from dozens or millions of simultan Gamers by making a single API call then azour Cosmos DV supports millisecond reads and rides to to help avoid any lags during the game play the Azure costos database automatic indexing allows for filtering against multiple different properties in real time for example locating players by the internal player IDs or their game center Facebook Google IDs or quering based on player membership in a guild this is possible without building complex indexing or shedding infrastructure social features including ingame that messages player Guild membership challenges completed high score leaderboards and social graphs are easier to implement with a flexible schema so Azor Cosmos database as a managed platform as a service require minimal setup and management work to allow for Rapid iteration and reduce time to market the last one is web and mobile applications so Azure cosos database is commonly used within mobile and web applications and is well suited for modeling social interactions iterating with third party services and for building Rich personaliz experiences the customer database sdks can be used to build Rich IOS and Android applications using the popular zarine framework so under applications first we have the social applications and then we have personalizations so in Social applications a common use for Azure Cosmo database is store and query user generated content means ugc so for web mobile and social media applications some examples of a user generated content are chat sessions tweets blogs posts rating and comments often the ugc in social media applications is a bland of free form text properties text and relationships that are not bounded by rigid structure content such as chats comments and posts can be stored in Cosmos DB without requiring Transformations or complex object to relational mapping layers data properties can be added or modified easily to match requirements as developers it iterate over the applications code thus promoting rapid development applications that integrate with third party social network must respond to changing schemas from these networks as data is automatically indexed by default in Cosmos database data is ready to be queried at any time hence these applications have the exibility to retri projections as their respective needs so the second thing in web mobile applications is personalization so nowaday modern applications comes with complex views and experiences these are typically Dynamic catering to user preferences or moods and branding needs hence applications need to be able to try personalization settings effectively to render UI elements and experiences quickly Json a format supported by the cosmos DB is an effective format to represent UI layout data as it is not only lightweight but also can be easily interpreted by JavaScript Cosmos R offers turnable consistency labels that allows fast reads with low latency rights hence storing UI layout data including personalized settings adjacent documents and Cosmos GB is an effective means to get this data across the wire so these were the use cases for Cosmos GB now that you have a theoretical understanding of azure database Services let's now see a simple deployment of a database service on Microsoft azure the simply type Microsoft Azure on Google what you can do is you can create a free account on Microsoft Azure you get a 12 months of free services and around 40,000 rupees of free credits also for using the services we can just directly open the console from here I just sign in so for deploying a simple database service so what we are going to deploy today we can like deploy Cosmos database like I have explained you what is cosmos database a new SQL database it is so we can just go to console to the portal I and go to the portal you can create a resource from here I can search for Cosmos Jo Cosmos yes you can see here like free credits I have a free trial account so it is showing that I have 14,500 free credits so this is the like credit amount you get for in a free trial okay so you can create a Azure Cosmos DB from here which one you want to create like you can create 4 SQL one so Resource Group can give a new one or we have existing we have a Rec Cosmos one resource Cosmos so if you want to choose the existing one or if you want to choose the new one okay so you can choose the existing one from here or if you want to create a new one then create new One Source One Cosmos I hope that works out yeah then give any unique name for this like uh I will give a demo underscore Cosmos 1 2 3 okay it cannot contain uh UND remember these things okay it does cannot contain this so 1 2 3 4 I will okay it's not available so I will give five also yeah it's available now and like choose your location whichever location you are located in can use nearby location so mine is Asia Pacific Central India so I've chosen this so free trial account is already there I applied for it then you can just review and create before getting deployed it will show you the review for it so remember that on the basis of the location we have selected the creation time will differ okay so you can review it all the information what you have inserted so now you can create so deployment is in progress it will take a few minutes you can see like how the resource has been created deployment is in progress will soon be created you can check details for it from here let's go to for again like it is already pinned here or you can search from here okay for Cosmos GB so just click here so it is showing that it is getting created this is the one I have created before only this one is creating resister in progress let's refresh one again you can see the processing going on here yeah so your deployment is complete showing you can go to Resource from here also you can just refresh it from here so yeah this is how it's been created you can open the resource from here and you can go to activity log or data Explorer you can create a database anything or you can see the consistency of it like default consistency and everything that's how Cosmos database has been deployed so I hope you have understood this deployment moving on to our next question which feature of azure SQL database helps to ensure High availability and disaster recovery and the options are a autos scaling B GE application C query store or D read replicas if you know the correct answer please comment Down Below in the previous module we have explored azure's database Services now we will focus on Azure SQL database where Learners will learn about the managed relational database service its features and how to use it effectively let's look into the family of azure SQL so first one in the family is SQL server on Virtual machines so with this you can lift and shift your SQL Server workloads to the cloud to get the combined performance security and analytics of SQL server with flexibility and hybrid connectivity of azure with 100% code compatibility access the latest SQL Server updates and releases including SQL Server 2019 register your virtual machines with SQL infrastructure as a service agent extension for automated virtual machine management at no additional cost SQL server on Azure virtual machines is part of the Azure SQL family which allows you to migrate existing apps or build new apps on the best cloud destination for a mission critical SQL Server workloads so its features are first of all best TCO that is total cost of ownership with Azure hybrid benefit with Azure SQL Server you can save up to 84% compared to Amazon web services migrating SQL Server databases with Azure hybrid benefit and get free extended support for SQL Server 2018 R2 images in Azure infrastructure as a service activate Azure hybrid benefit when you provision SQL server on Azure virtual machines images from the Azure Marketplace second feature is high performance virtual machines for SQL server on Linux and windows so you can take advantage of SQL Server virtual machines with industry leading performance choose from images with Windows Server redhead Enterprise Linux suc Enterprise Linux server or you been to Linux gain collocated integrated support for your SQL workloads with the redhead and suc the third feature is built-in security and manageability so you can ease maintenance with automatic security updates and restore your database to a specific point in time with Azure backup help protect your data addressed and in motion with the database stated as least vulnerable over the last 9 years in the cloud with the most global national and Industry certifications second member in the family is azure SQL managed instance so part of the Azure SQL service portfol Azure SQL managed instance is the intelligent scalable Cloud database service that combines the broadest SQL Server engine compatibility with all the benefits of a fully managed and ever platform as a service with SQL managed instance confidently modernize your existing apps at scale by combining your experience with familiar tools skills and resources and do more with what you already have Azure Arc enabled SQL manage instance is now in preview you can run the server on premise on any infrastructure of your choice with Azure Cloud benefits like elastic scale unified management and a cloud billing module while staying always current some of its features are always operate on the latest version of SQL so SQL manage instance is built on the SQL Server engine it's overgreen meaning it's always up to date with the latest SQL features and functionality never worry about updates upgrades or end of support again second is fully managed and optimized for DBA product tivity so boost productivity and operate more efficiently by letting the service perform time consuming and complex tasks on your behalf features like built-in High availability disaster recovery and automated backups ensure your data is available when you need it while AI power automatic tuning optimizes performance for you SQL manage instance combines all the best of SQL server with the financial and operational benefits of the platform as a service and the third feature is maintain SQL Server application compatibility so application modernization with the latest SQL Server capabilities in the cloud SQL managed instance provides an entire SQL Server instance within a managed service so you can continue to use familiar tools and SQL Server features like cross database queries and Link servers SQL managed instance maintains the highest compatibility labels so you can move your on premises workloads without worrying about application compatibility of performance changes and the third member in the family is azure SQL Edge so Azure SQL Edge is an optimized relational database engine Geared for iot and iot Edge deployments it provides capabilities to create a high performance data storage and processing layer for iot applications and solutions Azure SQL Edge provides capabilities to stream process and analyze relational and non-relational data such as Json graph and time series data which makes it the right choice for a variety of modern iot applications Azure SQL Edge is built on the latest version of the SQL Server database engine which provides industry leading performance security and quering processing capabilities since azour SQL Edge is built on the same engine as SQL server and Azure SQL it provides the same transact SQL programming surface area that makes development of applications or Solutions easier and faster and makes application probability between iot Edge devices data centers and Cloud straightforward so there are two different deployment models in Azure SQL Edge so the first one is connected deployment through Azure iot Edge Azure SQL Edge is available on the Azure Marketplace and can be deployed as a module for Azure iot Edge second is disconnected deployment so Azure SQL Edge container images can be pulled from Docker Hub and deployed either as a standalone doer container or a kubernetes cluster so some of the features of azour SQL EDR built in data streaming and time series within database machine learning and graph features for low latency analytics then data processing at the Edge for online offline and hybrid environments to overcome latency and bandwidth constraints next is deploy and update from the azur portal or your enterprise portal for consistent security and trunk key management last one is simplified pricing with no upfront cost and subscription offers as low as us $60 per year per device so the fourth and the major family member of azour SQL family is azure SQL database which we're going to briefly understand further first let's understand why we need an Azure SQL database extensively I will tell you the top five benefits that companies are realizing with SQL database so the first one is scalability and Beyond flexible service plans for SQL database meet the need for both big and small business users SQL is no longer Way Out Of Reach for smaller operations because the pricing structure allows users to pay as little as $4.99 per database per month with a maximum storage set at 150 GB per database that's a lot of space for very small cost second is high speed and minimal downtime so high availability architecture mean high speed connectivity and data retrival as well as low downtime at your organization there's nothing worse than stopping business because your technology can't keep up and that is no longer a problem with SQL database secondly companies can add application instances as needed through sharing for examp example shedding is a type of database partitioning that separates very large databases into smaller faster more easily managed Parts called Data shards not only can you spin notes up and down on demand you can leverage a federation infrastructure to scale more easily without affecting other areas of the server SQL Azo Federation data migration visard can further automate this process which impacts your organization and the employees much less lastly there are multiple levels of implement mentation that you can benefit from if you just need a website and a database you can hitch a SQL Azure instance to an Azure website and you are done if you need a full-blown virtual machine now or even down the road you can get that as well you can even use a locally deployed instance of SQL server in the virtual machine instead of SQL Azure these implementation options help make your company more adaptable to the inevitable changes it under goes on a regular basis with SQL Azure you are not stuck you are a foundation that encourages growth while working with it third is improved usability so SQL developers are familiar with all things SQL and SQL database can be updated with SQL CMD or the SQL Server management Studio better yet there is no coding required using a standard SQL it's much easier to manage database systems without having to write or update a huge amount of code fourth time is on your site with no administrative duties on your physical location employees can take time for strategic work to advance grow all around business success when your database is hosted in the cloud you don't have to deal with setting up SQL Server appropriating databases and dealing with physical machine maintenance and upkeep all of this results in better alignment of your organization and ultimately more time on your site Fifth and the last one is easy to use migration tools ramp up time with SQL database is now easier than ever and free SQL data synchronization allows you to either synchronize your SQL Server store data or migrate that data without having to worry about the cost to migrate by syncing gigabyte size tables now that you know why we need azour SQL database let's briefly understand what actually it is the basic fundamental building block that is available in azour is the SQL database Azure SQL database is fully managed platform as a service database engine that handles most of the database management functions such as upgrading patching backups and monitoring without user involvement Azure SQL database is always running on the latest stable version of the SQL Server database engine and ped OS with 99.99% availability platform as service capabilities that are built into Azure SQL database enable you to focus on the domain specific database Administration and optimization activities that are critical for your business with Azure SQL database you can create a highly available and high performance data storage layer for the applications and Solutions in azour SQL database can be the right choice for a variety of modern Cloud applications because it enables you to process both relational data and non-relational structures such as graphs Jon spatial and XML Azure SQL database is based on the latest stable version of the Microsoft SQL Server database engine you can use Advanced query processing features such as high performance inmemory Technologies and intelligent query process processing in fact the newest capabilities of SQL Server are released first to SQL database and then to SQL Server itself you get the newest SQL Server capabilities with no overhead for patching or upgrading tested across millions of databases SQL database enables you to easily Define and scale performance within two different purchasing models and that we will discuss further so Microsoft handles all patching and updating of the SQL and operating system code you don't have to manage the underlying infrastructure so now let's understand the deployment models Azure SQL databas provides the following deployment options for a database so the first one is managed instance this is primarily targeted towards on premises customers in case if we already have a SQL Server instance to a on premises Data Center and you want to migrate that into Azure with minimum changes to our application and the maximum compatibility then new will go to the manage instance second is single database so single database presents a fully managed isolated database you might use this option if you have modern Cloud applications and microservices that need single reliable data source a single database is similar to a contained database in the SQL Server database engine last one is elastic pool so elastic pool is a collection of single databases with a shred set of resources such as CPU or memory single databases can be moved into and out of an elastic pool now let's understand the purchasing models so SQL database offers the following purchasing models you can see here so the first one is vcore based purchasing model which is new and it offers a totally different approach to sizing your database it is easier to translate local workloads to a vcore based model because the components are what we are used to the vcore based model lets you choose the number of VOR the amount of memory and the amount and speed of storage the Vore based purchasing model also allows you to use Azure hybrid benefit for SQL Server to gain cost savings then the next one is DTU based purchasing model so the DTU based purchasing model offers a bland of compute memory and input output resources in three service tires to support light to heavy database workloads computer sizes within each tire provide a different mix of these resources to which you can add additional storage resources as you can see from the following diagram the DTU model offers a pre-configured and predefined amount of compute resources Vore is all about independent scalability where you can look into a specific area such as the CPU core count and memory resources something that you cannot control at the same granular level when using the DTU based model so the third one is the serverless model which automatically scales compute based on workload demand and builds for the amount of compute used per second the serverless compute Tire also automatically pauses databases during inactive periods when only storage is built and automatically resumes databases when activity returns now let's see the service triers for Azure SQL database so the first one is general purpose or standard model it is based on a separation of computing and storage service this architecture model depends on the high availability and reliability of azure premium storage that transparently copies database files and guarantees for zero data loss if underlying infrastructure failure happens second is business critical or premium service Tri model it is based on a cluster of database engine processes both the SQL database engine process and underlying MDF or ldf files are placed on the same node with locally attached SSD storage providing low latency to a workload High availability is implemented using technology similar to SQL Server always on availability groups the third one is hyperscale Service Tire model it is the newest Service Tire in the VOD based purchasing model this tire is a highly scalable storage and compute Performance Tire that leverages the Azure architecture to scale out the storage and compute resources for an Azure SQL database beyond the limits available for the general purpose and business critical service tires now let's understand the selfcontain services in azour SQL database so azour SQL database is a database as a platform service designed for applications that will use database as self-contained service databases can be grouped together to simplify management options or share the resources there are different options that can be used to bound databases in the group so the first one is databases in logical server so logical server is a default container for Azure SQL database logical server enables you to perform administrative tasks across multiple databases including specifying Regions login information firewall rules auditing thread detection and failover groups all databases within the server are self-contained with independent service triers that can be specified for each database each database can be independently scaled up or down by changing performance tries on the database which will not affect other databases databases cannot share resources and each database has guaranteed and predictable performance defined by its own service stle some server level specific features such as cross database quaring linked servers SQL agent service broker or CLR are not supported in Azure SQL database placed in logical servers second one is databases in elastic pool so databases need to share resources can be stored in elastic pools instead of The Logical server all databases within the elastic pool share the same resources associated with the elastic pool label currently there are three service tires in the elastic pools basic standard and premium databases within the elastic pools cannot have different service ties because they share resources that are assigned to the entire pool resources usage in one database might affect others however you can specify Reserve performance for the database in the pool that will guarantee a minimal amount of resources that the database can have this model is a good choice for databases that have performance Peaks or heavy usage in different time periods because the amount of resources associated with the pool can be assigned to the databases that need them while the others are inactive elastic pool model is designed for resource sharing and it still does not support server level features such as SQL agents ser broker Etc these are other mechanisms that can be used as a replacement of these features such as elastic jobs and elastic queries instead of some server label features so now let's look at some of the key features of azure SQL database the first one is extensive monitoring and alerting capabilities so Azure SQL database provides Advanced monitoring and trouble shooting features that help you get deeper insights into workload characteristics these features and tools include the built-in monitoring capabilities provided by the latest version of the SQL Server database engine they enable you to find Real Time Performance insights also platform as a service monitoring capabilities provided by AO that enable you to Monitor and troubleshoot a large number of database instances query store a built-in SQL Server monitoring feature records the performance of your queries in real time and enables you to identify the potential performance issues and the top resource consumers automatic tuning and recommendation provides advice regarding the queries with the regress performance and missing or duplicated indexes automatic twinning in SQL database enables you to either manually apply the script that can fix the issues or let SQL database apply the fix SQL database can also test and verify that the fix provides some benefit and retain or rever the change depending on the outcome in addition to query store and automatic tning capabilities you can use Ed DMVs and XC event to monitor the workload performance Azure provides built-in performance monitoring and alerting tools combined with performance ratings that enable you to monitor the status of thousands of databases using these tools you can quickly asset the impact of scaling up or down based on your current or projected performance needs additionally SQL database can emit metrics and resource logs for easier monitoring you can configure SQL database to store resource usage workers and sessions and connectivity into one of these Azure resources so these resources are first one is azure storage so for achieving vast amounts of telemetry for a small price so the second one is azure event hubs for integrating SQL database Telemetry with a custom monitoring solution or hot pipelines and the third one is AZ monitor logs for a built in monitoring solution with reporting alerting and mitigating capabilities so the second feature is a availability capabilities so Azure SQL database enables your business to continue operating during disruptions in a traditional SQL Server environment you generally have at least two machines locally set up these machines have synchronously maintained copies of the data to protect against a failure of a single machine or component this environment provides High availability but it doesn't protect against a natural disaster destroying your data center Disaster Recovery assumes that a catastrophic event is geographically localized in a to have another machine or set of machines with a copy of your data for far away in SQL Server you can use always on availability groups running in asynchronous mode to get the capability people often don't want to wait for application to happen that far away from committing a transaction so there's potential for data loss when you do unplanned failovers so databases in the premium and business critical service ties already do something similar to the synchronization of an availability group dat bases in lower service tires provide rency through storage by using a different but equivalent mechanism like built-in logic helps protect against a single machine failure the active Geo replication feature gives you the ability to protect against disaster where a whole region is destroyed Azure availability zones strives to protect against the outage of a single Data Center building within a single region it helps you protect against the loss of power or network to a building in SQL database you place the different replicas in different availability zones in fact the service label management of azour powered by a Global Network of Microsoft managed data centers helps keep your app running 24/7 the Azure platform fully manages every database and it guarantees no data loss and a high percentage of data availability Azure automatically handles patching backups replication failure detection underlying potential Hardware software or network failures deploying Buck fixes failovers database upgrades and other maintenance tasks standard availability is achieved by a separation of compute and storage layers premium availability is achieved by integrating compute and storage on a single Lo for performance and then implementing technology similar to always on availability groups in addition SQL database provides built-in business continuity and Global scalability features these include automatic backups so SQL database automatically performs full differential and transaction log backups of database to enable you to restore to any point in time for single databases and pool databases you can configure SQL database to store full database backups to Azure storage for long-term backup retention for managed instances you can also perform copy only backups for long-term backup retention second is point in time restor so all SQL database deployment options Support Recovery to any point in time within the automatic backup retention period for any database third one is active Geo replication the single database and po databases option allow you to configure up to four readable secondary databases in either the same or globally distributed Azure data centers for example if you have a service as a platform application with a catalog database that has a high volume of concurrent read only transactions use active GE replication to enable global read scale this removes BX on the primary data due to read workloads for manage instances use autof fail groups fourth is autof failover groups all SQL database deployment options allow you to use failover groups to enable High availability and load balancing at global scale this includes transparent GE replication and failover of large sets of databases elastic pools and managed instances failover groups enable the creation of globally distributed Service as a platform applications with minimal Administration overhead this leaves all the complex monitoring routing and failover orchestrations to SQL database Fifth and the last one is Zone indendent databases SQL database allows you to provision premium or business critical databases or elastic pools across multiple availability zones because these databases and elastic pools have multiple rendent replicas for high availability placing these replicas into multiple availability zones provides higher resilience this includes the ability to recover automatically from the data center scale features without data loss so the next feature is built-in intelligence with SQL database you get buil-in intelligence that helps you dramatically reduce the costs of running and managing databases and that maximizes both performance and security of your application running millions of customers workloads around the clock SQL database collects and processes a massive amount of telemetry data while also fully respecting customer privacy various algorithms continuously evaluate the Telemetry data so that service can learn and adapt with your application so in its process of work first step is automatic performance monitoring and tuning so SQL database provides detailed insight into the queries that you need to monitor SQL database learns about your database patterns and enables you to adapt your database schema to your workload SQL database provides Performance Tuning recommendations where you can review tuning actions and apply them however constantly monitoring a database is hard and tedious task especially when you are dealing with many databases intelligent insights does this job for you by automatically monitoring SQL database performance at scale it informs you of performance degradation issues it identifies the root cause of each issue and it provides performance Improvement recommendations when possible managing huge number of databases might be impossible to do efficiently even with all available tools and reports that SQL database and Azure provide instead of monitoring and tuning your database manually you might consider delegating some of the monitoring and tuning actions to SQL database by using automatic tuning SQL database automatically applies recommendation tests and verifies each each of its tuning actions to ensure the performance keeps improving this way SQL database automatically adapts to your workload in a controlled and Safe Way automatic Twining means that the performance of a database is carefully monitored and compared before and after every tuning action if the performance doesn't improve the twinning action is diverted many of our partners that run Service as a platform multitenant apps on top of SQL database are relying on automatic Performance Tuning to make sure their applications always have stable and predictable performance for them this feature tremendously reduces the risk of having a performance incident in the middle of the night in addition because part of their customer base also uses SQL Server they are using the same indexing recommendations provided by SQL database to help their SQL Server customers two automatic tning aspects are available in SQL database so first one is automatic index management which identifies indexes that should be added in your database and indexes that should be removed second one is automatic plan correction which identifies problematic plans and fixes SQL plan performance problems the next step in the process of work is artive query processing you can use adaptive query processing including interl execution for multi statement table valued functions bash mode memory Grant feedback and bch Moree adaptive joints each of these adaptive query processing features apply similar learn and adapt techniques helping further address Performance issues related to historically inable optimization problems so the next feature is Advanced security and compliance SQL database provides a range of buil-in security and compliance features to help your application meet various security and compliance requirement note this that Microsoft has a certified Azure SQL database against number of compliance standards for more information see the Microsoft Azure press Center where you can find the most current list of SQL database compliance certifications so buil-in security and compliance features so there are certain buil-in security and compliance features so the first one is Advanced threat protection aure vender for SQL is a unified package for advanced SQL security capabilities it includes functionality for managing your database vulnerabilities and detecting anomalous activities that might indicate a threat to your database it provides a single location for enabling and managing these capabilities so there are two kinds of assessment in advanced protection the first one is vulnerability assessment this service can discover track and help you remediate potential database vulnerabilities it provides visibility in to your Security State and includes actionable steps to resolve security issues and enhance your database fortifications second is threat protection so this feature detect anomalous activities that indicate unusual and potentially harmful attempts to access or exploit your database it continuously monitors your database for suspicious activities and provides immediate security alerts on potential vulnerabilities SQL injection attacks and animous database access patterns threat detection alert provide details of the suspicious activity and recommend action on how to investigate and mitigate the threat so under this the second sub feature is auditing for compliance and security so auditing tracks database events and writes them to an audit log in your Azure storage account auditing can help you maintain Regulatory Compliance understand database activity and gain insight into discrepancies and anomalies that might indicate business concerns or suspected security violations so third one is data encryption SQL database helps secure your data by providing encryption for data addressed it uses transparent data encryption for data in use it uses always encrypted fourth feature is data Discovery and classification data Discovery and classification provides capabilities built into Azure SQL database for discovering classifying labeling and protecting the sensitive data in your databases it provides visibility into your database classification State and tracks the access to sensitive data within the database and Beyond its borders and the last one is azure active directory integration and multifactor authentication so SQL database enables you to centrally manage identities of database user and other Microsoft services with azured active directory integration this capability simplifies permission management and enhances security Azure active directory supports multiactor authentication to increase data and application security while supporting a single signin process the last major feature is easy to use tools SQL database makes building and maintaining applications easier and more productive SQL database allows you to focus on what you do best building great applications so you can manage and develop nsql database by using tools and skills you already have so the first tool is the Azure portal a web- based application for managing all Azure Services second is azure data Studio across platform database tool that runs on Windows Mac OS and Linux third is SQL Server management Studio a free downloadable client application for managing any SQL infrastructure from SQL Server to SQL database fourth is SQL server data Tools in Visual Studio a free downloadable client application for developing SQL Server relational databases databases in AQL database integration service packages analysis service data models and Reporting Services reports and the last tool is Visual Studio code a free downloadable open-source code editor for Windows Mac OS and Linux it support extensions including the mssql extension for quering Microsoft SQL Server Azure SQL database and Azure SS gentics now let's look at some of the use cases for Azure SQL database so the first one is developer or test environment it's an important use case for replicating or migrating data to SQL hosted on Azure is for developer or test environments before deploying to the production environment it is pertinent that the data is tested against developer or test environments Azure SQL databases can act as a target for just such environments the life production environment can be replicated to the developer or test environment using a database copy second is business continuity one of the most important use cases for SQL on is using it as a Dr Target to maintain business continuity aure SQL databases can provide an SLA of up to 99.99% by maintaining several copies of the data this provides business continuity as it allows you to restore Geor redundant copies of the data or use active GE redundant copies as failover points in case of outages at data centers or in regions besides Azure SQL databases you can also use availability groups to fulfill business continuity demands not only can you use availability groups in azour SQL virtual machines but also use azour SQL virtual machine instances as a target for high availability and Disaster Recovery third is scaling our readon workloads apart from providing BC or Dr capabilities active GE application can also be used to offload readon workload such as reporting jobs to secondary copies you can also extend on premises SQL Server instances using readable always on replicas fourth is backup and restore Azure SQL databases are bagged up automatically on a regular basis and there are no storage costs for up to 200% of the maximum provision database storage you can restore backups to any point in time going back to the retention period which is determined by the Azure SQL Service Tire in use on premises SQL Server databases and transaction locks can also be bagged up directly ly to Azure using the backup to URL feature and stored in Azure storage Azure SQL databases can also be stored on local storage by exporting them to backpack files means BC PS files and the last use case is Advanced analytics another important reason for hosting SQL in Azure is to make use of azure's advanced analytics platforms such as Azure storage blob and Azure data L store Common scenario with Advanced analytics is when users reference data from various data sources using Azure data Lake store as the stacking area perform transformation activities using Hive or spark and finally load the data into Azure data warehouse for business intelligence and Reporting now that you have a theoretical understanding of azure SQL database let's now see a deployment of a SQL database service on Microsoft Azure we will also connect this database with SQL management server as well as with Azure data studio so let's move ahead we just simply go to the Azure portal if you don't have an account on Azure what you can do is you can create a free account like can start free here and you will get a 12 months free services access as well as if you are from India the currency will be around like 14,500 INR means Indian rupees free currency we will get for free use that's the amount you will get so I already have an account so I will just log in from here no I don't want to I will just go to Microsoft as here this is the portal so I don't want to buy see these are the credits it is showing you get 14,000 500 credits I have used some of the credits and this much I left so what you can do is you can go to SQL databases right now because is I have a pin here but what you can do is you can go to SQL databases from here like here or you can search it from here SQL database yeah now let's create the database from here so fre trial is here we have chosen this then we can choose the resource Group can create one group so let's create with a new one so let's give the name as resource SQL demo okay so a new resource has been created then enter database name can give some name to like demo SQL so yeah that's okay so give the server name also so create a new server server name so let's give it like remember these names SQL demo server I have given so the specified server name is already in user we can give like SQL demo server 662 but you have to remember this don't forget this like login information okay so you can give some server admin login so Azure admin will be okay I think yeah now let's move ahead to give a password don't forget these information we will require these information later on in this process and choose your region also so my region is Central India so it will be Asia Pacific Central India just a minute yeah so okay the next process if you want to use SQL elastic pool right now we don't require so we will choose no if you want to then you can choose yes then we can configure the database and storage also so these are the plans like I have explained you basic standard premium the service tires are there similarly purchasing models are there voco purchasing models are there general purpose hypers scale business critical the also have explained to you so like whichever you choose on the basis of that it shows the price so right now it's costing 29,9 but we don't require that much bigger so you can just you can choose it from here also like not available because it's the free tire account that's why we don't have a premium version so can use the standard one so in this standard one if you decrease the storage then it will also decrease the price not in this manner like from you can go to basic here you get 2 GB and the price decreased you can choose for 1 GB also but price will remain the same so you can go to 2GB so we will choose the basic one because we don't require that much of configuration so just apply from here just review and create a simple review plan so yeah create so it's getting created the process it takes certain time so that's why taking a Time few minutes it take so you can see the deployment is complete so you can go to Resource SQL demo so we have came to Resource here you see that server is created but you have to refresh once again just a second where is the database let's refresh again I don't know it's get created usually yeah so you can see have to refresh it so it will takes a little time so database is also created now so what we can do from here is first step is you have to create the firewall so you can go to database so you have to set the firewall set firewall servers you can go from here set the firewall so you have to choose the like all these client IP is given from start IP and end IP so you can just you have to add the client IP so it has been added so you have to give it from 0 to 255 complete you have to give so save and just go back then you can see different cool features are given on the left side of this dashboard so you go to the query editor this is here you can like connect your database to the browser and you can configure and then there is comput and storage if you have like you can yeah not a problem so computer and storage is given if you want to like change your plan purchasing model and everything then you can change it from here again like all the things are given here you can again change it it's not a problem then there are like connection Str no problem then you can go to connection strings where you can connect your uh and make connection to the strings like there's ado.net or if you are working on Java then jdbc is there obbc is there PHP go everything is there it's quite cool feature then there are synchronous to other database where you can synchronize to other databases also then also you can like add Azure search also then there are different security features are there Advanced security features are there for data also Advanced Data security features are given here now what we can do major thing we have to do is we have to connect it to the management SQL Server so if you don't have a management SQL Server you can download it from here you can just give management SQL so just go here and download it from here it's given you can just click here and it will get downloaded it will ask for a restart it will restart the computer and all the settings will be saved in your computer then you can start using SQL Server I already have downloaded it so I already have downloaded it so we can just go here I have it print here so Microsoft SQL Server management so yeah you have to give your server name so let's see the server name first here it is know from database only so my server name is given so you can just copy it from here paste it here you can just select the SQL Server authentication here and give you a lockin also so remember the login I told you to remember the login password so my is your admin password can just connect it from here so yeah we got it here so we can just ose the databases can see the demo SQL we can find the our database here and there different tables and everything is there so we can just start the new query from here can create the table so let's create the table create table with the name Person persons so yeah it will be enough let's now execute it you can see command completed successfully we can now insert values into it so like before that we can view in tables Also the table might we created the name of persons expending so yeah you can see persons is created so let's now insert values you can give like give my name so6 from person let's execute now oh sorry it's a little mistake let's execute now so yeah you can see we have got the table contents everything is there now we can also see for Azure data studio also so we will just go to a studio it's like you have a visual studio code also similarly there's a studio can start the new connection then we have to give the server name here just like we have the server name there so what was the server name let's copy it from here so yeah just give here and then we will chose the SQL login username Azure admin the password yeah we can just remember the password not a problem then database we have to select so what was the name of the database demo SQL that's why it's showing here yeah so we connected from here yeah we here like different features are given like new query new notebook just like you have in Visual Studio code in that way we is given and all the other like it has a very good UI also here we go to the database first on it is showing just maximize it okay yeah so it's open now just like you can give it here Al all the details it is showing like it's been executed and details are shown here this is how it looks like also one cool thing about a studio is you can go here and you can see like it actually remembers all the databases like this one it remembers this one I have closed actually this database I have closed I have created it some time back but it still remembers databases if you give the aure like server ID now so like it will remember all the databases like from where initially we have selected now from there on so this is how it looks like let's go back to Azure one we have understood how to do it in Azure data studio also and how to connect it with the management SQL Server also so there are certain other features also you can see here for performance overview like how it's working and to track the performance and everything also the auditing and things are given here just like these are the things but I have mainly told you about the purchasing models and deployment models and how service ties are there and how we can create the database connected with management SQL Server as well as a studio I think it might be a little hectic but I have explained you in a little simpler manner so please try it with your hands-on experience like take your hands-on experience also by trying it on deploying as your SQL on this azzure portal after understanding Azure SQL database our next topic is azure data link here you will learn about the scalable data storage and analytics Services which allows for the handling of large amount of data Azure data Lake storage so what is azure data Lake storage Azure data Lake storage is a repository that stores large amount of raw data in its natural format until it is needed for analytics application so why is it named as data link well James Dixon the chief technology officer of pentaho is the person who has generally been credited with the coining of the term data link according to him he described a data M that is a subset of a data warehouse as a keing to a bottle of water which is cleansed packaged and structured for easy consumption while a data lake is more likely a body of water in its natural States data flows from the streams that is the source of the system to the lake users have access to Lake to examine take samples or dive in so a data lake is a centralized repository designed to store process and secure large amount of structured semi-structured and unstructured data it can store data in its native format and process any variety of it ignoring the size limits next is how to create a data Lake storage for this we'll have a practical demo to have a better understanding so the first step to work on any Azure Services is to first sign in so first we need to sign in make sure you do have an Azure account so that you can have the access to different Azure services so let let's sign in stay signed in so once you sign in you enter to this dashboard so you can see here this is a dashboard of your account and you can see different kinds of azure services like storage accounts monitors virtual machine Resource Group SQL database SQL manage instances and also you can see your subscript itions and different types of Resource Group you have created and all other stuffs so let's quickly get started first we need to go to create a resource so once you come here you can see popular res services so like virtual machine cuberty services Cosmos DB and rest other so for us we need to go to storage account once you click here so you enter to create a storage account here we need to create a resource Group so first of all what do you mean by Resource Group a resource Group is a container that holds related resources for an Azure solution it can include all these resources for the solution or only those resources that you want to manage as a group so let's create a new Resource Group let's name it as demo data link one okay and once you create a resource Group so you come to a storage account M so a storage account contains all of your assure storage data objects including blobs file shells cues tables and diss so let's name it as marshmallow 123 so once you have named your storage account we go to the advanc section here we will directly move to data Lake storage generation 2 so what do you mean by data L Storage generation 2 Data L Storage generation 2 is designed to deal with this variety and volume of data at exhibit scale while securely handling hundreds of gigabytes of through output with this you can use data leg storage generation 2 on the basis of both real time and back solution so here we will enable the hierarchical name space once we enable it we'll click on review plus create so it will take few minutes to validate all your details once validation is passed you just click on create so it's now getting ready to deploy so here you can see your marshmallow 1 2 3 is created and now it is getting ready for deployment here you can see deployment is in progress once it gets ready so we will be working on it here you can see your deployment is complete now what you going to do is we'll move on to go to resources now here you can see your details your resource gr details the whole storage Account Details basically and even you can see the properties enabled with it so here we have data L Storage file service Q service table service networking security so after this we'll quickly go to containers and we'll create a new container for our storage account so let's click on container and we give it a name to it make sure you give a name it should be in lower case because they don't accept the upper case so let's keep it demo and create so yeah you can can see your container has been created let's click on it so here you can see there are no results as we haven't added any stuff so let's upload for this you can either upload it from aure Portal or else you can also go to storage Explorer so let's see how do we do on the Azure portal so once you come here you can just click on select a file let's take any picture let's see we'll upload a picture so here you can see AWS 3.png let's upload it so here you can see AWS PNG has been uploaded in your container same as we can do on storage Explorer too for that you need to download a storage Explorer in my case I have already downloaded the storage Explorer so let's quickly go into it so here also make sure you are signed in so that you can see all your containers and Resource Group which you have created so here you can see my storage account that isow 23 which we have created recently under this we'll go and find our container yeah so here we go to blob containers and we can see a demo so here you can see the image file which we had uploaded through our a portal now we'll upload a file and a folder both let's see so let's upload a file first so you click on upload upload files and then you select a file let's take any file let's take a picture again so A4 let's take this an image so now here you can see your image is transferring from your path to our demo yeah so here your image file is uploaded so as we said we can upload any kind of datas maybe structured or unstructured so let's check out by uploading a folder so here you go with the same process and you you select your folder let's take any folders let's see if we do have any folder okay I'll just take any one of my folders and just click and upload so your folder is also being uploaded over here and once your folder is uploaded now you can see the inside resources into it so here there were different files text image all of them are there you can access it from here itself and you can see other operation as well if you want to download any of the file or folder you can do it or if you want to open it let's open it any of the uh files see let's see if it is getting open or no okay yes so yeah you can see this file has been open now if you want to download this file let's download sequence so we'll download it and and just put it in downloads let's see if it is downloaded or no apply to apply so yeah it tells your download is completed let's see if you find it or no we go to our file and downloads and here see now I had downloaded this is the file which I had downloaded through our storage Explorer so this is how you manage your datas through your data leg storage so apart from this you can also give permissions to different users like whichever file if you want to give an access to a particular user then you can also give permission to them that they can either read write and access the whole file or a folder so for that you just need to click on a file or a folder and right click and you just see here manage Access Control list so once you come here so you can here you can see there are different owners so user owner and all so here I can add an owner like for this file who can just get an access to it so you can find out any name if you find out any relatable person or a user then you can give an access to that currently I don't have anyone so I won't be able to show you that but yes this is how you add or give permission to different users you can also do it with a folder or else you can also do it with the whole container as well if you want to share your containers with different people or different users then you can easily share them so here also you just go right click and just come to manage access control and just give them the access click on ADD and find out the person whom you want to give the access to and just after that once you give them the permission here you can see if you want to permit them for only read or only write or if you want to give read and write or all of the three so you can just give them the permissions accordingly and click on okay and then the particular user gets the access to all your files and folders so this is how we create and work with Azure data Lake storage now let us see the comparison between Azure blob storage and the data L Storage so here are some of the comparisons between Azure data L Storage and Azure block storage Azure data Lake storage is a technique of planning and control of the time whereas Azure block storage is an object stored with a flat name space AO data lake is an optimized storage for big data analytics workload whereas AZ your blob storage is basically a general purpose Object Store for for a wide variety of storage scenarios which also include big data analytics in Azo data Lake storage the apis are over https only whereas in Blob storage the rest API is over the HTTP as well as the https in Azure data Lake storage there is no limits on the account but in Azure blob storage there are specific limits for container sizes and the files in the blog so this was the major points which differentiate Azure block storage with Azure data Lake storage at last we come to the use cases there are many use cases of data Lake storage out of which we'll discuss about four so at first business intelligence on data Lake storage so data Lake storage dramatically improves the speed for ad hoc queries dashboards and remotes you can run existing bi tools on lower cost data laks without compromising performance or data quality it also avoids costly delays adding new data sources and the reports at second we see cloud data Lake migration here we can optionally deploy new applications to the cloud using data Lake storage such as S3 or ADLs you c

Original Description

🔥𝐄𝐝𝐮𝐫𝐞𝐤𝐚'𝐬 𝐃𝐚𝐭𝐚 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠 𝐂𝐞𝐫𝐭𝐢𝐟𝐢𝐜𝐚𝐭𝐢𝐨𝐧 𝐓𝐫𝐚𝐢𝐧𝐢𝐧𝐠 𝐂𝐨𝐮𝐫𝐬𝐞 (𝐔𝐬𝐞 𝐂𝐨𝐝𝐞: 𝐘𝐎𝐔𝐓𝐔𝐁𝐄𝟐𝟎) : https://www.edureka.co/microsoft-azure-data-engineering-certification-course In this Data Engineering Course, you will learn what is Data Engineering, How to Become a Data Engineer, Why Data Engineering is important, What is Big Data Engineering, the Importance of Big Data, the Difference between Data Engineer and Data Scientist, Big Data Engineering, What is Hadoop, Apache Spark Tutorial, AWS Elastic Map Reduce Tutorial, Azure Data Tutorial, Data Engineering Career Roadmap and lot more. This Data Engineer Full Course will help you to understand Data Engineering from scratch. 00:00:00 Introduction 00:02:14 What is Data Engineering? 00:20:08 Quiz 00:20:33 How to Become a Data Engineer? 00:22:18 Introduction to Big Data 00:54:26 How to Become a Big Data Engineer? 01:09:04 Big Data Engineer Salary 01:12:27 How to Become Azure Data Engineer? 01:29:21 Azure Data Factory 01:59:27 Quiz 01:59:49 Azure Database Services 02:30:40 Quiz 02:31:03 Azure SQL Database 03:14:20 Azure Data Lake 03:29:05 Quiz 03:29:30 Advanced Data Modeling with Power BI and Azure 03:33:52 Azure Databricks 04:07:21 Quiz 04:07:38 Introduction to Hadoop 05:13:40 Quiz 05:14:02 Hadoop Ecosystem 05:32:06 How to Install Hadoop on Windows 10 05:46:41 Apache Sqoop 06:04:41 Quiz 06:05:05 Apache Pig 06:26:24 Quiz 06:26:45 Apache Hive 07:27:23 Quiz 07:27:43 Hadoop Projects 08:02:53 What are Kafka Streams? 08:16:08 Quiz 08:16:31 Big Data Hadoop Interview Questions and Answers 🔴 Subscribe to our channel to get video updates. Hit the subscribe button above: https://goo.gl/6ohpTV 🔴 𝐄𝐝𝐮𝐫𝐞𝐤𝐚 𝐎𝐧𝐥𝐢𝐧𝐞 𝐓𝐫𝐚𝐢𝐧𝐢𝐧𝐠 𝐚𝐧𝐝 𝐂𝐞𝐫𝐭𝐢𝐟𝐢𝐜𝐚𝐭𝐢𝐨𝐧𝐬 🔵 DevOps Online Training: http://bit.ly/3VkBRUT 🌕 AWS Online Training: http://bit.ly/3ADYwDY 🔵 React Online Training: http://bit.ly/3Vc4yDw 🌕 Tableau Online Training: http://bit.ly/3guTe6J 🔵 Pow
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from edureka! · edureka! · 0 of 60

← Previous Next →
1 ChatGPT Not Working - 4 Fixes | How To Fix ChatGPT Not Working | Why Is ChatGPT Not Working |Edureka
ChatGPT Not Working - 4 Fixes | How To Fix ChatGPT Not Working | Why Is ChatGPT Not Working |Edureka
edureka!
2 Advanced Java script Tutorial | JavaScript Training | JavaScript Programming | Edureka Rewind
Advanced Java script Tutorial | JavaScript Training | JavaScript Programming | Edureka Rewind
edureka!
3 Java script interview question and answers | Java script training | Edureka Rewind
Java script interview question and answers | Java script training | Edureka Rewind
edureka!
4 OpenAI API Tutorial using Python | How to use OpenAI GPT-3 API - Ada Babbage Curie Davinci | Edureka
OpenAI API Tutorial using Python | How to use OpenAI GPT-3 API - Ada Babbage Curie Davinci | Edureka
edureka!
5 What is Unsupervised Learning ? | Unsupervised Learning Algorithms| Machine Learning | Edureka
What is Unsupervised Learning ? | Unsupervised Learning Algorithms| Machine Learning | Edureka
edureka!
6 Top 10 Applications of Machine Learning in 2023 | Machine Learning  Training | Edureka Rewind - 7
Top 10 Applications of Machine Learning in 2023 | Machine Learning Training | Edureka Rewind - 7
edureka!
7 Machine Learning Engineer Career Path in 2023  | Machine Learning Tutorial | Edureka Rewind - 6
Machine Learning Engineer Career Path in 2023 | Machine Learning Tutorial | Edureka Rewind - 6
edureka!
8 10 Must Have Machine Learning Engineer Skills That Will Get You Hired   | Edureka Rewind - 7
10 Must Have Machine Learning Engineer Skills That Will Get You Hired | Edureka Rewind - 7
edureka!
9 Data Structures in Python | Data Structures and Algorithms in Python | Edureka | Python Live - 5
Data Structures in Python | Data Structures and Algorithms in Python | Edureka | Python Live - 5
edureka!
10 Python Lists | List in Python | Python Training  | Edureka  Rewind
Python Lists | List in Python | Python Training | Edureka Rewind
edureka!
11 Predictive Analysis Using Python | Learn to Build Predictive Models | Python Training | Edureka
Predictive Analysis Using Python | Learn to Build Predictive Models | Python Training | Edureka
edureka!
12 Machine Learning Tutorial | Machine Learning Algorithm | Machine Learning Engineer Program | Edureka
Machine Learning Tutorial | Machine Learning Algorithm | Machine Learning Engineer Program | Edureka
edureka!
13 How to use Pandas in Python | Python Pandas Tutorial  | Python Tutorial  |  Edureka  Rewind
How to use Pandas in Python | Python Pandas Tutorial | Python Tutorial | Edureka Rewind
edureka!
14 Parameters in Tableau | Tableau Parameters Examples | Tableau Tutorial  | Edureka Rewind
Parameters in Tableau | Tableau Parameters Examples | Tableau Tutorial | Edureka Rewind
edureka!
15 Top 10 Reasons to Learn Tableau in 2023  | Tableau Certification | Tableau | Edureka Rewind
Top 10 Reasons to Learn Tableau in 2023 | Tableau Certification | Tableau | Edureka Rewind
edureka!
16 Tableau Developer Roles & Responsibilities | Become A Tableau Developer | Tableau | Edureka Rewind
Tableau Developer Roles & Responsibilities | Become A Tableau Developer | Tableau | Edureka Rewind
edureka!
17 Deep Learning With Python | Deep Learning Tutorial For Beginners | Edureka  Rewind
Deep Learning With Python | Deep Learning Tutorial For Beginners | Edureka Rewind
edureka!
18 Realtime Object Detection  | Object Detection with TensorFlow | Edureka | Deep Learning Rewind - 2
Realtime Object Detection | Object Detection with TensorFlow | Edureka | Deep Learning Rewind - 2
edureka!
19 Top 20 Tableau Tips and Tricks in 20 Minutes | Tableau Tutorial | Tableau Training  | Edureka Rewind
Top 20 Tableau Tips and Tricks in 20 Minutes | Tableau Tutorial | Tableau Training | Edureka Rewind
edureka!
20 Climate Change Prediction using Time Series | Python Projects | Edureka | DS Rewind -  5
Climate Change Prediction using Time Series | Python Projects | Edureka | DS Rewind - 5
edureka!
21 ReactJS Installation Tutorial | ReactJS Installation On Windows | ReactJS Tutorial | Edureka Rewind
ReactJS Installation Tutorial | ReactJS Installation On Windows | ReactJS Tutorial | Edureka Rewind
edureka!
22 Phases in Cybersecurity  | Cybersecurity Training | Edureka | Cybersecurity Rewind - 2
Phases in Cybersecurity | Cybersecurity Training | Edureka | Cybersecurity Rewind - 2
edureka!
23 What Is React | ReactJS Tutorial for Beginners | ReactJS Training | Edureka Rewind
What Is React | ReactJS Tutorial for Beginners | ReactJS Training | Edureka Rewind
edureka!
24 Cybersecurity Frameworks Tutorial | Cybersecurity Training | Edureka | Cybersecurity Rewind- 2
Cybersecurity Frameworks Tutorial | Cybersecurity Training | Edureka | Cybersecurity Rewind- 2
edureka!
25 React vs Angular 4  | Angular 2 vs React | React & Angular | ReactJS Training | Edureka Rewind - 5
React vs Angular 4 | Angular 2 vs React | React & Angular | ReactJS Training | Edureka Rewind - 5
edureka!
26 ReactJS Components Life-Cycle Tutorial  | React Tutorial for Beginners  | Edureka Rewind
ReactJS Components Life-Cycle Tutorial | React Tutorial for Beginners | Edureka Rewind
edureka!
27 Ethical Hacking using Kali Linux | Ethical Hacking Tutorial | Edureka | Cybersecurity Rewind - 3
Ethical Hacking using Kali Linux | Ethical Hacking Tutorial | Edureka | Cybersecurity Rewind - 3
edureka!
28 Types Of Artificial Intelligence | Artificial Intelligence Explained | What is AI? | Edureka
Types Of Artificial Intelligence | Artificial Intelligence Explained | What is AI? | Edureka
edureka!
29 Top 10 Applications Of Artificial Intelligence in 2023 | Artificial Intelligence| Edureka Rewind
Top 10 Applications Of Artificial Intelligence in 2023 | Artificial Intelligence| Edureka Rewind
edureka!
30 The Future of AI | How will Artificial Intelligence Change the World in 2023? | Edureka Rewind
The Future of AI | How will Artificial Intelligence Change the World in 2023? | Edureka Rewind
edureka!
31 What is Artificial Intelligence | Artificial Intelligence Tutorial For Beginners | Edureka Rewind
What is Artificial Intelligence | Artificial Intelligence Tutorial For Beginners | Edureka Rewind
edureka!
32 Google Cloud IAM | Identity & Access Management on GCP  | Edureka | GCP Rewind - 5
Google Cloud IAM | Identity & Access Management on GCP | Edureka | GCP Rewind - 5
edureka!
33 Google Cloud AI Platform Tutorial | Google Cloud AI Platform   | GCP Training | Edureka Rewind
Google Cloud AI Platform Tutorial | Google Cloud AI Platform | GCP Training | Edureka Rewind
edureka!
34 Projects in Google Cloud Platform  | GCP Project Structure  | GCP Training | Edureka Rewind
Projects in Google Cloud Platform | GCP Project Structure | GCP Training | Edureka Rewind
edureka!
35 How to Become a Data Scientist | Data Scientist Skills | Data Science Training  | Edureka Rewind - 3
How to Become a Data Scientist | Data Scientist Skills | Data Science Training | Edureka Rewind - 3
edureka!
36 Agglomerative and Divisive Hierarchical Clustering Explained | Data Science Training | Edureka Live
Agglomerative and Divisive Hierarchical Clustering Explained | Data Science Training | Edureka Live
edureka!
37 Climate Change Prediction using Time Series | Python Projects | Edureka | DS Rewind -  5
Climate Change Prediction using Time Series | Python Projects | Edureka | DS Rewind - 5
edureka!
38 Data Science Project - Covid-19 Data Analysis | Python Training | Edureka | DS Rewind - 6
Data Science Project - Covid-19 Data Analysis | Python Training | Edureka | DS Rewind - 6
edureka!
39 What is Honeycode? | Introduction to Honeycode | Edureka
What is Honeycode? | Introduction to Honeycode | Edureka
edureka!
40 Difference between Amazon AWS and Google Cloud | GCP Training Google Cloud | Edureka Live
Difference between Amazon AWS and Google Cloud | GCP Training Google Cloud | Edureka Live
edureka!
41 DevOps Lifecycle | Introduction To DevOps | DevOps Tools | What is DevOps? | Edureka Rewind
DevOps Lifecycle | Introduction To DevOps | DevOps Tools | What is DevOps? | Edureka Rewind
edureka!
42 Introduction to DevOps | DevOps Tutorial for Beginners | DevOps Tools | DevOps | Edureka Rewind
Introduction to DevOps | DevOps Tutorial for Beginners | DevOps Tools | DevOps | Edureka Rewind
edureka!
43 How to Create Login System using Python | Python Programming Tutorial | Edureka Rewind
How to Create Login System using Python | Python Programming Tutorial | Edureka Rewind
edureka!
44 Python Developer | How to become Python Developer | Python Tutorial  | Edureka Rewind
Python Developer | How to become Python Developer | Python Tutorial | Edureka Rewind
edureka!
45 How to become a Data Engineer | Complete Roadmap to become a Data Engineer| Data Engineer |  Edureka
How to become a Data Engineer | Complete Roadmap to become a Data Engineer| Data Engineer | Edureka
edureka!
46 Azure Data Engineer Certification [DP 203] | How to Become Azure Data Engineer [2023] | Edureka
Azure Data Engineer Certification [DP 203] | How to Become Azure Data Engineer [2023] | Edureka
edureka!
47 Data Analyst vs Data Engineer vs Data Scientist | Data Analytics Masters Program  | Edureka Rewind
Data Analyst vs Data Engineer vs Data Scientist | Data Analytics Masters Program | Edureka Rewind
edureka!
48 DevOps Engineer day-to-day Activities | DevOps Engineer Responsibilities | Edureka Rewind
DevOps Engineer day-to-day Activities | DevOps Engineer Responsibilities | Edureka Rewind
edureka!
49 How to Become a DevOps Engineer?  | DevOps Engineer Roadmap | Edureka | DevOps Rewind
How to Become a DevOps Engineer? | DevOps Engineer Roadmap | Edureka | DevOps Rewind
edureka!
50 How to Become a Data Engineer? | Data Engineering Training | Edureka
How to Become a Data Engineer? | Data Engineering Training | Edureka
edureka!
51 How To Become A Big Data Engineer? | Big Data Engineer Roadmap | Edureka Rewind
How To Become A Big Data Engineer? | Big Data Engineer Roadmap | Edureka Rewind
edureka!
52 Python Integration for Power BI and Predictive Analytics | Power BI Training | Edureka
Python Integration for Power BI and Predictive Analytics | Power BI Training | Edureka
edureka!
53 Power BI KPI Indicators Tutorial | Custom Visuals In Power BI | Power BI Training  | Edureka Rewind
Power BI KPI Indicators Tutorial | Custom Visuals In Power BI | Power BI Training | Edureka Rewind
edureka!
54 Apache HBase Tutorial For Beginners | What is Apache HBase? | Big Data Training | Edureka Rewind
Apache HBase Tutorial For Beginners | What is Apache HBase? | Big Data Training | Edureka Rewind
edureka!
55 Big Data Hadoop Tutorial For Beginners  | Hadoop Training | Big Data Tutorial  | Edureka  Rewind
Big Data Hadoop Tutorial For Beginners | Hadoop Training | Big Data Tutorial | Edureka Rewind
edureka!
56 Big Data Analytics  | Big Data Analytics Use-Cases | Big Data Tutorial | Edureka Rewind
Big Data Analytics | Big Data Analytics Use-Cases | Big Data Tutorial | Edureka Rewind
edureka!
57 What Is Power BI? | Introduction To Microsoft Power BI | Power BI Training  | Edureka  Rewind
What Is Power BI? | Introduction To Microsoft Power BI | Power BI Training | Edureka Rewind
edureka!
58 Triggers in Salesforce | Salesforce Apex Triggers | Salesforce  Tutorial  | Edureka Rewind
Triggers in Salesforce | Salesforce Apex Triggers | Salesforce Tutorial | Edureka Rewind
edureka!
59 How To Become A Salesforce Developer | Salesforce For Beginners| Salesforce Training  Edureka Rewind
How To Become A Salesforce Developer | Salesforce For Beginners| Salesforce Training Edureka Rewind
edureka!
60 Java ArrayList Tutorial | Java ArrayList Examples | Java Tutorial | Edureka Rewind
Java ArrayList Tutorial | Java ArrayList Examples | Java Tutorial | Edureka Rewind
edureka!

Related Reads

📰
I Built My Second ETL Pipeline. This Time, I Started Thinking Like a Data Engineer
Learn how to build a production-ready ETL pipeline with Python, Docker, PostgreSQL, and Kestra by thinking like a data engineer
Towards Data Science
📰
JuiceFS Sync for PB-Scale Data Transfers: Resumable Sync, Encryption, and Bandwidth Control
Learn how to efficiently transfer large volumes of data using JuiceFS Sync, which offers resumable sync, encryption, and bandwidth control, ideal for PB-scale data transfers.
Dev.to AI
📰
How Airflow is using AI to make data engineering more resilient, not more complex
Airflow uses AI to make data engineering more resilient by detecting data drift, resuming failed pipelines, and fixing issues automatically, reducing complexity and improving reliability.
Medium · AI
📰
What Can We Do When Memory Becomes the New Bottleneck in Data Engineering?
Learn how to overcome memory bottlenecks in data engineering using Pandas chunking, Dask, and Polars, and why it matters for processing large datasets
Towards Data Science

Chapters (32)

Introduction
2:14 What is Data Engineering?
20:08 Quiz
20:33 How to Become a Data Engineer?
22:18 Introduction to Big Data
54:26 How to Become a Big Data Engineer?
1:09:04 Big Data Engineer Salary
1:12:27 How to Become Azure Data Engineer?
1:29:21 Azure Data Factory
1:59:27 Quiz
1:59:49 Azure Database Services
2:30:40 Quiz
2:31:03 Azure SQL Database
3:14:20 Azure Data Lake
3:29:05 Quiz
3:29:30 Advanced Data Modeling with Power BI and Azure
3:33:52 Azure Databricks
4:07:21 Quiz
4:07:38 Introduction to Hadoop
5:13:40 Quiz
5:14:02 Hadoop Ecosystem
5:32:06 How to Install Hadoop on Windows 10
5:46:41 Apache Sqoop
6:04:41 Quiz
6:05:05 Apache Pig
6:26:24 Quiz
6:26:45 Apache Hive
7:27:23 Quiz
7:27:43 Hadoop Projects
8:02:53 What are Kafka Streams?
8:16:08 Quiz
8:16:31 Big Data Hadoop Interview Questions and Answers
Up next
A Moment Frozen in Time | Arnav Iyengar | TEDxJenks Youth
TEDx Talks
Watch →