The DataHour: Energy Data Science Project from Scratch

Analytics Vidhya · Beginner ·📐 ML Fundamentals ·3y ago

Key Takeaways

The DataHour: Energy Data Science Project from Scratch covers the fundamentals of data science in the oil and energy sector, including data processing and project development from scratch using tools like Python, Pandas, and scikit-learn.

Full Transcript

hello and welcome everyone we are really happy to have you all here this evening for a session full of action-packed learning i am priyanka for a part of the data science team of analytics vidya and i'll be the moderator of this session uh my colleague anand will be helping you with any queries that you will be having and so for those who have joined us for the first time a brief introduction to the data session the data r is a series of webinars conducted by analytics and led by top industry experts it is a fun way to understand the concepts of data science from the leading players in the data tech domain and as the name suggests it's one hour dedicated to data we are hopeful that these sessions are going to be a great source of enrichment and value adding to our community members i hope you are excited to attend this data hour with us and before we kick start uh things i will uh just help you with a quick recap of the housekeeping items we are recording the session and we will make the recording available in few days on our youtube channel please use the q a section for asking any questions you might have during the session and we will do our best to answer them as the data hour progresses or towards the end also we will share a poll about the feedback towards the end of the session which i request you all to kindly fill up now on to our session today which is energy data science project from scratch in this data hour you will get a good sneak peek into what kind of data processing and projects happen in the oil and energy sector you will gain an understanding of raw data from an oil producing field that is the oil well data as well as the water injector well data we will talk about null handling feature engineering outlier detection descriptive statistics etc at the end introductory machine learning will be explained to conclude the session and now on to our speaker in this session of data hour we have divyanshwagyas with us he is the the energy data scientist and founder at petroleum from scratch diviancho is a machine learning and data science consultant at oil and energy projects and a community member for python and data sciences for the energy industry with the moto to be the perfect mix of physics and mathematics over to your diversion the virtual stage is all yours ah thank you thank you for that introduction uh thank you everyone for joining in we really don't have too much time and i'm expecting that a lot of you are from the domain which not from the domain which i'm from so i come from the petroleum engineering domain i come from the oil and gas industry and i'm exploring the various opportunities that lie in in the intersection of the data science data sciences with the oil and gas industry the various problems that can be solved uh using various machine learning algorithms for the oil and gas industry and uh with with the experience that i've had i've realized that uh the boom is just very new and a lot of more opportunities are growing as we as we step into this in this industry uh and as far as i can look at it uh the energy industry has been a bit slow to adapt to the data uh revolution and which is again a blessing in disguise for someone who wants to you know get get into the data science industry uh if energy industry fascinates you oil and gas electric vehicles all these things they fascinate you this is a good market to start your leg in so i'll share my screen let me know once it is visible and we will start things um i'll i'll i saw the poll i i i saw that a lot of people are new to the data science industry and some of them are really experienced so i hope uh the session brings good light to both kind of audiences uh for those who are well versed with data sciences maybe an introduction to a typical oil and gas project can help for those who do not know either of them i mean this is a i hope this becomes a good session to start with so i'll share my screen let me know once you can see it uh good to see some people from oil and gas industry sharing my screen let me know once you can see it yeah can you see my all right thank you okay so uh basically let's let's just start with uh the typical possibilities uh what happens in the oil and gas industry uh a quick two minute introduction about uh what basically it is all about for a layman so as as the part of the industry there are various teams that work in uh our basic aim is to figure out what rock has oil in it and how can we extract oil from like thousands of feet downwards uh in the most efficient manner so if you if you just imagine that a piece of rock have has oil in its pores it does not just give away oil right we have to carry out a lot of things in order to extract as much oil as possible from uh and the problem is that rock is lying thousands of feet below the surface so there is a team of geologists and geophysicists they work together uh in order to build a simulation a static simulation of the entire uh we call it the subsurface which is a layer of rocks above which you are standing and the various signatures tell us that where oil can be possibly located and then we as oil and gas may be drilling engineers we drill holes up to that point and our aim is to you know extract oil from the hole which is drilled now the problem is that initially at least let's say five to ten percent of oil will just flow onto the surface because of the pressure that it has in the reservoir right we call it the reservoir whenever i use the word reservoir i'm just talking about the rock which is containing oil lying thousands of feet below us right now some of the oil will just flow up because it has a lot of pressure and you just open up a vent it's almost as if you open up a coca-cola bottle and the fluid just starts popping out because it has a lot of pressure but slowly it keeps on dying the pressure the flow of the fluid it keeps on dying right and then we need to use a lot of other ways to keep on pushing the oil to the surface we drill other wells through which we inject some water we drill some other wells through which we inject some surfactant uh the detergents and also that the oil feels easier to move so we try a lot of things and that totally uh sums up the entire oil and gas industry and i'm not using any any uh terminologies that is hard for a non oil and gas engineer to understand but this is the entire value chain and then once the oil comes up to the surface uh we try to refine it uh we try to you know extract the usual components out of it and then we try to transport it which is called the downstream side of the oil and gas industry now in the entire value chain if you if you think about the oil and gas industry if you have seen pictures on the internet you must have seen a lot of huge rigs right you must have seen a lot of processing power plants just like a water processing power plant you must have seen a lot of uh you know rigs uh by by various industrial giants right now initially the the thing was uh if i talk about now if i talk about the context of what data science can bring into this industry it can bring in safety let's start with that okay let's start with with talking about the possibilities that data science can bring into the energy industry possibilities suppose you have a huge uh refining power plant or maybe some offshore structure which is lying like thousands of uh you know feet away or some miles away from the surface within center of the ocean uh you have to make sure that you follow a predictive maintenance path right predictive maintenance path right you cannot just expect that some failure happens and you will just reverse to trying to repair the entire thing because a failure in oil and gas industry if anyone knows about the deep water horizon the failure in the oil and gas industry can mean a lot of lives and it can mean a lot of you know non-productive time a lot of loss to the industry so uh the initial way of approaching any failure in the industry was reactive maintenance that we would see some sort of failure happening then we would repair it but now with the advent of huge computational power we can use the sensor data which is coming to uh to the data scientists of the energy industry we can develop on stream predictive maintenance model that can suggest that hey within one hour from now based on the historical data that i have trained my algorithm on i can say that some failure is going to happen uh maybe 40 minutes from now one hour from now two hours from now so this is a major major dominant uh you know project in in this industry it can be in other industries as well uh and the project that we are today we are going to look at uh is is something using which we will use oil production prediction uh with the interest of time i will take the questions later on um maybe after i i complete a good chunk of what i want to cover and then we will look at the few the questions in the chat so i hope that's okay yes that is absolutely fine okay thanks so the the project we are going to look at today is predicting oil production rates from the oil field data now the aim behind this is like i said we are we are producing oil from a huge field of rock and we are trying to produce as much oil as possible everyone knows about the oil prices so every barrel of oil that we are producing it is very very very very it's like gold if the oil prices are huge so oil we need to make sure whatever we do we make sure we produce as much oil as possible now why does a machine learning model that can forecast or predict oil production help how can it help it can help in the fact that we can accurately you know let me write it down accurately estimate the results for example i i run a forecasting algorithm that tells me that hey uh over the next one month or over the next one year this is the oil production which is going to happen as as per my machine learning model uh that tells the entire economic analyst in the oil and gas industry how much cash flow we can expect in the coming times now if that cash flow uh if they do the break-even analysis if that cash flow uh tells them that the the operation of this field is going to be profitable then they carry on with the production of the field if that cash flow expected in the future is is going to cost them some loss because they are investing too much then they would stop producing from that field so this this is a huge opportunity for people uh working in in in the in the simulation space to use machine learning as a quick solving algorithm that can tell them uh one it can tell them the accurate uh forecast it can guide them to what kind of oil production is expected based on what kind of operations and it can help them decide whether they should carry on with uh the production in this field or not right so let's start with uh with uh the basic stuff in in python if if no one knows about anything about python you are you are listening to a data science video for the first time uh i'll try to introduce three things that are useful uh for uh for very useful for data science projects uh as far as raw python goes so a bit of python two minutes just two minutes you need to know about lists what is a list for example uh a list is a collection of numbers for those who already know this please bear with me this is for those who do not know this is a list what is a dictionary dictionary is something which will store the the data but in in a labeled format a has values one comma two comma three and b has values four comma five comma six if you wanna store uh a data in some other format uh let's say you want to store your name it will be in the format of a strings right it will be let's say dv or something like that okay that's as easy as uh it is in python uh i don't think we need to go deep initially as we go on in the session we will figure out what is needed we'll need a few packages we'll need numpy which is uh uh the it's it's it's like the calculator in python we'll need pandas it's like the microsoft excel in python we'll need matplotlib which is a way to visualize things in python and cbon which is again a way better way a fancy way of visualizing things in python so we will import a few libraries uh and these are the few libraries that we need to import right so we import numpy as np pandas as pd you can import numpy as numpy itself the the pet name is not required you can use your own pet name but the the data science industry it works on conventions so this is a standard convention people use numpy as mp and so on right okay now let's come to the important part the data that we will be using in this session data set that we'll be using in this session uh i've stored that data in my github repository and of course i mean for whoever who could not grasp much of it i have all these projects uh stored in my github i have a huge number of projects so if in case you want to come back and look at it this industry interests you feel free to have a look at it and uh the link will be shared right so the the path i'm copying the path to the file okay this is the path to the file all right so this is the data set path i've located this in uh this is from uh open source field uh which which is uh you know called the wall field a field is nothing but a hectare hectares of area underneath which uh the geologists have figured out that there is some oil residing there over with some producing production is happening so the data set path is this let's import a data using uh python pandas okay so the way to do that pd.read csv because it's a csv format file we'll just path parse the data set path and this is our data set okay now let's talk about what kind of data you can expect in a typical data science project and what data set this is this is a time series data set every recording is observed on a particular day and which is a very common data set in the manufacturing energy or any sort of hardware industry wherever sensors are located or flow meters or whatever wherever recording meters are located uh there is always a timestamp associated with the values which are recorded right other kind of data sets can include a non-time series kind of data set where the order of sequence of elements does not matter in this case uh whatever happened on this day and whatever happened on this day the sequence is important the next day happens after the first day you cannot assume that these two dates are independent of each other i mean they are independent uh they don't they don't necessarily have a formula that defines that tomorrow will formula for follow this formula but the pattern is important the pattern follows a particular decline that is important to us right so the sequence is very important as far as time series data sets are concerned all right and one of the first things that we do in in importing a time series data set from a producing field is to make sure that the the date column is at the index and the way to do that is to tell python that hey can you keep the index column which is the zeroth column as your index uh python does that as i run it you will see that the date column just shifted its way to the left hand side as an index it's in the bold now now it's still an issue right now because uh this is by default assumed to be a string in python because it's having underscores in it so you have to tell python that can you please understand or infer the index column to be of the date time format because if you let it be the way it was if you let it be a string that it was it it would uh not help you in plotting it will not help you in various time series related processing steps almost uh you can subtract dates you can figure out a date range etc but all that kind of processing requires python to understand that this is a date time object so the way to do that is to tell python that hey can you pass that index to be a datetime index kind of thing parse dates equals to true and once you do that all sorted now you have to store that in your df now the next thing which will make things easier for you is to explain what are all the columns in this data set right so this is the data set coming on from a typical oil and gas field okay and it is uh if i can annotate it let me just use a drawing i'll be using my mouse so pardon me so i'm drawing it okay so okay it's not working never mind so a typical oil and gas field is nothing but uh you have a huge hectare or a field of it's it's almost like a farm which is isolated from the from the outer universe or the outer world you you barricade it and within that field what you do is you set up a rig and that rig drills uh and and then once the hole is drilled you start producing and various equipments that are used in in that production of oil has a stipulated name given so for example to start with you must have realized that you must have realized that oil is flowing that means some kind of opening must be there some kind of choke must be there almost like a tap through which water flows so you can see there is a choke size feature here this tells us how much have you opened the surface choke it's almost as if how much of the tap you have opened to produce from in oil and gas you never open the choke fully you open it someday you will open it half somewhere you will not open it too much uh pressure maintenance is very important so that's why this column is nothing but the choke size okay on stream injection hours tells you in one day in that particular day how many number of hours you were on production how many breaks did you take you cannot take a break in oil and gas field it keeps on running throughout the day as is the case with other fields as well more water injection volume is the other column very understandable it it is the data from other well associated with this particular well that but that injector well is pushing water down the bottom hole down the rock it's almost like you are throwing water and flooding water into the rock so that that water once it goes in it pushes oil to the surface so that more water injection volume is how much water you are injecting okay how much injection hours you were doing on that day how much production arts you were doing on that day average downhill pressure that's important because it tells you if oil was flowing to the well with what pressure was it flowing to the well with what temperature was it flowing to the well right with what uh how much was your opening how much was your diameter of the of the tubular through which you were producing so all of these factors are very very relevant for an oil and gas engineer to model something to model something which which will be used for reserve estimation and economic analysis so in this case let me tell you right off the front in this case we will use all these other features as our inputs and we will try to forecast or predict the bore oil volume as you might have guessed so our why in this case is the the the thing that we want to produce is the bore oil volume predict is the more oil volume and the thing that we want to use as inputs is everything else and that can be optimized so basically we are assuming that there is a physical uh you know physical uh system that we are trying to model using machine learning uh all the other parameters they constitute they can be used as good predictors of the oil production volume that's our primary assumption the algorithm we have not yet thought about but we are assuming that all the things like the choke size the on stream injection hours how much injection rate is there how much is your pressure temperature all these things if you know you can put that thing as an input to a model and that model can give you the oil production rates right so that's our assumption now let's look at the data set first df.head helps you look at the first few rows of the data set right so you can see initially it feels like the entire data set is empty so you can use the pandas default plotting utility df dot plot provides subplots equals to true okay and it's coming up let me provide a fig size equals to 12 comma 4. this is the simplest way to have the first hand visualization of your data set twelve comma twelve let me increase the tallness of the figure as well okay now the first information is pretty clear the fact that initials were zeros but uh but the data set is not empty which is good now let's look at the important column which is the y column bohr oil volume okay where is the bohr oil volume this is the bohr oil volume the orange column okay like i said oil and gas production always follows a particular decline okay the sequence is very important very typical of time series data sets you cannot place the 2011 data in front of 2009 data because 2011 is later in the decline stage as compared to 2009 so the decline the decline when i say you can expect this to be an exponential kind of decline initially the oil production rates daily production rates were very high because initially it's almost like the coca-cola example when you open the bottle the flow initially is very very fast but slowly that flow decays down right so that's why similarly 2009 is when you open the coca-cola bottle and then slowly the the rates keep on kept on declining and finally at this stage this forecasting model can come off to be really good because now your authorities will be asking you some questions that hey can you tell me what is the forecast for this particular field if the forecast is good then we kept keep on producing from this field if the forecast is not good then we stop right here because we cannot waste any more money and remember guys 2014 to 2018-19 was a dark period for the oil and gas industry uh the economics went down terribly so it became very very important for everyone to realize whether it is uh economically viable or not right all right so these are the few few few columns that we have in our data set now if you if you observe that oil production kept on declining up to 2014 and all of a sudden if you see if i zoom in all of a sudden after 2014 it spiked a bit it spiked a bit even the gas that was produced along with oil that also showed a tiny spike now you need to have a physical understanding of why this might have happened if there is no explanation to that maybe machine will not be able to understand either so the explanation rise in the top curve the choke size you can see the choke size has increased around the same time which was shut in for a long time you can see the choke size which is the opening of the surface walls was zero the choke size was zero which means the choke was closed for a lot of long period of time and all of a sudden it increased the size increase as if someone opened the tap all of a sudden and that's why the spike came up so right off the bat we can understand that this particular feature called the choke size feature almost has a very very important relationship with the oil production so whenever we get to making a machine learning model or something this feature has to be there and similar analysis can be done with all the other features this is called exploratory data analysis you need to spend a lot of time with the data right now i have picked up picked up the best features that we have from the from the uh various sensors that's why we are seeing good good signatures that correlate with the output but normally you will have to spend let's say months and months to come to the good set of features which are representative which can be used as good inputs for the output so right now that's not the case okay now let's move on let's move on i have already told you about the output column but let's store that output column name in a variable output column equals to bohr oil volume okay now one of the very important things in in machine learning let's assume that we are going ahead let's assume let's assume that x is linearly related with y it's a good practice uh that people follow in in data sciences that it's not necessary to go for fancy models like a neural network or something if you are dealing with the data set for the first time it's always good to start at the bottom and then keep on improving and maybe bring in more complexities but today let's focus on assuming that it's a it's a linear regression problem and right off the bat when we assume that we need to make sure that the the multi-co-linearity problem which is a very very typical problem in linear regression is is not existent in at least in this data set and one of the ways to look at the inter feature correlation is uh to look at this thing a heat map of correlations we can use df.core if you use df.core it tells us how each of the two features are correlated with each other i know this is not the best correlation magnitude to look at but in the interest of time let's assume we go ahead with this for an idea okay and if we if we look at the heat map of this again not the best practice but if you look at the heat map of the correlations it will give us an idea that hey some features are intercorrelated with a huge magnitude so you can either remove one of them because why would you go with two features which are almost identical when you can just go ahead with one you want to avoid redundancy in your in your data set you want to remove all the duplicate features in your data set also you want to remove all those features which are very very highly correlated with the output as well it's almost as if we want to provide prevent any data leakage as well right so let's look at which of these features have a huge correlation uh with the output feature okay so to save some time i'm copying these commands so the what i'm doing here is i'm using this above data frame df dot core and i'm filtering out for the correlation magnitudes of each feature only with the output feature right now each features correlation with every other feature is being displayed but let's focus on each features correlation with the output feature so once we look at it we'll see that the the more oil volume and the more gas volume are almost 100 percent correlated with each other now the reason behind this is gas is an associated fluid which is produced along with oil because they are flowing out from the same outlet so if you are increasing the oil production the associated gas production will almost always be high as well so it's only a matter of a bit of physical calculations that you can use if you know the gas production you can know the oil production if you know the oil production you can know the gas production so you will not use gas production as an input because that's almost as if you are cheating in your machine learning model your machine has to serve some value your model has to serve some value and if you provide gas volume as an input to your machine learning model your machine learning model will almost give you a zero mean squared error almost there 99 accurate results etc which is not the practical thing okay so we will try to skip this particular feature we can set a threshold that hey if your uh relation with the output feature is more than 85 or 90 percent there is some leakage there is some correlation between the output and we don't want to use that so we will exclude those features right now let's just exclude the gas volume feature right okay so if in case you want to look at the magnitudes of of these correlations you can use that and you can create a data frame and you can see that gas volume is almost 100 correlated with oil volume and hence we will exclude this particular column again like i said uh a good practice that you can follow in other data science projects as well maybe there is some inherent correlation between the output and the given feature that was not conveyed to you by the domain understanding and hence you might want to confirm with the sme that hey are these two duplicate features are these two very very closely related with each other if yes then i need to skip one of these right so that's why we are skipping over gas volume okay now similar algorithm you have to use for input columns as well okay but right now we are seeing that the input columns are not very very extremely interrelated with each other so let's assume that all the other features are good so we go ahead with input columns with this particular condition okay we go with this particular condition that hey my input columns will be all those which are at least more than 0.2 percent correlated with my output and at least less than 0.9 percent correlated with my output i hope this is understandable we do not want those features which have absolutely no relationship with my output and we also do not want those features which are almost identical to my output we do not want to create a model which has output as the input we don't want any kind of data leakage right so this condition once we apply we will have input columns so the input columns that we have figured out okay uh we have these input columns we will need downhole temperature we will need downhole pressure uh delta p of tubing uh choke size like i said choke size very very important uh water injection volume because more you inject more you provide the pressure more the oil will be produced from the other well right uh on stream hours of course because the number of hours you will be producing for is directly proportional to the total production obtained right temperature pressure so basically all in all what we have concluded by now is whatever we did as a data scientist nicely resonates with what nicely resonates with the physical understanding of the domain and which is very very important you need to make sure you're not shooting in the dark you're not you're not shooting in the dark as in whatever you're doing must resonate with the with the field personnel as well you might have to tell the story to the field people which is again a very important thing your manager will ask you hey why did you pick up all these features uh you you can tell them two kind of evidences one the data driven evidence that sir this is more than twenty percent related and less than ninety percent related that's a data driven evidence and this is my domain-based evidence that of course this is the proportionality and so and so so we have concluded with the input features okay now let's create the metrics let's move towards ml till now whatever we did was a bit about feature engineering a bit about understanding what are the features which are important uh what are the algorithms very very basic stuff uh nothing nothing extraordinary now let's move to ml and then once we complete the entire thing we will look at what better we could have done right because of course this is the grounded model this is the this is the you know very simplest assumption kind of model and we might not strike gold on the first attempt right so x and y we have the first step is to define what is your x i mean what is your input and what is your output now no matter what algorithm you use you are trying to create a a simulation of this particular thing y will be a function of x okay this is what machine learning is right you are trying to create y as a function of x there can be some error as well but this is the physical system you are trying to model now this f depending on what algorithm you are using right now we are using a linear regression algorithm so we are assuming that f is a linear uh model which is mx plus c kind of thing but if you are using a neural network we are assuming that f is a very very complex kind of model so overall philosophy remains the same but the the f keeps on changing that's machine learning right now the first thing you have to do like every machine learning model you have to keep one portion of your data untouched you have to keep one portion of the data untouched as in you will do all the modeling stuff on some part of the data and the remainder part of the data you will use for validation okay so you have now because this is a time series data we cannot do the random shuffling thing that we do in other machine learning algorithms yes we are looking at regression right now but i will transform the thinking to a forecasting model in just a bit but always remember the ideal practice in any time series modeling algorithm is to never use shuffling because shuffling is almost equivalent to data leakage and time series problems so if your total data set size let's look at the total data set size what was the total data set size we can look at that using df.info okay 3291 you know rows or days of data so we can use 3000 as our training 3000 days as our training data and 291 days as our testing data right uh we will not look at the later 291 days uh and we will validate it further on okay now one more thing that we have to be very very cautious about is whether our data currently has outliers or not whether it has nulls or not right now this is an ideal data set no null values you can see everywhere it's non null but if in case there were some null values you would have to interpolate for example if i pick up one column let's say oil production volume uh more oil volume okay more oil volume and if i plot it dot plot fig size equals to 12 comma four okay now if i do something like this okay if i if i use the same column more oil volume and i replace a few normally most data most data sets already have the null problems right now because this is an idealized data set we are not seeing nulls but let's let's replace some let's say uh from 2000 up to uh 2200 let's replace them with or let's say 2500 let's replace them with np.nan okay not the best way to do that of course but let's now visualize it did it happen so you can see this this uh this let's assume that our data originally was something like this we had a not a lot of null values lying around this 2013 mark now what would you do in this case there are a few options that you can use normally if you are using a non-time series example some people go ahead with trying to use the average to fill in for the missing values some people uh follow some other tricks and technologies but right now the best way you can use is physics right you can use this kind of decline you can model using curve fitting or whatever you can you can model the decline in this area because i see that this is an exponential decline why not fit a exponential decline model using scipy or numpy or whatever just random curve fitting that you can do in excel as well and use the predictions by that curve fitting to fill that to fill this particular zone so you can use that that would be much much better but the in the interest of time let's use something like dot interpolate and you can see it's okay it's not the best i mean it's bad way to do it but this is something that you can do as well okay so let's let's replace the df output by uh the next thing you have to do once you interpolate you have to replace the original more oil volume by this one okay more oil volume interpolated so this is just an introduction to how you would handle nulls of course every data column would have a different range using which you can you can do follow various other ways to fill for the nan's but right now the output column can be filled using this it's always good to use physics wherever you can okay never keep yourself away from the domain okay so we will have this and to check for outliers etc and what kind of what are the columns that we need to check for what are the columns which have a lot of outliers you can use df.describe a lot of people use they at least type this command in their notebooks but what is the best way to use the information given by descriptive statistics i'll tell you one thing that i use df.describe for so for example i know the 75th percentile of every column and i know the 25th percentile of every column follow this thumb rule that if a particular column your minimum is very very far away from 25 percentile or your maximum is very very far away from the 75th percentile then you need to focus on that column for outlier removal i will repeat pick up a column observe if your minimum is very very far away from the 25th percentile or your maximum is very very far away from the 75th percentile if that's the case you might have to check for outliers in that particular column for example average downhill pressure the 75th percentile is 235. uh max is 317 looks good but this one the 75th percentile is 6851 max is looks bad 25th percentile is 3972 minimum is zero looks terrible so this is definitely a water injection volume is definitely a column that you have to handle for outliers okay but right now in the interest of time we are not doing that but the way we what we would do is you would at either remove these outliers or maybe replace them with more representative values uh in the in the interquartile range okay let's move forward with our modeling algorithm now one interesting problem with this data originally you can see is the range of every column is very very different uh the injection hours in is in 0 to 25 this is in thousands there is a column for gas which we have removed right now but it has millions almost millions so that's that's a problem if you feed this data directly to the machine learning model because of the bias and the magnitude the machine will treat this as if uh the high magnitude features are very very important as compared to the low magnitude features so it's always good practice to you know scale it and and if you have attended some courses you must have realized that you must have realized that uh the feature scaling uh makes your machine learning models convergence very very fast so it's always good to you know use some sort of scaling algorithm to scale your data data to a particular range and you would never scale uh a fit a scalar on the entire data set because that would lead to some sort of because the scalar algorithm what it does is scaling what it is let's assume that you are scaling by just by dividing by the max so you have to calculate the max right so if you calculate the max on the entire data set there's data leakage because how can you know the max of unseen real world data maybe the max that you calculated here and the data that you found out in reality was way different so you cannot use a scalar scaling algorithm on the entire data set you have to use a scalar on the training data only that's why we did the train test split first and now we will do the scaling okay so let's do the scaling so you can see i'm using standard scaler here you can use min max scaler as well but stand anyway it's fine it's all the idea is to chunk down every feature to a common range min max scalar will bring everything to a zero to one range or minus one to one range but standard scalar it brings everyone everything to a standard normal distribution which is sent centered around zero with a standard deviation of one that's a fancy way of telling that everything comes to a narrow margin okay so you can see i'm i'm fitting on the training data and i'm transforming on the test data yes you can transform what you fitted on the trained data but you never fit on the test data okay all good now let's train a machine learning model like i said we will start with a linear regression algorithm okay we'll start with the linear regression algorithm so from escalon dot linear model you are importing linear regression first you have to tell python that hey i want to use all the powers of the linear regression class so python arranges for that so this is class model instantiation so when you do that you are creating an object of the linear regression type now this object will fit on the data right now this is a very very generic model right now this model is not data specific this is just tell us telling python that hey we are going to use a linear regression way of modeling we still don't know what the data looks like so the next step that we will do is to fit the training fit this particular generic model to the data to make it more specific to adapt this algorithm which was initially very generic to adapt it to the particular oil and gas data so when we do that there is some issue let's see what issue input contains nan yes because uh we did some uh nan creation here right so let's rerun that x comma y part again where is it where is it okay just give me one second i'm finding out where yup okay dot fill and a let's fill for zero we have looked at better ways to do that but let's just to prevent that error let's fill with zero reminder this is not something that you would ever do so don't fill with zero okay but this is just to complete our entire project now if i run it again extreme y train has to be run again again again again now if i run it you see now it's running okay now you see that you have fitting on the training data x scale version of x because otherwise your model would have performed badly and your model would have trained slowly that's the two disadvantages if you do not uh if you do not scale your data set and you you would have lost the surety whether your model would converge to a good solution let's predict our y predicted okay lm dot predict now whenever one thing that i want everyone to know is given this is a linear regression algorithm whenever you fit a linear model on the training data i told you that this makes the model data specific which means that the coefficients and the intercepts are now found out for this particular data if you run the same step for some other data the slope and the intercept will keep on changing for the particular data that's why machine learning model mathematically it's a very generic thing but when you fit on it when you fit on a particular data that's when it adapts to that particular data right now we will plot out our predictions along with the actual data okay so we can see on the x-axis i'm taking the time stamps on the y-axis i'm taking the predictions right now we are uh plotting the actual training data and the predicted training data you can see it's fine it's still fine very very close the blue lines are very close to the actual orange line orange dots but let's do the same thing for testing data let's do the same thing for testing data and now when we plot it now when we plot it you see the testing data is way apart way of the green dots the black black is our prediction and the green dots are the original oil rates and you can see the predicted values are way off and you can also see that during some time stamps the model is predicting negative oil rates which is laughable in fact it's not something that can ever happen so that's that's that's like that's a slap in our face that your model is not performing well but we know it would we already knew it would not perform well we did not handle multi-collinearity we did not handle non-linearity as well maybe what if some of these features were not linearly related for example what if if we look at df columns what if in x we needed the square of the annulus pressure what if we needed the pressure multiplied by temperature as a feature so it's not always the case that what you are assuming is what the physical system will follow and from experience i can definitely tell that oil and gas subsurface systems and the physical system that you are trying to model right now it's highly non-linear and that's why you are you are breaking the assumptions of physics that's why your model is not performing well even if the test performance was good your model would still fail in the reality so you you might have to bring in some intersection with physics maybe some domain based corrections on the data we did not even check the magnitudes of all the data sets all the features maybe some of the features were in negatives where they were not allowed so you have to be in very very close touch with with the domain personnel in order to make sure what you are doing is correct okay now if i if i use the same thing over a random forest at least because right now one data driven problem that we are facing is we are breaking the assumptions of linear regression okay we are breaking the assumptions of linear regression so why not if we do not know how to correct for that why not use something like a random forest which we do not have time to understand but just believe me that it is a model which generalizes well it has a neck of understanding the entire crux better if you feed data to a random forest what random forest does is it takes in the data it it creates multiple mini models and the final prediction is an average of all those mini models right now we have the prediction from only one model which can be bad but random forest is it's almost like you are taking the advice of all the family members some of them will be supporting you some of them will be against you but overall whatever decision you make will be a combination of all of them so that's the benefit of random forest in machine learning terms we call a model which generalizes better so let's do that quickly and let's see whether the predictions are better in a random forest okay this might possibly be the first machine learning webinar you are attending where i am presenting towards you a model which is terribly failing okay so if i plot now if i plot the entire thing now you can see we already came close again warning no no guarantee that this will perform well in reality but like i said this has understood the crux better maybe some limitations were there some magnitude one benefit is we did not even have to do a lot of feature processing just like we had to do in linear regression we did not even use scaled version of the data so random forest definitely is doing better in this case but again you still have to be very very close uh in in with the physics people people who are working in the field you have to check that sir i think in the in the 2009 period there were a lot of production droppings can you tell me why that was happening if that was an outlier maybe i'll have to remove it so we have only looked at 10 percentage of the story but assuming if this was a correct thing what we would do next would be to use this black forecast right now this is not a forecast but if we modify uh if we introduce some lags in our inputs we train some some forecasting algorithm i have i will share a link with which talks about forecasting as if it was linear regression uh which is a better and easier way to understand forecasting but if we if we are forecasting that this black line is our forecast for the next few months the area under this curve what is the unit of this black line it is in barrels per day and what is the what is the unit of x-axis days so if you take the area under this curve this is barrels so it tells you that after this red line how many barrels are remaining from this point onwards that's called a reserve if you multiply the barrels which are remaining in the oil field and if you multiply that with the oil prices in in the current scenario which is let's say 92 brent for brent it is 92 barrels uh 92 dollars per barrel if you multiply them that tells you that hey in the future in the next six months uh this and this uh million dollars you can expect as the cash flow and then you will do your economic analysis that hey i am going to forecast my investments over the next five months and that tells me that my investments are going to be way off the charts as compared to the returns i'll be receiving that tells me that hey i need to stop right now and you would not use this model standalone you would use the other physics based simulations and maybe use a hybrid prediction for the forecast maybe a 60 40 weighted average or something like that but this is why we are doing that okay and i hope this was a good introduction which was a good combination at least of of what is happening in the oil industry uh and what uh what how how well it is connected with what machine learning workflows follow so now i will open the floor for uh maybe some discussions uh and maybe look at the flooded chat box that i see uh hi there franco so uh there are questions in the q a section uh we have told the people to put their questions in the q a section you can start from there if you want i'll read out the questions and you can answer no i uh yeah i can do that so i'm looking at questions one by one so rajesh has asked us that is the choke same as walls of course yes absolutely so it's almost you can assume it to be very similar to a tap or a wall so a choke size if it is more it means you have opened the tap more or the valve you have opened it more and if you if you see the choke size to be zero uh the tap is closed okay uh data set name it's the wall field data let me paste it out in the chat for you i will paste that uh i will i will share the relevant files uh uh in the chat after i answer the questions don't worry about that okay let me share the actual project with you in the chat box [Music] okay i'm sending let me open the chat box first okay pasting it here okay uh i'm sorry to interrupt you just to inform everyone so i've started a feedback poll uh please do answer that as well it will help us to you know uh bring on some more sessions like this and yes you can continue to run so thank you okay thanks yeah so someone asked me about the oil volume and gas volume looking very very similar yes we dealt with that in the session like i told you they are coming from the same wall so gas is associated with the oil and hence they are almost identical they are the fluids produced off of the same wall that's why they were looking very very identical so that's there uh for the data set again i'm sharing the data in the chat data is copied in the chat so you can directly import the data uh in the chat right uh so uh when you show the above code what kind of projects you deal with has asked a very interesting question can't you step in to make understanding with data without having the domain knowledge i mean of course you can do a lot of things with data but sometimes to improve whether the data itself is correct you might need domain domain understanding i mean most of the data science teams that are created at least in the industry they definitely have one or two domain smes they they keep on attending the the regular calls you ask them whether this feature is important whether these values are expected whether they are representative of the physical system so it's very important for your model to perform well so either you you do the domain reading on the background or you connect yourself or you involve people from the domain in your in your project team either works okay so yes i have shared the notebook uh i mean rajesh the threshold value of correlation features it depends from a use case to a use case also depends on the uh on the company policies and the sme discussions that you have but normally feature should not be too interco related and if you if you are a statistician you might want to look at whether they are statistically significantly correlated with each other i recently shared a mathematical background about linear regression on my linkedin so please go and check that out that explains how to handle multi-co-linearity in a way better way than i could discuss today so there are better ways to do that right so beginning we can use ml instead of end of the oil reservoir ml can be used to have a guidance as far as i know i mean this will not replace any any pre-existing workflow but this will definitely be a good suggestion mark that hey uh we can use uh this suggestion i mean in fact this was well appreciated by people from the oil industry because uh the simulation outputs they take a lot of time to to generate themselves because that is dependent of partial differential equations and to solve them it takes a lot of time and this is a replica of that but it solves way faster due to computational power okay so kashish has asked whether this time series is stationary of course i could not get a chance to look at that in this session time series forecasting can be for some other session but of course yeah that's a good point to check if you are looking at it from the forecasting angle you need to make sure stationarity is there chemical engineers they definitely are very important for petroleum industry a lot of chemical engineers are working as great data scientists in various oil and gas companies in fact in my last company almost every day i would interact with people from the chemical engineering domain so they came right right from the field working as the consultants in the data science industry so yeah i mean it's always okay to ask questions about the data like i did today like like when i showed you this plot i asked you why this is happening then i look for for some clues in the data itself either ask for clues from the data the data will confess to things or you ask ask for clues from your from your seniors or people in the in the team okay all right so i think that's uh to select the features again i would like you to refer to a linkedin post i shared i will uh you can you can find that uh to check that on my linkedin i shared about linear regression there are i mean random forest if you are using random forest algorithm random forest has a way i will show you it tells you the feature importances so i will rf dot rf dot feature importances you can see directly you can obtain the importance of all the features so if some features let's say this feature and this feature is very very less important random forest is suggesting that i do not need these features uh you can use that but right now we cannot decide purely based on this okay so various algorithms have different ways to select features okay yeah forecasting article i'm sharing the link it's my attempt to you know explain forecasting as if it was linear regression but again can find that out in linkedin as well okay i've shared it uh splitting the data yes uh uh i i like the answer to that by andrew ng uh he he tells us that uh if your data size is huge and if you can afford to you know have one percent as the test size that also works so 80 20 is not the standard benchmark if you have million samples in your data even one percent would mean tens of thousands of rows right so that's enough for testing set so it depends on the use case to use case right so no no no okay i think that's all about it for and i would definitely paste my linkedin as well for whoever has more questions that they could not ask today i will i will share my linkedin in the chat please feel free to come up with questions i think we are almost there at the end of this session yes here we are we have surely shared your linkedin profile link as well in the chat about it you can also go ahead and share it perfect perfect perfect all right thanks thank you so much adrian show thanks a lot on behalf of analytics video as well this has been an amazing session and we hope to have another session with you as well like even the people or the attendees today would love that so thanks a lot it was really informative for me as well thank you thank you everyone thanks thanks for the positive feedback see you soon bye you guys bye bye good night

Original Description

The DataHour: Energy Data Science - Project From Scratch In this DataHour, you’ll get a good sneak peek into what kind of Data, Processing & Projects happen in the Oil & Energy sector. You’ll gain an understanding of Raw Data from an Oil Producing field i.e the Oil Well Data as well as the Water Injector well data. The data processing steps will be the major focus, which is the heart of many Industrial Data Science workflows. We'll talk about Null Handling, Feature Engineering, Outlier Detection, Descriptive Statistics etc. At the end, introductory Machine Learning will be explained to conclude the session. 🔗 More action pack session here: https://datahack.analyticsvidhya.com/contest/all/ Stay on top of your industry by interacting with us on our social channels: Follow us on Instagram: https://www.instagram.com/analytics_vidhya/ Like us on Facebook: https://www.facebook.com/AnalyticsVidhya/ Follow us on Twitter: https://twitter.com/AnalyticsVidhya Follow us on LinkedIn:https://www.linkedin.com/company/analytics-vidhya
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Analytics Vidhya · Analytics Vidhya · 3 of 60

1 The DataHour: Data Science in Retail
The DataHour: Data Science in Retail
Analytics Vidhya
2 The DataHour: Anomaly detection using NLP and Predictive Modeling
The DataHour: Anomaly detection using NLP and Predictive Modeling
Analytics Vidhya
The DataHour: Energy Data Science Project from Scratch
The DataHour: Energy Data Science Project from Scratch
Analytics Vidhya
4 The DataHour: Explainable AI Need and Implementation
The DataHour: Explainable AI Need and Implementation
Analytics Vidhya
5 The DataHour: Google Cloud AI/ML
The DataHour: Google Cloud AI/ML
Analytics Vidhya
6 Prediction to Production in Machine Learning #machinelearning #prediction
Prediction to Production in Machine Learning #machinelearning #prediction
Analytics Vidhya
7 Practical Applications of Data science in Ecommerce
Practical Applications of Data science in Ecommerce
Analytics Vidhya
8 How to tackle Overfitting?#machinelearning #overfitting
How to tackle Overfitting?#machinelearning #overfitting
Analytics Vidhya
9 Building Data Pipelines on GCP #googlecloud #datapipelines #data
Building Data Pipelines on GCP #googlecloud #datapipelines #data
Analytics Vidhya
10 Hands-on with A/B Testing #abtesting #datascience
Hands-on with A/B Testing #abtesting #datascience
Analytics Vidhya
11 Efficient Implementations of Transformers #transformers #cnn  #machinelearning
Efficient Implementations of Transformers #transformers #cnn #machinelearning
Analytics Vidhya
12 Modern Deep Learning Architecture #deeplearning  #architecture #deeplearningtutorial
Modern Deep Learning Architecture #deeplearning #architecture #deeplearningtutorial
Analytics Vidhya
13 Key steps for Designing Artificial Neural Network (ANN) for Image classification #machinelearning
Key steps for Designing Artificial Neural Network (ANN) for Image classification #machinelearning
Analytics Vidhya
14 5 things you should know about Azure SQL #azure #sql #datahour #datascience
5 things you should know about Azure SQL #azure #sql #datahour #datascience
Analytics Vidhya
15 AI & ML in the Automotive Industry #machinelearning #ai
AI & ML in the Automotive Industry #machinelearning #ai
Analytics Vidhya
16 Building Machine Learning Models in BigQuery
Building Machine Learning Models in BigQuery
Analytics Vidhya
17 NLP aspects in Telecommunication Industry
NLP aspects in Telecommunication Industry
Analytics Vidhya
18 Practical Time Series Analysis
Practical Time Series Analysis
Analytics Vidhya
19 Fundamentals of Quantum Computing
Fundamentals of Quantum Computing
Analytics Vidhya
20 A DAY IN THE LIFE of a Data Scientist (From waking up to working on algorithms)
A DAY IN THE LIFE of a Data Scientist (From waking up to working on algorithms)
Analytics Vidhya
21 Classification Machine Learning Model from Scratch
Classification Machine Learning Model from Scratch
Analytics Vidhya
22 Knowledge Graph Solutions using Neo4j
Knowledge Graph Solutions using Neo4j
Analytics Vidhya
23 Model Guesstimation (MLOps)
Model Guesstimation (MLOps)
Analytics Vidhya
24 ETL Pipelines in Google Cloud Platform
ETL Pipelines in Google Cloud Platform
Analytics Vidhya
25 Key steps for Designing Convolutional Neural Network(CNN) for Image Classification
Key steps for Designing Convolutional Neural Network(CNN) for Image Classification
Analytics Vidhya
26 Getting Started with AWS EC2 #amazon #aws
Getting Started with AWS EC2 #amazon #aws
Analytics Vidhya
27 How to Use Azure NLP and Graph Databases for Intelligent Knowledge Mining
How to Use Azure NLP and Graph Databases for Intelligent Knowledge Mining
Analytics Vidhya
28 Certified AI & ML BlackBelt Plus Program #shorts
Certified AI & ML BlackBelt Plus Program #shorts
Analytics Vidhya
29 Visualizing Data using Python #machinelearning #visualization #python
Visualizing Data using Python #machinelearning #visualization #python
Analytics Vidhya
30 DCNN for Machine RUL Prediction using Time-series Data #timeseries #machinelearning #datascience
DCNN for Machine RUL Prediction using Time-series Data #timeseries #machinelearning #datascience
Analytics Vidhya
31 M in ML stands for Math & Magic
M in ML stands for Math & Magic
Analytics Vidhya
32 An Unsupervised ML approach using Clustering
An Unsupervised ML approach using Clustering
Analytics Vidhya
33 Customizing Large Language Models GPT3 for Real-life Use Cases #gpt3 #datascience
Customizing Large Language Models GPT3 for Real-life Use Cases #gpt3 #datascience
Analytics Vidhya
34 Model Parameters vs Hyperparameters - Techniques in ML Engineering #machinelearning
Model Parameters vs Hyperparameters - Techniques in ML Engineering #machinelearning
Analytics Vidhya
35 Practical MLOps #mlops #datascience
Practical MLOps #mlops #datascience
Analytics Vidhya
36 Data Engineering with Databricks #dataengineering #databricks
Data Engineering with Databricks #dataengineering #databricks
Analytics Vidhya
37 Multi-Objective Optimisation
Multi-Objective Optimisation
Analytics Vidhya
38 When Airflow Meets Kubernetes
When Airflow Meets Kubernetes
Analytics Vidhya
39 AI in Banking
AI in Banking
Analytics Vidhya
40 Learn Convolutional Neural Network for Image Recognition
Learn Convolutional Neural Network for Image Recognition
Analytics Vidhya
41 Extracting Value from Data
Extracting Value from Data
Analytics Vidhya
42 How to measure Marketing Channel Effectiveness
How to measure Marketing Channel Effectiveness
Analytics Vidhya
43 Transforming Lives | Data Science Immersive Bootcamp
Transforming Lives | Data Science Immersive Bootcamp
Analytics Vidhya
44 Stock Market Analysis - AI driven approach
Stock Market Analysis - AI driven approach
Analytics Vidhya
45 Become a Data Engineering Professional in 2022 | Future Trends + Skills Required
Become a Data Engineering Professional in 2022 | Future Trends + Skills Required
Analytics Vidhya
46 Ensemble Techniques in Machine Learning #machinelearning #ensemble #datascience
Ensemble Techniques in Machine Learning #machinelearning #ensemble #datascience
Analytics Vidhya
47 The Power of Visualization | Tableau Full Course | Analytics Vidhya
The Power of Visualization | Tableau Full Course | Analytics Vidhya
Analytics Vidhya
48 Demand for Data Engineers is on the Rise | Data Engineer | Analytics Vidhya
Demand for Data Engineers is on the Rise | Data Engineer | Analytics Vidhya
Analytics Vidhya
49 Data Visualization in Data Science | DataHour | Analytics Vidhya
Data Visualization in Data Science | DataHour | Analytics Vidhya
Analytics Vidhya
50 Role of Optimization in Machine Learning & Deep Learning | DataHour | Analytics Vidhya
Role of Optimization in Machine Learning & Deep Learning | DataHour | Analytics Vidhya
Analytics Vidhya
51 Solving any Machine Learning Problem | Approach and Steps Involved
Solving any Machine Learning Problem | Approach and Steps Involved
Analytics Vidhya
52 Topic Modeling Explained with Implementation | Using LDA in Python | DataHour by Arpendu Ganguly
Topic Modeling Explained with Implementation | Using LDA in Python | DataHour by Arpendu Ganguly
Analytics Vidhya
53 Data Engineering in E-Commerce | The Best Case Study
Data Engineering in E-Commerce | The Best Case Study
Analytics Vidhya
54 Introduction to Classification using Azure Machine Learning | DataHour | Analytics Vidhya
Introduction to Classification using Azure Machine Learning | DataHour | Analytics Vidhya
Analytics Vidhya
55 Introduction to Federated Learning | DataHour | Analytics Vidhya
Introduction to Federated Learning | DataHour | Analytics Vidhya
Analytics Vidhya
56 Diffusion Models for Generative Arts | DataHour | Analytics Vidhya
Diffusion Models for Generative Arts | DataHour | Analytics Vidhya
Analytics Vidhya
57 Master Google Analytics in 1 Hour | DataHour | Analytics Vidhya
Master Google Analytics in 1 Hour | DataHour | Analytics Vidhya
Analytics Vidhya
58 Learn Hypothesis Testing | DataHour | Analytics Vidhya
Learn Hypothesis Testing | DataHour | Analytics Vidhya
Analytics Vidhya
59 A Practical Approach to Kaggle Competition | DataHour | Analytics Vidhya
A Practical Approach to Kaggle Competition | DataHour | Analytics Vidhya
Analytics Vidhya
60 Making AI work for Business | DataHour | Analytics Vidhya
Making AI work for Business | DataHour | Analytics Vidhya
Analytics Vidhya

This video provides a beginner's guide to energy data science, covering the basics of data science and machine learning in the oil and energy sector. It demonstrates how to process and analyze oil well data and water injection data using Python and relevant libraries. By following this video, viewers can gain hands-on experience with data science project development from scratch.

Key Takeaways
  1. Import necessary libraries like Pandas and scikit-learn
  2. Load and preprocess oil well data and water injection data
  3. Apply data visualization techniques to understand the data
  4. Develop and train supervised learning models
  5. Evaluate model performance and refine the pipeline
💡 The oil and energy sector can benefit greatly from data science and machine learning, particularly in optimizing oil well data analysis and water injection data processing.

Related Reads

📰
Opinion: An AI Feature Should Pass a Model Swap Test Before It Touches Production
Ensure AI features pass a model swap test before production to prevent silent regressions
Dev.to AI
📰
How I built a self-correcting pricing tracker with Apify and Notion
Build a self-correcting pricing tracker using Apify and Notion to automate price monitoring and update your database accordingly
Medium · Machine Learning
📰
My Model Said Germany Had (Almost) Zero Income Inequality
Learn how to merge multiple government inequality datasets into one MLOps pipeline using Python and identify potential discrepancies in the data
Medium · Python
📰
Laporan Praktikum 1: Membuat Program Identitas Sederhana Menggunakan Python
Create a simple identity program using Python and learn the basics of programming
Medium · Programming
Up next
Generative vs Discriminative Models - Explained
DataMListic
Watch →