Solving Real-World Problems with Data | Dr Vikas Agrawal, Sr Principal Data Scientist@Oracle | LWD03
Key Takeaways
Dr. Vikas Agrawal, Sr Principal Data Scientist at Oracle, discusses the importance of understanding the problem before solving it, and how design thinking and technology can be used to solve real-world problems with data. He also talks about the limitations and capabilities of Generative AI, and the need for human oversight and judgment to ensure trustworthiness of AI outputs.
Full Transcript
what uh you know is not so commonly known to the customers often is that the tools and techniques and the state of the art in many fields has not reached a point where you can always get correct answers for example [Music] [Applause] [Music] hi vikas and welcome to the show uh leading with data and thanks for taking time out from your schedule for doing this discussion uh privilege yeah for our audience uh vikas has been extremely supportive of analytics with there and any Community initiative which we have done over the last few years he is currently a research scientist at Oracle and doing some great work and uh he's someone I find uh you know who really connects both the pieces the technology as well as the business aspect and the impact of the work very very closely so so I really look up to him to understand the intersection of these things so I hope that through this discussion today because we all get to learn that from you so so welcome to the show because it's been you know six years of our association it's been so much fun yeah so I'll actually you know uh start the discussion from uh the point which I mentioned right so I've always found you to have great technical depth at the same time you know uh having that macro view as well that how some of the solutions which are getting built not only impacts the business problem but also the overall let's say ecosystem so uh uh so how do you uh do this or how how have you been able to you know uh maintain that consistently because there is so much to learn technically in the domain there is also so much to you know kind of reflect on and research at the same time and that's something which I you know always find uh really great with you so so how do you do that uh uh in your day-to-day work well so one is I have to credit my mentors uh you know at IIT Delhi at Caltech at Intel Corporation and in Infosys and here at Oracle uh because the the the kind of idea that I'll use technology to do something simply doesn't resonate with my mentors what resonated with my mentors was we have a problem to solve we'll do whatever it takes to solve the problem so now as a part of that the first step that I was taught and over a period learned is that you have to first understand the problem you're trying to solve very well in fact we spend 90 percent of the time understanding the problem and the remaining in person solving it so that is the critical part of uh you know what I do even now is if if we if a customer brings a problem or one of our product managers brings a problem we spend most of the time actually understanding the problem whether it is how others have solved it before or whether pieces of it have been addressed before or what do we want at the end what does the customer actually want at the end no so and you know this also goes back to the whole idea of design thinking yeah I learned from our friend Vishal sikka that if if you are starting from technology you're not getting very far right we stay at the Forefront of it that's fair but what are we using it for that's our first order many times sometimes customers say I have a lot of data what you know what can you do with it right wow we really love you as our customers because that's unlimited money for us but then I say but no that's not useful for you let us figure out what do you want correct because I can do a lot of video data yeah yeah please go ahead uh so see day to day the way it manifests is really in terms of you know for for example if you want human Capital Management problem solved right I have textbooks on that sitting here yeah if if I'm going to solve an accounting problem I have my ncert textbook sitting there 11 12 ncert District I skipped those in 11th and 12th but I'm studying them now now if I'm solving a manufacturing problem I go to that I have had colleagues you know uh at a previous company I worked at they were working with a Pharma company that product manager he went ahead and read a biochemistry textbook in 15 days wow so you know that's how you learn from your colleagues around you that you know the end of that determination that you know the subject as well as the meaning of course you in 15 days you cannot learn the whole subject but the idea is please know what they're talking about so part of it is communication sorry I interrupted you yeah no I was actually trying to get uh deeper in this so so while you're you know let's just spending this time with the customer do you also start structuring uh let's see your thoughts and your framework to solve the problem and also maybe start mapping some of these uh tools and techniques or you first I mean if you have to break it down one step further yeah a lot more yeah there's a lot more breakdown there right so so one of the things we do is once we identify that there is a problem worth solving do we even have data to solve that problem or how do we get the data to solve the problem that's number one yeah right if we don't have data to solve the problem then we are not going very very far yeah right number two we say do we have the technology to solve the problem today if I'm going to have a technology 10 years later not useful for my customer today so now if I have a there in one or two years I may still take it up I may say okay let's do a POC with what we have and we create new intellectual property along the way new algorithm right new tools new techniques uh but if I don't see a path in five years I don't take it up okay as a academic Community as a industry research Community we don't have a clear path then you know at least in the industry environment I don't take it out in the Academy of course right so now then we say okay let's do a honest to Earth POC which will have serious amount of data you know depending on the size of the problem it could be terabytes it could be petabytes it depends depends on this what problem you're solving right and then we say okay let's put the data in one place which is accessible to our team now what we're going to do is we're trying to look at the data see what are the outputs expected and spend maybe three months coming up with something that properly works and I'm not talking about merely getting a UI right or merely getting you know a notebook working I'm talking about end-to-end data pipeline all of that stuff okay okay now it may not scale to at this stage it may not scale right but it will solve and you know I do not believe in those slice version of MVPs that I'll show you because see slice I can do I can get a Graduation to do the slice for me in the industry I have to have the whole cake slice will not work exactly I just don't I'll give you the icing on the cake in the POC that I have a clear path to what algorithms I'm using what are the data sources I'm using and what is the nature of the output yeah and if I am perhaps 70 of the way there in terms of the mutual nuclear expectations then I'll say okay let's do a proper pilot you know put it on your Cloud you know and run you know if you you know maybe are people testing it you're people testing it and usually we expect the test the testing folks to be not somebody you pick off of Mechanical Turk right I mean it would not be like usually not the end end users people who are experts in the area who know if I am giving the wrong output they know I'm giving a wrong output which may not always be possible for uh end-end user they may simply go with the output so that's a very crucial step I mean nowadays they call it you know the PV we have fancy names for it you know reinforcement learning with human feedback the idea has been around for all along right because changing your rates back there in some form or the other we I mean that's absolutely required correct then comes the optimization phase see so far I have not optimized my model so far I have not worried about whether it will work in the next country whether it will work in the next year whether it will work uh you know scalably because I may have to retrain the model I believe every week will it finish in eight hours the training I am not worried about all of those things I've just done a pilot yeah right correct so now I said 90 of the work is in the research now I still have 90 left in the ml Ops now because now I have to make sure that my model um when it you know if it's running in India you know for the data in India it'll operate one way it goes to Norway the business process characteristics change even within the same company correct yeah their Factory characteristics will change now I know one company where this doesn't happen in for example Intel corporate operate they they follow a internal policy of copy exactly so their offices are all the same everywhere in the world their factories are all the same the characteristics are all the same and so on but most companies don't do that you know they customize things according to geography and so on so especially you know retail world or finance and so on you will see happen very often so we have to create the models in such a way that the model automatically adopts adapts itself yeah yeah and the model should also know when it is going out out of your distribution so correct itself so very interesting so so there are multiple steps here yeah yeah yeah so but if I have to just uh put it at a slightly simplistic level you are seeing in POC you would want to get to a level where the algorithm so you are essentially proving the concept without worrying about scalability and ml Ops so you're you're getting that yeah a little bit right but okay yeah I'll still worry about a little bit because see I can show you a very good POC right but if it doesn't fit your budget for example I'll give you a simple example right or something you know even we are doing now right now see I can do certain things with old school natural language processing which will basically fulfill your requirement it will not cost you GPU time on the cloud GPU time on the cloud is very expensive yeah right so uh now I can choose to create a solution for you which uses GPU and state-of-the-art llm state-of-the-art computer vision systems that will cost you about every inference will cost you a bomb right or I can get you something which is quite fast and its compute you know compute cheap okay that so yeah I think that trade-off is very important so there are customers that we have that will not care about the money and I haven't seen those yet so so we have to you know we have to worry about the scalability I got it got it interesting and uh typically uh you know you if you mentioned that uh 10 of the uh the time later on is all about optimization and ml Ops now in this journey right I said 90 percent ninety percent there and then uh 10 percent in the uh POC stage then yeah in the POC says 10 percent yeah I agree right okay got uh in this cycle when do you see the highest risk of a project failing or what is the most critical step to get right so if you have to for example uh be extra cautious at some of these steps to ensure that the project actually uh has higher chances of delivery where do you kind of spend more time on or or the other way to look at it is a mistake in which of these steps is the most costly usually I think the the the the the mistakes that are most costly are around AI hype as in where we communicate something and the customer understands it to be something else whatever this conference so we need to make that make it make sure that the expectations are mutual and clear okay okay so so that is where what usually happens is you come up with a POC and you think you've done the right thing yes and what happens is the customer is usually this this wasn't the case let's say 10 years ago 20 years ago right because people are preserving AI for a while um today the expectations tend to be very very high yeah okay and what uh you know is not so commonly known to the customers often is that the tools and techniques and the state of the art in many fields has not reached a point where you can always get correct answers for example perfect that is so true that is true that expectation setting the expectation and of course the problem definition which is the very first step right if you don't get the problem definition right you're definitely going down the drain guaranteed meaning you're losing your customers money because they'll Place faith in you and you know because you know they they expect a certain professional uh quality and they'll place the place then what you find out is that if you define the problem in such a way that either it is not directly addressing the customer's problem or you define it in such a way that it is trying to boil the ocean right both and do you also see scenarios do you also see scenarios where customer had has this latent problem in the sense they are saying something else right right yeah yeah the fourth now I'll give you an example of that one that latent story so uh there was a customer from a very large Pharma company you know very large and you know we probably used the products every day perhaps now that the challenge that uh they threw at us or we we thought they threw at us is I want to repurpose my drugs so they have their drugs and what they want to do is take the side effects of the drug as the main effect and we know some drugs which okay do that right something was developed for a cardiovascular purpose and then it's reused for something else right so now the challenge is that um that we don't even they don't understand it we don't understand it how that's done and somehow they have a conception that AI can do it somebody told them at a conference that it can be done right so okay so then I asked to ask them the question that um do we know the entire metabolism the pharmacokinetics of the drug do we know the side effects how far they go and they said well we haven't done all the studies right meaning the science is not even there so if the science is not even there what will AI do for you right so and then sometimes they are trying to invent signs yeah yeah yeah but or see what they're thinking well I'm going to look through all the published work in the world and somehow use some kind of natural language processing large language models create a large language model all the papers and kind of make it happen the challenge is we're not there yet as a you know human kind we're not quite there truth so so that's what we have to say is let's what we can do for you is organize your information in a way that you can do your science got it correct and so that's a real use case for them and that also requires AI that is true that is true uh in fact that actually Rings a bill there were quite a few people this year specifically at data hack Summit who came back and then said the same thing because because of the uh you know hype around generative Ai and what it can do and cannot do uh a lot of people or a lot of companies are making these internal promises around uh what can happen without realizing that that we are not there yet and uh I know at least three people who approached me and said that you know the summit was very enlightening to at least bring out where we stand today and they said that they would actually now need to go back to their companies and say that you know some of the uh project Investments which we had thought we would be making uh we should hold them or or at least change the expectations from them as opposed to what was doing so so it is it is actually uh you know increasing problem in the in the domain uh coming to uh you know that topic while we are discussing that uh how it's been your kind of uh you know interaction with generative AI what was your first aha moment and and uh is there something which you're already using in in your workflows uh using generative AI well so generative AI you know externally right using that is generally not being done in most Enterprises you know IP contamination reasons and so on we do not know what's right so but but um you know we do of course uh take in material which is commercial you know available open source material which is uh commercially with available with commercial licenses right now there the key thing that I have found is that where we want to not what we weren't able to do previously with respect to for example classification problems right so you you have you have open-ended classification problems nowadays now you know or with respect to text summarization or with expanding the text or with providing explanations to users for certain things I think there we have made a lot of advancement now where we haven't played a lot of advancement is in trustworthiness and that is something you know and I'm going to talk about this in a local talk here at Hyderabad in next week or so where I will you know talk about techniques that one can use to convert the output from for example llms to a more trustworthy output or at least be able to filter it and say this is trustworthy and this is not trustworthy and that is very critical for an Enterprise see for playing with it you can do it right so I and I saw this difference right that you know where you see uh you know like you know programming for example right A lot of people are really excited by the benefits in you know augmenting coders and now what we've seen is only those quarters are augmented who really understand computer science who really understand algorithms um when this thing gives you an output any any of the tools right you take Bard or you know Microsoft's version of chat gpd4 or you take Cloud to or any of these right or gpd4 you take any of these they make mistakes right yeah and see it's actually easier to write your own code than to find those rookie mistakes mistakes because it is just you know connecting together you know next best uh so people are saying you know uh coders are out of business uh you need actually even better training computer scientists to understand this yes right what is the API for this API call for this that that time it will save but it is perpetuality for that you need you know you know people who know what they're doing correct and then the uh Associated effect is okay yeah yeah the associated which right right and uh uh is it so in terms of uh you know reflecting on uh its impact on let's say the final products which you ship or the final products which go into the Enterprises uh where do you see it would make the biggest impact would it be in the let's say uh user interfaces would it be in uh information retrieval would it be in uh just make a uh you know uh making sure that a lot of uh rudimentary stuff and not rudimentary in the sense that uh it's not required rudimentary in the sense it's something which is very well defined and is just getting uh repetitive where do you see the biggest use cases of generative Ai and it making a biggest difference in let's say the Enterprise Solutions so if we talk about specific to so so there is already large workflows being set up for vision so that I don't have to pay anything right yeah you know governments are setting up surveillance systems you know if operations are setting up surveillance systems and uh you all you also have you know for example Adobe released their fire eye is is a generally available uh software now right so the question thing is all set at this point the the language models that we see today uh yeah there are multiple use cases which we thought you know or at least I thought and actually many people who feel thought we wouldn't be able to solve in the next even next 10 years we thought you know um I told CEOs in 2010 that come back in 20 years and I have to change my words I have to say okay I can do that for you now right so so for example this this open-ended classification problem right but I don't even know the class like for example if I am on Amazon so if I'm Amazon or Amazon equivalent right and if I want to classify a piece of clothing right I have a long description of it and I want to say what clothing is it what category should it go on right I'm a retailer I have 10 000 products what category should I do under right now otherwise what I would have to painstakingly do is I have to create a ontology and you have I have to sort of or a proper taxonomy at the very least if not an ontology and then have somebody set up some examples right and then you'll drain a system and go with that right that's the old school way of doing it now with the llms what I can actually do is something a lot more clever I can just say I'll take any standard taxonomy I will have my llm generate whatever it thinks is the classification of that clothing piece of clothing and then I can match this classification to my taxonomy and got it you'll be amazed it's you know very very accurate with something you know people thought this would not never be possible this open-ended classification problem or um in in terms of resolution or understanding text again logic no logic doesn't work math doesn't work you know just yesterday I was doing this with Bard right um I asked it a simple question two times one by two plus one okay okay so what is the answer you can say it immediately right two times one by two plus one do you know the answer okay first time it gave me answer one I said no your answer is wrong it says no your answer is wrong and it gave me a justification okay then I said no your justification is also wrong so then it gave me another justification with another answer which is 2 into 1 by 2 plus 1 is 3. and again it Justified that okay so now there and I would say not much but if anything that has to do with English language yeah today I read a paper that anything to do even with what we would call human creativity right in the sense of you know literature history all of this kind of creativity that is basically overtaken by llm for 99.9 percent of the population you know all the books have been put in here the whole Wikipedia is here so it can generate those to you on the Fly product descriptions for example you know once it has done the classification now it knows what what you know it can at least produce a skeleton for you you can fill in the blanks if you need you know product item list and it can produce a description for you and on by the same time the customer also it could produce a summary yeah so yeah I think a lot of these workflows will significantly change anything in that that involves running text no math no logic don't ask it to reason okay but true another thing is the teaching teaching part which of course Salman Khan you know from Khan Academy is using it to greatly yeah you know the multiple choice questions other than je and all of that you know because those are not a lot but anything that has English like you know TOEFL type of questions GRE type of question it can just do just like that yeah now the advantage there is that you do not have to employ a teacher to tell you that because the teacher may come up with some back expectation right but this thing will give you the actual reason why this should be the answer and it's like this again so then the question is are these tests valuable and so all of that you know so the Human Society will change because of that because the work workflows will change very true very true and uh and again uh you know uh from a use case perspective again in Industry coming back right so so some of these things uh how does it impact the let's say deliveries or how does it impact some of these Enterprise Solutions so so so uh I think all of these things that you mentioned we are experiencing now in fact I mean uh to what you said right so to if you would have asked me even five years back right between there is a field which is completely logical and there is a creative field which one is more likely to get automated I bet almost everyone would have bet on the logical field right and what we are seeing is almost uh you know diametrically opposite effect so uh so how does this you know kind of translate to the solutions you are building yeah I think I should have addressed that part when you asked me that right see the the way the the way this translates is that wherever we have Solutions which require reasoning systems we have to continue to use the good old-fashioned AI systems right right but whether it's AI systems or or you know there are there are systems which I have used where we had differential equations so you know you can call them AI or not but you know so those systems will continue to be the way they are right but then okay the systems which require as you mentioned right into information retrieval right so for example Google tried to solve the problem of Enterprise search today if I take a vector database I I put my entire um you know corporate information into a vector DB right and then I have my you know whichever flavor of a large language model I have right I can go ahead put my FAQs put my um you know manuals and people can retrieve information just like that in the Enterprise which was never possible before you know Google had Google Desktop for example right it worked well but nothing like what you can do now simply because um previously I was simply doing similarity search between strings or documents and so on that's all I was doing that's all I could do right but today what I can do is I can find semantically similar pieces of text for which which I could also do previously but I had to do a lot more work on that yeah I have to do a lot more work today I don't have to do right I I simply have to say okay this is in my Vector DB now based on my large language model I can bring bring back everything that is similar to what's in my Vector DB all the documents that are similar to this I can bring back and those will have similar meaning and I can sort of rank them in terms of you know how they are again llms don't solve this whole problem but this the information retrieval problem in the Enterprise can be dramatically solved is quite important when people ask questions of databases right now we had been doing this for a while frankly okay to create natural language interfaces Financial language Generations right in analytics for example right all your power bi is in Excel tableaus of the world right these essentially become extinct okay yeah because if you have this Warehouse or database okay all you all you have to do is literally ask a natural language question ask a natural language questions I can ask the natural language question and as long as I have fed my database schema into the llm right and perhaps I can provide some examples right here's a question here are some here's an SQL corresponding to that right once I have done that I'm good to go require some uh you know factual checking because we've done the study and what we found is that hallucination still plays a role here so and this is another thing I'll talk about the talk that's coming up for next week is that how do you get around the hallucination problem because although I will train it till my you know fine tune my llm with my question SQL pair but then how do I make sure that this is the right SQL now I cannot keep checking as a human every time a customer will come in the last time right it's not like I can go and check before the customer reads the answer is good right so so for that I have to make sure that the SQL actually makes sense right and you know one of the ways some people trying are trying that is you take the SQL convert it back to natural language and see if how similar those two are you know yeah yeah meaning what what yeah it's like property this correct correct different things but the idea is that you have to be able to put this factuality filter at least for the Enterprise right but once you have done that the problems which we thought were very hard to solve now we can solve and now you can solve it though by the way I have to say we have a compute problem because even at inference time um I need multiple gpus the a hundreds right which you know which I didn't mean previously so you have to see you know how much value this brings to the customers and you know whether people are willing to pay for this so that we have to see I'm sure people are willing to pay for it just you know yeah yeah and I think two areas which are related to what you're uh mentioning just kind of taking it a step forward one is uh you know how do you make sure that uh or let's say how do you evaluate llm because a lot of it and I know it's a you know very open question but uh uh you know it's a different regimen compared to what it used to be before right earlier if I had to do a classification model evaluation I would either use a confusion metric and then it was simple right it was very definitive with uh tell me about that I I I'll show you how that's quite completed but yeah okay yeah yeah so that was one part how do you address that and the second one was how do you handle things like model upgrades right so so the model was giving you one output uh and then let's say you used it in your Solutions but later on uh updated version is available and then some of the workflows which you had built get changed because of that and is that a problem with you kind of encounter 4C or is it something which can be solved it is the second one first right sure see the second one um is a hard problem right in two ways right one is if you have created a workflow in such a way that it depends on the particular way that your llm is giving you an output right then you have not created an Enterprise worthy solution that means you are meaning you have not done a good job short answer it's it's probably not even so POS yeah it's not even a POC if you are if your model is depend if you're if your Downstream workflow on you know what order this thing gives you an output okay I'll give you a simple example how to solve it right for example sure right um I could say um if I wanted a deterministic count of output I could say give me this output in an XML form with this this this this and it works a lot better than asking for sentences okay interesting right so with that XML or Json whatever format you your your whatever format you can even decide a format in fact with most of the llms that I find today if you do fine tuning with a you know maybe 20 30 50 examples of this you know question and then you know you give it an XML that you know represents the output of that then what happens is it knows that when that kind of question is asked I'm going to give that kind of output now yeah that is something that is something you can immediately figure out how to use and it will be consistent if anything the only thing that is well that might change is that um it may either not give you the right XML format then you already know it's not doing the right thing right yeah so which is not a common problem and happens quite a lot actually yes yes right so now um so so this will happen right but then we cannot make our workflow dependent on different versions right right now it is possible that some version may make your workflow obsolete and it makes a lot of other products obsolete simple thing just like you know when uh open AI release the Enterprise version of uh chat repeating basically a lot of startups became absolute instantly correct right true so that can certainly happen right that if you uh if your workflow is fairly simple right and it's just depending on the output of the llm and you just repackaging it then obviously you know there isn't that much value you're adding on top and if you're not then you know I guess you should if it if the new version provides you the functionality nothing no harm in that right if the new version breaks your functionality you only have a choice to stay on the old version and correct see that is why people do not want to put their Enterprise story with somebody who is not stable in this so I'll give you an example Oracle supports their products from 10 years ago 20 years ago and if you have any old product they'll support it for I think five years five years yeah correct correct AI support for two years one year not even two years actually again they have a certain number of people and how much can they do right so perfect well the the the short answer to that question is that we cannot make ourselves dependent on third parties like that if you are a big player in the business you have to have your own models correct correct and and the first one you you were mentioning that uh for the audience uh the first one was uh essentially how do you evaluate uh so now so see a lot of papers have been published on this right even in the last six months or nine months a lot of papers have been published and depending on the interests of the authors um they have done different kinds of evaluation right some evaluations work in this thing pass a test right some evaluations were of the type can it pass you know there's a different natural language processing benchmarks natural language understanding benchmarks um now the good thing good thing about the models that we have today is they pass those benchmarks now what we have to do as you know I'm again I'm talking about the Enterprise point of view right not talking with the academic point of view right academic evaluations will go on for sure from a Enterprise point of view what I have to see is is this particular version of my language model or any model right is is it fulfilling my specifications with respect to the output so there might be an llm which simply doesn't do give me that XML that I wanted right it does very good at generating my text but it doesn't do very good with structured out right they might see we don't have too many llms out there right now there's hardly two or three which are good enough right you know gamma 2 is good but then again you have to train it a lot to get something serious out of it is very simple the evaluation is I have 12 outputs I want from it right now which one gives me those outputs number one that's the minimum criteria yeah I want a new certain format I want it in a with certain level of accuracy can am I getting that that's number one and today I'm not asking for logic I'm not asking for math in the future I'll ask for that sure if you have that discussion I'll ask for that I'll say if I have I have Incorporated uh Bull from alpha into your llm as as one of the right it is an add-on but it only transfers the problem to an element it's in the llm transfers the problem tool from alpha but it doesn't actually use the logic the second very important so one is minimum requirement second is does it give me an output in a certain amount of time right because a lot of problems I need nearly instant latency right so there you know whether I use a 64-bit 32-bit you know four bit weights right all of those choices and whether with four bit four bit weights or with the sparse model is this model still able to do yeah so that scalability problem is a huge problem basically interacting with vikas maybe it can do right but there are a million customers I have correct right so you know what kind of compute story I have to create to be able to enable to serve that that many people yeah so that is a very critical piece because in fact in here accepting or rejecting uh models just based on that let's see if it doesn't meet my latency requirement I cannot deploy it in any case eight minutes later it comes up with the SQL not useful not to use it yeah customer asked the question 10 milliseconds SQL 100 milliseconds you know go to database come back then natural language generation for you know maybe 20 milliseconds that is acceptable Google but if I if you do this eight minutes right or you know or even one minute today the latencies of the order of one minute yeah yeah any things which are I'm running on my cloud on my GPU right still is a it's of the order of one minute true true so then you know so that's where I have to see will I do I take a scaled down version instead of 70 billion do I take a 7 billion one version of it do I have to deploy it on my on people's mobile phones yeah yeah so I think in the Enterprise those questions become quite important and then does it give me those correct outputs that we discussed so I think the correctness is critical in the scalability is critical both we will have more requirements when these things become more today this is plenty to filter it today great great uh because uh one question which is you know let's say someone uh so in in this slightly changing direction because of uh time uh for people who are getting into the field today right the field for them uh in the next let's say five years ten years would be very different from let's say what we have experienced in the last 20 years right so so how do you see that Evolution and as a person who is just entering it what would be your you know pieces of advice or what what should they keep in their mind uh as they navigate through their careers so um so one thing is this is a very exciting time to be learning material yeah but what what I do suggest to uh colleagues you know younger colleagues and to you know many friends [Music] you know there are many students who reach out right or you know there's a triple it here and so on one of the things I really call out to uh people is that you absolutely have to know the math yeah meaning there's there's plenty of people who are selling courses out there they say you just need to know how to uh run so how to just you know call an API the thing is in one year chat GPT will do it for you today I'm just using a name but there will be enough rules to do that for you that for you right so uh if if you want to be see again I'm I'm not saying that there will not be space for humans to run some of those things right but that those will not remain higher level jobs okay or very well paid jobs so if one wants to be in the situation where one is actually solving problems in the field um and you know I have seen this with my daughter you know you know she's a triple it they they absolutely have to know the math they absolutely have to know um you know what they call DSA in you know the all these interviews right so people think that you know I don't know why they you know why do they ask these questions the thing is if you have to understand things algorithms if you have to change these algorithms if you have to augment them right I cannot just use a llama too out there somebody else created it right I have to do a ton of things before it gets to anywhere near production will not do it for you right now when attribute does that for you then of course then we'll see the next thing but today um I mean I have seen plenty of people giving this advice you get in the field oh yeah you you did your graduation in certain thing and you don't uh you know don't like math you can still get into this and I I say please think about it carefully because it will be painful yeah because it will be very hard people will ask you questions to speak to you which will not be able to answer so it is better yeah that uh you know that one gets a good grounding in Millennial algebra one gets a good grounding in the fundamentals one does not need to know Linguistics one does need to know why these things connect with each other similarly for vision to be able to properly understand common National neural networks for example you really need to understand how Vision Works yes you can run a script okay but so can a million high schoolers so who who do we hire them right somebody who can create a new algorithm or then I can learn that script right yeah copy the API put it in my code true there are times the focus on the fundamentals is critical that is people five years ten years 15 years down the road you know coding goes with it the math is Chef true true thanks thanks for uh that uh because uh just one last piece uh before we conclude the podcast and uh this is just to know you better as a person so so what I'll be doing is just a rapid fire round and then uh whatever comes as first thought to your mind I would like to listen that right so uh uh chai or coffee MNC MNS MNC okay uh favorite ml algorithm entry boost Cricket or football I would like to walk fast [Music] spectator sports and Linkedin or Twitter or X I think LinkedIn Twitter is X now uh Kindle or physical books physical books okay favorite movie favorite movie I think the one I liked the most in my son watches it all the time is Gandhi okay okay but there are things from 1930s which our family likes a lot but anyway that's you know that's black and white so by Gandhi is and uh favorite music so I I really liked uh carnatic music previously I had very little appreciation but you know I've lived in Hyderabad for you know 12 30 years so that's that's that's the that's my chaotic music sure great thanks thanks vikas for all those answers discussion and you know I think a lot of things that you mentioned that that almost condenses you know years of experience perspective and then you know spending time in the field so so thanks a lot for bringing this out for the audience and I'm pretty sure that the amount of learning this conversation has can can go a long way in helping people in their own Journey so thanks a lot for sharing these insights your questions help me clarify things in my mind too thank you [Music] [Applause]
Original Description
In this third episode of Leading With Data, Kunal Jain (Founder & CEO, Analytics Vidhya) is in conversation with Dr. Vikas Agrawal.
Dr. Vikas is a Senior Principal Data Scientist for Oracle Analytics Cloud. His role is to design, develop, and deploy AI applications in ERP, SCM, HCM, CX, and MFG across Fusion and NetSuite customers. He is an Electrical Engineer turned Computer Scientist from IIT Delhi, having worked at CalTech, Intel Corporation, and Infosys.
Watch as Dr. Vikas delved into:
👉 How he stays up-to-date technically & keeps a macro view of business solution
👉 Why is so important to figure out the problem before building a solution for your client?
👉 What steps are followed in building a solution for the client?
👉 How have LLMs made it easier to build solutions?
👉 Why it is so important to build trustworthiness in models?
👉 Why coders aren't out of business just yet!
Leading with Data - Podcast with Kunal is now available across leading podcast platforms as well. Check it out here & subscribe:
1️⃣ Spotify 🔗 https://tinyurl.com/2av72fpn
2️⃣ Apple Podcasts 🔗 https://tinyurl.com/bdcmzcae
3️⃣ Google Podcasts 🔗 https://tinyurl.com/mhe8977d
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Analytics Vidhya · Analytics Vidhya · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
The DataHour: Data Science in Retail
Analytics Vidhya
The DataHour: Anomaly detection using NLP and Predictive Modeling
Analytics Vidhya
The DataHour: Energy Data Science Project from Scratch
Analytics Vidhya
The DataHour: Explainable AI Need and Implementation
Analytics Vidhya
The DataHour: Google Cloud AI/ML
Analytics Vidhya
Prediction to Production in Machine Learning #machinelearning #prediction
Analytics Vidhya
Practical Applications of Data science in Ecommerce
Analytics Vidhya
How to tackle Overfitting?#machinelearning #overfitting
Analytics Vidhya
Building Data Pipelines on GCP #googlecloud #datapipelines #data
Analytics Vidhya
Hands-on with A/B Testing #abtesting #datascience
Analytics Vidhya
Efficient Implementations of Transformers #transformers #cnn #machinelearning
Analytics Vidhya
Modern Deep Learning Architecture #deeplearning #architecture #deeplearningtutorial
Analytics Vidhya
Key steps for Designing Artificial Neural Network (ANN) for Image classification #machinelearning
Analytics Vidhya
5 things you should know about Azure SQL #azure #sql #datahour #datascience
Analytics Vidhya
AI & ML in the Automotive Industry #machinelearning #ai
Analytics Vidhya
Building Machine Learning Models in BigQuery
Analytics Vidhya
NLP aspects in Telecommunication Industry
Analytics Vidhya
Practical Time Series Analysis
Analytics Vidhya
Fundamentals of Quantum Computing
Analytics Vidhya
A DAY IN THE LIFE of a Data Scientist (From waking up to working on algorithms)
Analytics Vidhya
Classification Machine Learning Model from Scratch
Analytics Vidhya
Knowledge Graph Solutions using Neo4j
Analytics Vidhya
Model Guesstimation (MLOps)
Analytics Vidhya
ETL Pipelines in Google Cloud Platform
Analytics Vidhya
Key steps for Designing Convolutional Neural Network(CNN) for Image Classification
Analytics Vidhya
Getting Started with AWS EC2 #amazon #aws
Analytics Vidhya
How to Use Azure NLP and Graph Databases for Intelligent Knowledge Mining
Analytics Vidhya
Certified AI & ML BlackBelt Plus Program #shorts
Analytics Vidhya
Visualizing Data using Python #machinelearning #visualization #python
Analytics Vidhya
DCNN for Machine RUL Prediction using Time-series Data #timeseries #machinelearning #datascience
Analytics Vidhya
M in ML stands for Math & Magic
Analytics Vidhya
An Unsupervised ML approach using Clustering
Analytics Vidhya
Customizing Large Language Models GPT3 for Real-life Use Cases #gpt3 #datascience
Analytics Vidhya
Model Parameters vs Hyperparameters - Techniques in ML Engineering #machinelearning
Analytics Vidhya
Practical MLOps #mlops #datascience
Analytics Vidhya
Data Engineering with Databricks #dataengineering #databricks
Analytics Vidhya
Multi-Objective Optimisation
Analytics Vidhya
When Airflow Meets Kubernetes
Analytics Vidhya
AI in Banking
Analytics Vidhya
Learn Convolutional Neural Network for Image Recognition
Analytics Vidhya
Extracting Value from Data
Analytics Vidhya
How to measure Marketing Channel Effectiveness
Analytics Vidhya
Transforming Lives | Data Science Immersive Bootcamp
Analytics Vidhya
Stock Market Analysis - AI driven approach
Analytics Vidhya
Become a Data Engineering Professional in 2022 | Future Trends + Skills Required
Analytics Vidhya
Ensemble Techniques in Machine Learning #machinelearning #ensemble #datascience
Analytics Vidhya
The Power of Visualization | Tableau Full Course | Analytics Vidhya
Analytics Vidhya
Demand for Data Engineers is on the Rise | Data Engineer | Analytics Vidhya
Analytics Vidhya
Data Visualization in Data Science | DataHour | Analytics Vidhya
Analytics Vidhya
Role of Optimization in Machine Learning & Deep Learning | DataHour | Analytics Vidhya
Analytics Vidhya
Solving any Machine Learning Problem | Approach and Steps Involved
Analytics Vidhya
Topic Modeling Explained with Implementation | Using LDA in Python | DataHour by Arpendu Ganguly
Analytics Vidhya
Data Engineering in E-Commerce | The Best Case Study
Analytics Vidhya
Introduction to Classification using Azure Machine Learning | DataHour | Analytics Vidhya
Analytics Vidhya
Introduction to Federated Learning | DataHour | Analytics Vidhya
Analytics Vidhya
Diffusion Models for Generative Arts | DataHour | Analytics Vidhya
Analytics Vidhya
Master Google Analytics in 1 Hour | DataHour | Analytics Vidhya
Analytics Vidhya
Learn Hypothesis Testing | DataHour | Analytics Vidhya
Analytics Vidhya
A Practical Approach to Kaggle Competition | DataHour | Analytics Vidhya
Analytics Vidhya
Making AI work for Business | DataHour | Analytics Vidhya
Analytics Vidhya
More on: LLM Foundations
View skill →
🎓
Tutor Explanation
DeepCamp AI