Topic Modeling using Transformer Based Embeddings | DataHour by Bharath Kumar Bolla

Analytics Vidhya · Intermediate ·🧠 Large Language Models ·3y ago

Key Takeaways

Introduces topic modeling using transformer-based embeddings and BERTopic

Full Transcript

so welcome everyone uh this evening so my name is Bharat I am your presenter today I am currently I am working as senior data scientist at Salesforce Hyderabad okay so this is some of our work with a vasudeva um uh and few other folks who did this work okay so this today's topic is about topic modeling on consumer financial production Bureau data an approach using bird-based Temperance okay um so this work we presented uh in an IEEE conference and it got accepted last year okay and this is the paper if you want you can explore it I can paste a link of the paper also if you are interested um so basically today basically our goal is I wanted to make it a bit more of a research purpose okay a research paper exploration okay so I will walk you through the motivational background like why we took this approach right what is the problem statement what are our aims and um objectives what is your methodology right some results and discussions conclusions and recommendations okay um coming to the motivational background right so basically the background is we did this analysis on a financial consumer production data so I hope by now most people know what is topic model topic modeling is nothing but extracting topics out of the free text okay so what is free text Suppose there is some text available so you wanted to know every text has something it is time to convey some meaning right there is some top topic it is uh trying to explain right so whether it's a fine uh the topic can be anything it can be about business topic Sports movies arts history or anything all right um uh yeah I would give you the link to the recess paper as well okay so the topic can be editing so if uh in such cases when your topic is anything right and you wanted to um explore all right what are the topics from without any labeling data so basically I hope everyone knows now uh what is a supervised model on supervised model OKAY a supervised model is something uh where we have labeled data uh suppose for example you wanted to say this text belongs to sports this text belongs to uh some movies this text belongs to something right so that is a supervised model so you uh right so when you have something already labeled right and then uh when new text comes and based and you train a model you can classify any new text based on the already the model you did it on a supervised data so that becomes a supervisator model so basically topic modeling is an unsupervised model I hope everyone understand um yeah right what is unsupervised model because we don't have labels here the idea and the task is to extract the labels from the text given text okay so this is a very very um this is not a very trivial task right but a great deal of amount have work has been done over the past two decades starting from LDA which is a very revelationary method to the current bird-based approach okay so this is device model and there is no guarantee that if your topics are really uh well separated well distinguished right so often you need some domain knowledge so who provides the domain knowledge right so probably you are a data scientist right you develop you say that these topics are good but someone needs to validate that uh topics so you can only compare two different models whether this is good or this is bad but there is no 100 right or wrong or something that says that these are the perfect topics this text contains and often it becomes very cumbersome when you have huge amount of text but you can get some kind of feedback from the domain experts who are related right in your company or in your industry or in your Academia who can give you some guidance about the quality of the topics that you extracted from them uh from the text okay so the motivation here is to create an intelligent NLP solution for finance organizations right to atom automatically assign the complaints received by the consumers to the subject matter experts right so and thereby reduce the response time and also see the manual efforts of the customer service professionals right so imagine you uh on a chart or something you send it um you send a message or a text something which says that you have a problem regarding the tally of your account or something like that or you have some other um uh some other problem that you are facing and you wanted to get it resolved right so that that particular thing that your particular concern should be addressed by a subject matter expert pertaining to that particular domain right so in a financial organization like a bank or like a like a trading organization or like a financial or like a Fidelity or Vanguard or these kind of organizations right so every complaint uh complaints needs to be handled by different set of people right who handle various brand domains right so it is very important for the topics to be diverted to that particular domain so uh and manually it is very it is impossible almost impossible to do that manually because you receive number of complaints and uh this becomes a very cumbersome task even to Route manually as well so can we have an intelligent solution right uh to label them and once you label them right you can train a model once you train a model next time you can automatically label these Solutions belongs to this thing uh you can do a certain domain expert right so the problem statement here is uh the to find out the best topic modeling technique to create the topics from complaints uh compliance Text data right this text data is publicly available okay um text data is public available Financial organization right this is a consumer financial production Bureau public level Text data and to find a way to project prioritize the complaints to save the organization from reputational logs right so the more uh responsive the organization is towards these customers the better it is serving the customers right and in a in a competitive world like this it is very important to be serving your customers a very responsible right so our objective one right so first of all our objective one is in a research you need to have some objectives so first our object is to investigate the utility and performance of existing word embedding techniques right so for example uh I I really hope that you guys know what is an embedding by now uh embedding is nothing but a representation of a word in an n-dimensional space right so for example uh you train a large unsupervised language model right and for example there is a word called king or a queen or anything and each word you are representing in the n-dimensional space that n-dimensional space is nothing but a 32 Dimension or 64 Dimension or 256 Dimension or generally even and multiplied by eight because of the GPU processing okay if right so you're representing a word in undimensional space right and the the two if two words are closer that means two words are very related and you can check if two words are closer by simple uh cosine similarity so now we wanted to see traditionally in a statistical way we have topic modeling techniques like um LSA and LDA right latent semantic um LSA and also LDA written delete uh all right LDA right which is um very two famous and conventional models right so we wanted to explore the current a Transformer based embeddings which are nothing but uh like Bird robot and distal but which are nothing but the birth variations trained uh domain are trained uh to make sure that we have a smaller language model right and versus filbert is nothing but a bird trained on financial text right so if we wanted to compare these four bird-based models with LSA and LDA and we wanted to see the quality of the topics okay so these are the two objectives we wanted to serve right so the data set we have is a CSV file uh right uh which is directly downloaded from cfpb official website all right as I told you cfpb is uh Consumer Finance production Bureau right um and uh the data identified for the studies in open source data um data set has columns like company product State uh date of complaint right because along with the consumer complete narrative basically the uh and the text of column okay so this consumer company narrative is the column on which we are doing the analysis okay so this is a quantitative study which evaluates performance of different topic models on financial Text data so for example this is a sample data set you can see that so this is one of the consumer complete narrative okay so uh you can see PayPal Resolution Center ignored their own policies is one complaint all right so you can see that because um you can see that some of that data is already masked because that can reveal personal information so you can see that uh some of the information regarding the account names or complaint number or these are already masked so anything with xx17 is something like a date or something right uh right so these are pertaining to these uh campaign right so as to protect the uh to protect the consumer personal in Privacy Information they mask already this kind of data so this is not a huge problem for our thing so I hope everyone is um clear with the data so the research methodology is uh basically we wanted to First do Lea and lsea okay and these topics are produced based on word frequency right so basically you know so basically LDA and LSA both are based on word frequency so you wanted to associate the frequency distribution of a word uh based on these are these are probably stick models right so and you can see that these fails to work on large purposes they have sparsity issues so that is one of the main problem and it also ignores the order of words right so one this is one of the main problem it ignores the water order of words because hence semantic relationship between the words is not captured so that is the main problem right for example uh if you say something like um bank and banking can be two different words right so it doesn't capture the relationship between these two words are being the same right for example in your in your topic in your text if you make a compared with some spelling errors that can also lead to of course you can do limitization stemming that still destroys the some of the world all right so if there is a spelling mistake or anything like that it starts to see it as a two different words and also it fails to capture the semantic relations between various words so these are the basic drawbacks of LDA and LSA right but whereas embedded space models like vector based model right the topics are extracted by grouping the similar meaning of words together right and these Works effectively are large partners and the last important thing is it consists the order of the words and captures the semantic relationship between the words so before getting into the diving or getting into the data and results or anything it is very important for us it's very important for us to analyze the data right this is called as exploratory data analysis so if you can just see this right so you can see that you can see that mostly if you just see the number of complaints Right company wise complaints so so we have a company column and you can see that most are related to credit reporting agencies right so 40 of the data has created reporting agencies and most of the complaints regarding credit reporting is probably there is you or someone has already paid the bill but it still shows not paid or something like that right that severely affects their credit score so 40 percent of the components are regarding credit reporting agencies right um all right and these three are only between between Equifax and you can see that Transunion and some other company right so that is the main thing and and also you can see that 40 points right under first these companies are basically uh credit reporting agencies and the second analysis is you can see that right um uh you can also see a simple a Time series based analysis like yeah you already have a date column in your data and you can see that uh when and how many complaints you have received right so this is the complaint number so you're just uh accumulating all the compiled numbers for each day and you can see that um overall the complaints trend is increasing right this shows that uh overall for various factors since 2015 the complaints are increasing and you can see that there is a spike in somehow there is a spike in probably 2017 and in uh between 2017 and 18. uh not sure what cost that particular Spike uh maybe there is some Equifax data leakage or something like that I'm really not sure but uh if you just do a moving average kind of thing you can see that uh certainly you can see that the components are increasing so it's very important for us to um for us to uh prefer most of the companies right so these complaints are increasing and they need to take this complaints very seriously so analysis right so for text cleaning we remove the white spaces and we did say the complaint test using regular expression uh we use punctuation from the text screen using string we press function and we we inbuilt the stopwatch from nhdk library and right and we did the tokenization especially we did the tokenization for uh these what you call as for LDA and LSA methods right the text is initially converted to lower case using a lower function and this text is limitized which means the converting the words to its root forms right because um stemming uh right so converting the words to root form is a good better method than stemming itself and uh we explore the complaint text to find out number of words characters hashtags stop words negative words within that text before converting them to token right so coming into the research methodology right so basically what we did was we did the data right uh collected the data and then we did the data cleaning all right basic data cleaning as I mentioned uh in the previous slide and we did the topic morning and evaluation right so this topic mounting and evaluation contains two steps one is non-transformer based and Transformer based right and non-transformer based um are the ones um uh the ones that I'm talking about where you use finboard uh the Transformer models trained on uh all right a bird-based Transformer models uh for getting the word wise uh Vector vectoral representation and uh right uh and also non-transform methods such as LD and LSA because you always have to um how did uh right uh yeah so we I I will ask you I will check get back to your question after the anomalies I just need to know pradeep what is the anomaly you meant okay so we use two metrics to compare right one is CV metric uh right basically it measures the inter cluster distance which means that it signifies the distance between two clusters and it should be maximum for the ideal model right so if the two clusters are separated well that means your topics are well separated okay and you must metric versus the representation the intra cluster distance within the cluster right so for example um you have within the cluster you have some uh data points right so within the cluster your data points are uh should be very close which means that it represents the distance between two topics within the cluster and it should be minimum for the ideal model right so your CV should be maximum you must should be minimum so these are the two metrics we evaluated apart from that qualitative analysis as well so the research methodology is like this as a as I mentioned in the previous slide we use two methods uh one is for Transformer based models and one is for uh non-transformer based models um a spike in number of components in particular so um we we didn't do I mean of course we since this is unsupervised analysis we don't need to do any anomaly detection and and we didn't take the complete data for our analysis I if I remember correctly I think we took uh I think last 365 days or two year data something to do the analysis okay so uh because even this uh sample data which we were extracting from the last one or two years we had some compute complete problems that's why we restricted ourselves to um a few years we didn't take the entire data because the entire data is very very huge I hope I answered your question for you but it's a very good question so the top architecture right the methodology is for um probabilistic based models like we take the tokenized component sentences basically we do all the data pre-processing like climatization removal stock ports and all these things right so basically we tokenize and we do limitization and we do most upwards and and then we apply tfidf method to get the weightage of each word in each circle right and we apply LSA and idea and we extract the topics so the bottom architecture is about the birth topic so this is uh this is a library created by Martin gordonhost I would introduce this library to you as well once the talk is done um and you can visit the library to know more about how it performs so here as you guys all know for transformable based models you should um right you don't need to do any data pre-processing because um the Transformer based model understands the sequence relationship between words punctuation marks and all these things right so it is um uh you should not apply any pre-processing methods on when you're applying a Transformer based novel so that's why we do we took untokeness complaint sentences and we uh so this is the birth topic package a library built by Martin Garden host so here the first step is to you vectorize this uh right and your vectorize each topic right each text right each uh each topic means what I meant is each complaint right so for this since you wanted to extract the n-dimensional representation at a complaint level you use sentence Transformer right so you use sentence Transformer again uh right and this sentence Transformers can be derived from about digital bird feedback right and robot and once you do sentence to have a sentence Transformers right you can but see at the end you have to reduce the dimension but because these these have a dimension of I think uh based on the model 512 or 256 or something right so each complaint you are projecting in a n dimensional space of 256 Dimensions let's say for example right to uh I mean uh right you already you are reducing the dimensions using U map right because uh that is for the reasons of uh that is for the reasons of making it more easy for you to manage and understand and visualize right so you can use umap or tsne but for some reasons umap is a better method because tsne requires more compute okay uh tsne uh stands for uh t uh stochastic neighborhood embedding models uh umap is a different uh model which is also which also works on non-linear dynamic reduction method I think most of the beginner you guys know that PCA is principal component analysis is one of the linear Dimension reduction method so on on the other hand U map and T sne are the non-linear dimensional reduction methods right so your data need not be in a linear fashion right so even these are placed in nonlinear so you we apply umap right you reduce the dimension so this is one of the hyper parameter if you control and you do clustering right you get and when you're doing clusting um you can choose hdb scan for the clustering you can do encoder but that is not available uh here in this package right uh you only have you map for the dimensional directions okay so once you have HTTP scan you get CTF IDF for the importance of the words right and then you get you generate the topics here okay so make sure it's a very good question so encoder it's not available in this package so now as I told you uh CB and metric and UMass metric are the two metrics we did so now let's see the results how we got so you can see that LSA topics are mostly like these are the topics we got uh depth collection mortise loans disputed and inquiries service fee and credit reports so we don't get these topics directly we still have to combine because actually what we get is some the sum of the words present in one topic right I I hope if anyone knows about topic right so these topics when you combine with your domain knowledge you you can assign a topic at a very higher level actually you get only words so for LDA you get mostly about Auto Loan mortgage loan uh account charge of account reporting fraud debt collection late fee payment credit card as I said before 40 of the topics are about credit reporting agencies right so and all these issues are mostly about credit reporting if you just take a look at it credit reporting right credit card reporting uh right so all these late fee payment credit card reporting uh account reporting loan all these are related to the trade uh loan and reporting agencies right uh this is actually this entire research is a thesis done by uh vasudeva uh if you want uh I can show you the Raw results as well but for now I will just show you this high level difference so now in this you can see this topic modeling so this is these are the results generated by the purely bird-based Models All right so now you can see that so this is the cluster diagram introduced in the topic distance map so now you can see that you can uh you can see that these topics how most of these uh each um each thing represents one cluster and you can see that most of the thing represents the bankruption fraud so we have to still so these topics foreclosure and these doesn't come out naturally you have to give a label based on the uh based on the words that come out of these topics okay so these are the topics you get from a bird-based model right and if you see uh this is the topic name we we assigned a topic name delayed and re-raced complaints cluster right so some of the major words in this cluster right for example uh in this delayed and relayed right for example this particular customer where I am pointing my mouse Here auto pay e-loan statement Banker penalty and alert and tape right so our the person who is doing this research or thesis right is a financial domain expert so he enabled these topics are delayed and re-res complaints okay and the second topic is about e-banking so this is all about payment judgments these are the major words list right so then the third topic is about auto pay right so now you can see that auto pay sometimes you you are not able to set up an auto pay or your auto pay is not activated so you can say that auto pay card holders role auto track already all are regarded to auto pay collection right so then the fourth topic is about delinquent depth collection deliquid debt collection if someone is not able to pay their loan and after 60 days uh 90 days if you're unable to pay your installment then your account goes to a telegrams so so when your account goes to delinquency then they do recovery and all that stuff so that is why this has very related words like delinquency depth recovery collection discrepancy major right so all these are related to delinquent collection the fifth topic is about foreclosure foreclosure is about so she was not able to pay if someone takes a home loan and not able to pay their loan for more than three months right then the uh the bank or the financial institution which provided the loan tries to retake the property from the uh from the loony right so that is called foreclosure so foreclosure income value freezer Equity deed specialist right so all this stuff so then there is a root customer support right so for example you try to call make a call and you someone gets harassed dialer you dial you rule there is a root talk email recipient clock right it's all these points of the root customers so and then uh frankly the last 10th one is um sadly bankruptcy cheated scam game all these are bankrupt sent for so you can see why I why we uh highlighted these two uh uh these two in bold are because uh these are the two unique topics that we got from bird-based model right so here we are comparing Bird versus um fin Bird versus um Roberta versus tissue world right so next I would walk through you through the fin Bird right so now if you just see you can see that it just times falling through Bert and fin birth you can see that fin bird has many more topics is this obvious and clear to everyone you see right so Finn bird has more top uh more topics and also more well separated right so now uh so these are the topics that feel better so now you can see that some of the uh some of the it got some very unique uh topics such as fcra ftpa of course we have FTB and FTC right um uh these are some Financial uh uh Charter okay and cash net issues right uh again uh Delete discharge duplicate disclosures right something like glitches violations and all these things so it has a very unique way of file because this bird model is not trained on all the Corpus yeah it is on the same text it is on the same text okay so this bird model is uh the embeddings that you get from this bird model are trained on a financial purpose that's why you get very specific uh very unique topics here okay and this is distilled birth right so you can see that uh see this is for example this is one cluster this is one cluster you can see that within the cluster um these are see for example within this cluster you can see if there's a very very close and you have a lot of topics this is one cluster this is one cluster but here digital but is also good in the sense that um I mean the problem with distribute is it's a inter intra cluster distance is a bit high right and it also gives some different types of topics okay so again these are the topics that you would get from the uh distributor okay and digital body is a model uh area it's a lightweight model uh it doesn't have huge uh the number of parameters for distribute are very very less similarly the robot as well right as the name says but when you have a distilled model you have a distribute mode okay so if you see robota is not doing well because you can see that it didn't identify many topics probably you can see this is one topic this is one topic this is one topic a third five right so this clustering figure itself reveals that roboti is not doing that great right and these are the five topics what we got okay and you can see that bonus reward points is surprisingly uh no other topic no other model discovered this topic uh right bonus Reward Points uh identification traffic is like a fraud thing it's already there so this is probably fraud it already directed but uh foreign funding FCI FMI and these are the regular ones right right so what does this points out so now uh now we have to compare Apples to Apples by putting everything on one frame what are the topics each uh each model extracted uh from this so now you can see that so if there is a unique topic right so we are putting it in a unique column here right so for example uh you can see that a delinquent depth collection is same as collections here right so only bird and finbert extracted these two topics right again e-banking and e-banking are digital fraud uh these two extracted so now if you see the auto pay auto pay only bird extracted but here if you see that finbot extracted all these five topics so one in one is related to specifically related to TransUnion FTB FTC credit reporting all these five unique topics and also right are reported by Finn Bird right and also the number of topics it uh coverage right for example it is mostly able to cover all the topics that most other models is able to cover and also it reported some unique topics that is why we can see that fin but uh because uh the domain knowledge right so the model uh the bird model that is trained on financial Corpus is a better model to use and overall these models are better than your LDA and LSA right so now when you compare the the CV right the CV and UMass metric you can see that so LSA and LDA so the CV is high is the higher it is better so now you can see that fin bird has a CV of 0.33 right better than even the birth model all right not not huge not huge difference uh but still it's all better than the feedback uh the bird and even LSA and LDA right at the same time UMass right which is uh supposed to be uh as low as possible even here uh you can see that uh finbot is a better model all right so your intercostal distance uh intercluster distance okay um sorry on both the Matrix finbot is a better model so here you can analyze quantitatively and this kind of analysis shows you a qualitative analysis so it is very important right sometimes quality quantitatively you may get better results but qualitatively they may not translate to what you want so both ways when we analyze uh what we see is uh write our basically our our inferences are at a two two uh we got two kinds of references one is your birth based model sir if you just put these bird-based models are having a better metrics quantitatively and also called schedule and within the bird-based models so if you have to rank on a on their basis right so you can see that fin bird is the better model followed by bird followed by distribut and followed by um robot right so so these are the uh findings from this okay um and sinbird is able type identify more topics as you can see here right and it also is I identify common topics from other models and also some unique topics as well right so so the conclusions are right so based on our experience and review of relevant literature right so Legacy standard topic you know models cannot capture the importance of context right as these algorithms ignore the order of words but sentence based Transformers capture the importance of the word or of text and capture the context okay so this is the take home message which I wanted to give it to you guys today okay and the second conclusion is that so by comparing by doing a quantitative analysis with which with a certain kind of uh with a with some degree of certainty you you can say that uh by comparing your UMass and CV metric but Transformer based models are better than non-transformed this topic models right in capturing the context at topics with acceptable inter cluster distance and intercostal distance and the third thing is out of Transformer based more embedding space topic models fin bird works better on financial data right because it has trained on it uh domain same domain purpose so filbert demonstrated good intercost and intercoms to bonding and it also able to identify unique Topics in finance when compared to other pre-trained birds okay so the future scope what you can do so uh bird topic packages is you map distance reduction techniques right to improve the efficiency of umap uh umap can be replaced with very short encoders um someone asked a question here about Auto encoders right so who is that guy who asked about uh uh we also have yeah uh suggested that can we use a um yeah Auto encoder technique right so yes we can use Auto imported techniques I think they are part of the package and recently a very latest version of the package has already been released right so you can use um uh you also encoders you can try different encoders and see and also uh you can play with the dimensional reductions right so as I told you uh you can reduce to 20 30 50 as a hyper parameter tuning right so you can again have to do the same aspect by comparing and the recommendation too is you need to find a way automatically assign the new companies to one of these clusters so now uh we we can bucket these complaints and in one of these topics and you can uh train a supervised model right and then you can see how your model is performing at a supervised level the third thing is there is some kind of time motor and out of memory right so yeah for example someone asked uh yeah we used only 2016 data one year data we couldn't use on entire data uh that can be worked to handle complete data sets which would be beneficial in the future studies okay so we usually one year data because we have some compute issues but if you can handle um using a cluster computer Cloud computer AWS credits if you have something you can handle at a higher level as well okay you can handle on the whole data right so these are some of the recommendations okay uh and um you can also do some work on recently on uh how the topics change over the period of time right so that is also an important research aspect that is being actively pursued everywhere so you can also uh pursue this topic later as well okay so how the topics change right so um it's just this is I can show you this is a very good package on this as well so this is one future thing if you want to explain so this is uh basically about our about the topic that uh I wanted to discuss today and uh some of the things I wanted to introduce so this is our paper um I would give you the paper link someone asked for a paper link let me copy this and get me the PayPal name so this paper is freely available on arxav and you can uh I'll give you both the links right so I will give you the original I trippy link as well as the ARX I will link as well and if anyone is interested to if anybody is in uh code we didn't open source and I'll try I'll try to put it as well uh if anyone is interested we can pursue uh extending this topic uh right I can I'm a thesis supervisor for upgrad as well I have so far I have guided 50 students if anyone is interested we can pursue this topic if anyone is a master's level student who is already pursuing NLP or something like that and um so this is a nice uh topic extension topic right we can pursue as well okay so uh ping me uh later on LinkedIn or connect with me so this is our paper okay coming back to the birth topic package so this is the birth topic package okay so now you can see that it does vectorize on basic based on a CTF IDF okay and on online converter as well okay okay I did did I not share to everyone okay let me share the paper links I have to everyone I think by default I shared only some moderation let me share it everywhere so this is the word topic package okay so now you can see that the architecture everything this package is open so you can use it right so these are some of the uh hyper parameters we can tune here so now you can see that um uh where is that various things right so now if you just see it I just wanted to show API both of your web pages backends yeah so now if you want you can have a cluster so how many topics you want to reduce there are a lot of things you can explore right so this is I just wanted to this is a great package and by the way if anyone is starting in data science the person who created this top uh this entire package is a uh is not a real I mean he's not a trained data scientist but out of some curiosity he created this package and it got quite popularized recently right you can see that it has 3.5 K stars and 449 uh mergings right um so everything is here so you can have our term scores topics per class and there are ways to do it you can do some plotting as well some documents right so one of the main problem is it has sometimes it takes a lot of time uh right so this you can get this kind of colored topics as well right so this is the package I wanted to introduce you which we used for our analysis okay um and um yeah any questions I would like to take so yeah let me start uh what is bird-based analysis so bird is a Transformer based model where um so mayank I think but I'll just show you I'll answer Lively all these questions so Bert is you can see that um so if you want you can see what is hurt and uh I will just post you some of the links here so that uh you can check this so it is a Transformer based language model where it is trained on unsupervised tasks where you can you get a better representation of the word in an n-dimensional space okay uh and how are Financial companies families it's a publicly available data set so probably I think you can search for uh cfpb data you should get it uh freely yeah see you can download the data as well right get data download the complaint data it's very easy I'll show you this link casual if anyone wants to use for your PCS or anywhere I hope you're getting my links whichever I'm sharing can anyone come from and did you also use the lower non-numeric kind of basic text cleaning for good no for both we did do uh for bird we didn't do any uh data cleaning because it's not required uh what advice will you give questions get into data science and machine learning so so so data science is a very rapidly changing field so uh I mean I would say that you need to have a strong computer science and statistical fundamentals acquire those fundamentals and come get into getting that knowledge and things in data science is not too hard but you need to put some extra effort because the field is changing very rapidly the kind of advancements that are happening uh are happening at a very rapid phase so you can cannot catch up everything first of all you have to think that you the fundamentals need to be strong everyone is everyone is looking for only fundamentals no one is looking for Extraordinary things so if you have this fundamental you can start as a data analyst or you can acquire some skills and then slowly you can have it build your own repository you can report on any of your own skills so this is the study with software and commands we have to write instead of python uh Fatima I didn't understand your question exactly instead of python I mean uh so with software and command we have to write in sub python no we use everything in Python itself so um okay so can you please give us some examples about how various industries use topic modeling uh uh topic morning to Target the business some so uh use case of it so it depends right so it depends on mostly um every industry is trying to understand whether you call it a topic or aspect recently also we did some analysis on weak supervised learning with respect aspect right so if anyone dealing with the text-based complaints right so everyone is trying to give some reviews so if someone is giving review what is that review is about they're happy about price they're happy about uh the quality are they happy about the customer service so you need to know so there is a second level of higher level research where uh the sentiment associated with aspect so that is the higher second level of resource that is everyone is actually pursuing okay yes so yeah it's a question from uh for sentiment analysis topic modeling industry for sentiment analysis you don't need topic modeling as I said uh for if you were assigning a sentiment for a particular aspect as I told you are you happy about price are you unhappy about quality which one you are unhappy about probably you you might be okay with the price but you're not happy about quality so there is a sentiment associated with both the things so associating sentiment with an aspect is an important research area is dimension reductions used to reduce the sentence that have been transformed is it right to say that umap does featured it yes um does feature reduction exactly aware sentiment analysis no sentiment analysis again uh you you there is a there are a whole lot of different approaches for sentiment analysis um how to test if data is leader though through EDR packaging most data is non-linear that you can trick 90 points many instances data you never find data in a linear model in any way okay unless under until in ideal cases most data is small linear most factorization one of the reasons we are not using PCS the data may not be linear can I explain how do we know whether data is so I mean see you see if you if you have PCA right so basically if you do PCA uh your your component should be separate separable linear right so for example if you take Iris data right Iris PCA so if you see Iris PC I'll just point out to um right so if it's linear right if you just check this figure right so for any linear if you should be able to for any linear data you should be able to draw a line like this and separate the data so if you can draw a line like this and separate the data then that's a linear method that is why K means is a linear linear clustering method right so whereas hdb scan hdb scan is a non-linear Dimension a non-linear clustering method right so what does that mean linear clusting and non-linear cluster I am going into a completely different topic but this is a very important thing that everyone needs to be aware of for example I'll just point out to the string uh this thing so see so this is a game instruction where you can separate by just a line a line is nothing but a linear this is a linear line you can separate but for example um yeah this is a wonderful picture from Psychic learn itself right so now uh if you have two circles like this can you separate by K means k-means always separates my line it just draw a line like this you can see that but it is not a drawing a circle to separate so that is the drawback of a k links that's why uh We've for DB scan why DB scan is able to separate these two clusters like this this is called non-linear separation whereas K means fails that is why in our in our topic modeling uh in our topic modeling we are using HTTP that's why Martin got a first event with hdb scan High dimensional DB scan because uh this is a non-linear uh clustering method okay I hope this makes you some kind of fear okay so k-means is only works for linear when your data is already linear right that's how you can easily separate right so wherever you can see when your data see came in works good in this when you can separate like nicely like this but it fails here it fails here it even fails here as well only so but it it is very perfect here right it's very perfect here so that is why caymances uh linear uh separate clustering method uh again coming back to I'm starting um please do you have examples of notebook I'll try to share this so those are with my uh with whatsoever I asked him to share it but uh see it doesn't matter you have the package available publicly for free and there are empty number of articles on it so you can get the code for free and everywhere A lot of people have done a lot of analysis on using this package that can be a very good starting point for you I'm star okay I'm starting out of an angle we should get hold over classical please record his classical is quite required uh yeah 1.5 years you can learn your classic NLP as well how to label data especially for turning the model it's a very good question so one of the uh topic I am actually first streaming is uh weak supervision so weak supervision uh please learn about weak supervision um if you have wanted to know uh about uh right snorkel is a very good package I we did some research and I'm planning to publish a couple of papers on a week supervision so this is a place where you need to learn about if you want to label your topics okay so this is something called weak supervision please learn about weak supervision uh if you want to learn about labeling your topics I hope I answered uh whatever you mapped did you implement have we used a packet API uh this is my answer to Marcella underscore uh how to label data as oxide how to label data for time that's onset uh do you combine segment analysis of topics models we didn't do it here but I did it in a separate analysis right can you please throw some light on how we can do context based sentiment analysis so that is a bit out of topic here probably uh I will do one more session on context space sequence analysis um there is there is one context based uh statement uh analysis I did um let me check probably I think you can check it here I have one more session uh looking for the session I did or I can post it when I get it I'm not able to see it but I would uh I can share that to you later yeah labeling is the biggest challenge I told you right so Simi supervision uh look into some of these methods image supervision right this is uh introduction to semi supervision for data labeling you should look into these research methods pseudo labeling semi supervision there's a very good block from analytics video itself so this is one you can check it um and the second thing is as I told you weak supervision right so this you can check it this is from Stanford a research recently they established a new company called snorkel so this is where you can learn from uh I hope I'm posting this in the chat itself it's because it's easy for me but I hope everyone can take these questions as a masters data handling students with lots of college products but no practical data science experience uh where do you start in the industry um this is a very subjective question so um I would say would you start in the end how do you start in this I would take it there so the keep your final see this is you always need to be a tea based expert so have some experience on LLP your computer science knowledge how to do basic data visualization data analysis so I would say um so even before you guys you jump into NLP or computer vision my experience is try to get good at Tableau Data modeling right earlier people used to get recruited for having data science knowledge for uh if you build a model building a model is very easy these days uh building your model is very easy these days but you should um uh what you should know is how to productionize that model how to put it into an API these are the skills that everyone is looking for you okay and how to analyze how to so modern training is easy to test you need to retrain the model uh that is more important right so how to retrain the model when to retrain the model how is your model how are you monitoring the model metrics are becoming these days and did you also I think I think all the questions has been answered from my level right yeah yeah any any other questions thanks I hope um this is a introductory talk for you to get excited about some of the current uh um some of the current recent advancements in NLP space uh in our supervised level well thank you very much for joining this session I hope you had a great um evening uh nice talking to you guys uh at a very interesting questions but uh keep learning uh that's the only way you can get thank you very much thank you thanks a lot

Original Description

Customers' reviews and comments are important for businesses to understand users' sentiment about the products and services. However, this data needs to be analyzed to assess the sentiment associated with topics/aspects to provide efficient customer assistance. LDA and LSA fail to capture the semantic relationship and are not specific to any domain. BERTopic, is a novel method that generates topics using sentence embeddings and is applied to Consumer Financial Protection Bureau (CFPB) data.  In this DataHour, Bharath will show how BERTopic is flexible and yet provides meaningful and diverse topics compared to LDA and LSA. Furthermore, he will explain how domain-specific pre-trained embeddings (FinBERT) yield even better topics. For more amazing datahour session, visit: https://datahack.analyticsvidhya.com/contest/all/ Stay on top of your industry by interacting with us on our social channels: Follow us on Instagram: https://www.instagram.com/analytics_vidhya/ Like us on Facebook: https://www.facebook.com/AnalyticsVidhya/ Follow us on Twitter: https://twitter.com/AnalyticsVidhya Follow us on LinkedIn:https://www.linkedin.com/company/analytics-vidhya
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Analytics Vidhya · Analytics Vidhya · 0 of 60

← Previous Next →
1 The DataHour: Data Science in Retail
The DataHour: Data Science in Retail
Analytics Vidhya
2 The DataHour: Anomaly detection using NLP and Predictive Modeling
The DataHour: Anomaly detection using NLP and Predictive Modeling
Analytics Vidhya
3 The DataHour: Energy Data Science Project from Scratch
The DataHour: Energy Data Science Project from Scratch
Analytics Vidhya
4 The DataHour: Explainable AI Need and Implementation
The DataHour: Explainable AI Need and Implementation
Analytics Vidhya
5 The DataHour: Google Cloud AI/ML
The DataHour: Google Cloud AI/ML
Analytics Vidhya
6 Prediction to Production in Machine Learning #machinelearning #prediction
Prediction to Production in Machine Learning #machinelearning #prediction
Analytics Vidhya
7 Practical Applications of Data science in Ecommerce
Practical Applications of Data science in Ecommerce
Analytics Vidhya
8 How to tackle Overfitting?#machinelearning #overfitting
How to tackle Overfitting?#machinelearning #overfitting
Analytics Vidhya
9 Building Data Pipelines on GCP #googlecloud #datapipelines #data
Building Data Pipelines on GCP #googlecloud #datapipelines #data
Analytics Vidhya
10 Hands-on with A/B Testing #abtesting #datascience
Hands-on with A/B Testing #abtesting #datascience
Analytics Vidhya
11 Efficient Implementations of Transformers #transformers #cnn  #machinelearning
Efficient Implementations of Transformers #transformers #cnn #machinelearning
Analytics Vidhya
12 Modern Deep Learning Architecture #deeplearning  #architecture #deeplearningtutorial
Modern Deep Learning Architecture #deeplearning #architecture #deeplearningtutorial
Analytics Vidhya
13 Key steps for Designing Artificial Neural Network (ANN) for Image classification #machinelearning
Key steps for Designing Artificial Neural Network (ANN) for Image classification #machinelearning
Analytics Vidhya
14 5 things you should know about Azure SQL #azure #sql #datahour #datascience
5 things you should know about Azure SQL #azure #sql #datahour #datascience
Analytics Vidhya
15 AI & ML in the Automotive Industry #machinelearning #ai
AI & ML in the Automotive Industry #machinelearning #ai
Analytics Vidhya
16 Building Machine Learning Models in BigQuery
Building Machine Learning Models in BigQuery
Analytics Vidhya
17 NLP aspects in Telecommunication Industry
NLP aspects in Telecommunication Industry
Analytics Vidhya
18 Practical Time Series Analysis
Practical Time Series Analysis
Analytics Vidhya
19 Fundamentals of Quantum Computing
Fundamentals of Quantum Computing
Analytics Vidhya
20 A DAY IN THE LIFE of a Data Scientist (From waking up to working on algorithms)
A DAY IN THE LIFE of a Data Scientist (From waking up to working on algorithms)
Analytics Vidhya
21 Classification Machine Learning Model from Scratch
Classification Machine Learning Model from Scratch
Analytics Vidhya
22 Knowledge Graph Solutions using Neo4j
Knowledge Graph Solutions using Neo4j
Analytics Vidhya
23 Model Guesstimation (MLOps)
Model Guesstimation (MLOps)
Analytics Vidhya
24 ETL Pipelines in Google Cloud Platform
ETL Pipelines in Google Cloud Platform
Analytics Vidhya
25 Key steps for Designing Convolutional Neural Network(CNN) for Image Classification
Key steps for Designing Convolutional Neural Network(CNN) for Image Classification
Analytics Vidhya
26 Getting Started with AWS EC2 #amazon #aws
Getting Started with AWS EC2 #amazon #aws
Analytics Vidhya
27 How to Use Azure NLP and Graph Databases for Intelligent Knowledge Mining
How to Use Azure NLP and Graph Databases for Intelligent Knowledge Mining
Analytics Vidhya
28 Certified AI & ML BlackBelt Plus Program #shorts
Certified AI & ML BlackBelt Plus Program #shorts
Analytics Vidhya
29 Visualizing Data using Python #machinelearning #visualization #python
Visualizing Data using Python #machinelearning #visualization #python
Analytics Vidhya
30 DCNN for Machine RUL Prediction using Time-series Data #timeseries #machinelearning #datascience
DCNN for Machine RUL Prediction using Time-series Data #timeseries #machinelearning #datascience
Analytics Vidhya
31 M in ML stands for Math & Magic
M in ML stands for Math & Magic
Analytics Vidhya
32 An Unsupervised ML approach using Clustering
An Unsupervised ML approach using Clustering
Analytics Vidhya
33 Customizing Large Language Models GPT3 for Real-life Use Cases #gpt3 #datascience
Customizing Large Language Models GPT3 for Real-life Use Cases #gpt3 #datascience
Analytics Vidhya
34 Model Parameters vs Hyperparameters - Techniques in ML Engineering #machinelearning
Model Parameters vs Hyperparameters - Techniques in ML Engineering #machinelearning
Analytics Vidhya
35 Practical MLOps #mlops #datascience
Practical MLOps #mlops #datascience
Analytics Vidhya
36 Data Engineering with Databricks #dataengineering #databricks
Data Engineering with Databricks #dataengineering #databricks
Analytics Vidhya
37 Multi-Objective Optimisation
Multi-Objective Optimisation
Analytics Vidhya
38 When Airflow Meets Kubernetes
When Airflow Meets Kubernetes
Analytics Vidhya
39 AI in Banking
AI in Banking
Analytics Vidhya
40 Learn Convolutional Neural Network for Image Recognition
Learn Convolutional Neural Network for Image Recognition
Analytics Vidhya
41 Extracting Value from Data
Extracting Value from Data
Analytics Vidhya
42 How to measure Marketing Channel Effectiveness
How to measure Marketing Channel Effectiveness
Analytics Vidhya
43 Transforming Lives | Data Science Immersive Bootcamp
Transforming Lives | Data Science Immersive Bootcamp
Analytics Vidhya
44 Stock Market Analysis - AI driven approach
Stock Market Analysis - AI driven approach
Analytics Vidhya
45 Become a Data Engineering Professional in 2022 | Future Trends + Skills Required
Become a Data Engineering Professional in 2022 | Future Trends + Skills Required
Analytics Vidhya
46 Ensemble Techniques in Machine Learning #machinelearning #ensemble #datascience
Ensemble Techniques in Machine Learning #machinelearning #ensemble #datascience
Analytics Vidhya
47 The Power of Visualization | Tableau Full Course | Analytics Vidhya
The Power of Visualization | Tableau Full Course | Analytics Vidhya
Analytics Vidhya
48 Demand for Data Engineers is on the Rise | Data Engineer | Analytics Vidhya
Demand for Data Engineers is on the Rise | Data Engineer | Analytics Vidhya
Analytics Vidhya
49 Data Visualization in Data Science | DataHour | Analytics Vidhya
Data Visualization in Data Science | DataHour | Analytics Vidhya
Analytics Vidhya
50 Role of Optimization in Machine Learning & Deep Learning | DataHour | Analytics Vidhya
Role of Optimization in Machine Learning & Deep Learning | DataHour | Analytics Vidhya
Analytics Vidhya
51 Solving any Machine Learning Problem | Approach and Steps Involved
Solving any Machine Learning Problem | Approach and Steps Involved
Analytics Vidhya
52 Topic Modeling Explained with Implementation | Using LDA in Python | DataHour by Arpendu Ganguly
Topic Modeling Explained with Implementation | Using LDA in Python | DataHour by Arpendu Ganguly
Analytics Vidhya
53 Data Engineering in E-Commerce | The Best Case Study
Data Engineering in E-Commerce | The Best Case Study
Analytics Vidhya
54 Introduction to Classification using Azure Machine Learning | DataHour | Analytics Vidhya
Introduction to Classification using Azure Machine Learning | DataHour | Analytics Vidhya
Analytics Vidhya
55 Introduction to Federated Learning | DataHour | Analytics Vidhya
Introduction to Federated Learning | DataHour | Analytics Vidhya
Analytics Vidhya
56 Diffusion Models for Generative Arts | DataHour | Analytics Vidhya
Diffusion Models for Generative Arts | DataHour | Analytics Vidhya
Analytics Vidhya
57 Master Google Analytics in 1 Hour | DataHour | Analytics Vidhya
Master Google Analytics in 1 Hour | DataHour | Analytics Vidhya
Analytics Vidhya
58 Learn Hypothesis Testing | DataHour | Analytics Vidhya
Learn Hypothesis Testing | DataHour | Analytics Vidhya
Analytics Vidhya
59 A Practical Approach to Kaggle Competition | DataHour | Analytics Vidhya
A Practical Approach to Kaggle Competition | DataHour | Analytics Vidhya
Analytics Vidhya
60 Making AI work for Business | DataHour | Analytics Vidhya
Making AI work for Business | DataHour | Analytics Vidhya
Analytics Vidhya

Related Reads

📰
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics
Learn to build production-grade LLM evaluation pipelines to catch hallucinations before deployment and improve model reliability
Dev.to AI
📰
Why Every AI Engineer Should Learn Hugging Face
Learn how Hugging Face simplifies AI development and why it's a crucial tool for AI engineers to master
Medium · Machine Learning
📰
A bug in Qwen3-TTS taught me voice is biometric
A developer's experience with a bug in a voice cloning model highlights the biometric nature of voice, emphasizing security and privacy concerns
Dev.to · Daniel Nwaneri
📰
What is LoRA and how it lets anyone fine-tune a massive AI model on a single GPU
Learn about LoRA, a technique that enables fine-tuning of massive AI models on a single GPU, making it accessible to individuals and small teams
Medium · LLM
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →