Text Pre Processing for LLM | Text Cleaning | Stemming | Lemmatization

RoboSathi ยท Beginner ยท๐Ÿง  Large Language Models ยท2mo ago

About this lesson

๐Ÿ“˜ Notes: https://robosathi.com/docs/natural_language_processing/text-pre-processing/ ๐ŸŽฅ NLP Playlist: https://www.youtube.com/playlist?list=PLnpa6KP2ZQxcDlHCeNiKbRhLWKVunQaxn ๐ŸŽฅ Evolution of LLMs: https://youtu.be/YLxYim_Kpbo โœ… This video describes the various pre-processing techniques in detail, viz., ๐Ÿ‘‰ Cleaning ๐Ÿ‘‰ Stemming ๐Ÿ‘‰ Lemmatization ๐Ÿ•” Time Stamp ๐Ÿ•˜ 00:00:00 - 00:00:28 Introduction 00:00:29 - 00:03:50 What is Text Pre-Processing ? 00:03:51 - 00:05:13 Text Cleaning 00:05:14 - 00:07:38 Stemming 00:07:39 - 00:08:58 Lemmatization 00:08:59 - 00:10:08 Stemming Vs Lemmatization - When to use what ? 00:10:09 - 00:10:27 Next: Tokenization

Full Transcript

[music] >> Hello and welcome. I hope you are having a great day. In this video, we'll understand what is text pre-processing in context of NLP or natural language processing. So, here we'll understand what do we do and why do we need to do the text pre-processing because the text is not so clean and all the words and commas, this, that, this is, off, and these kind of things should be removed. How it's done? Let's understand. So, what is text text pre-processing in NLP? So, the context is that a raw text is messy and inconsistent. The raw text that we have it's it has a lot of things that should not be there and it is not in a proper format also. It's inconsistent. So, it needs some cleaning. Okay. And this is this is because we need to feed it to the models, so deep learning models for them to learn. So, in the way they need it, in context of that it is messy. It it has a lot of things which is not needed and it is inconsistent. It is not in a consistent format. Okay. Because say for example, I am something something. This is also there and there is am I something something something. So, both of these are correct, right? So, doesn't we cannot find like there's some consistency in the language. Cuz there are so many forms that same things can be used. Because of that, so but the machines need some patterns, some consistent way that they can understand. So, how do we we transition from this inconsistent way to a consistent format, cleaning all the things that is not needed? That is the whole idea of text text pre-processing stage. Okay. So, there are two things that we do. We do the cleaning part. Remove all the punctuations, comma, like uh smiley. There's a smiley like that or there are other emojis that we use. All those things are irrelevant. Mhm? Then we lower case when the case is like name with a capital letter, etc. We just want to know what is noun, what is verb. That will work fine for the machine to understand. It doesn't need capital letter, small letter, etc. So, everything's converted to a lower case. Stop word removal. This is of. These kind of things which are always there and don't give any value or meaning to the sentence. They're just like punctuations and uh this this kind of things. These are stop words. No value. Not much value, basically. Then stripping of the special characters, special characters, emojis, etc. Everything should be removed. So, that's the idea of cleaning. Get rid of everything that is not language, not not proper word, that doesn't give me any like every sentence will have a subject, the doer, the predicate on which it is doing, like that, some verb that will happen. So, sentence has a certain format in the grammar. So, all these things are relevant. Rest everything is adjectives, adverbs. Those are things that are important, but rest of things we'll get rid of. Okay? So, that's the whole idea of cleaning. Then we do stemming and lemmatization. I'll explain this later. Okay. Reducing words to their root form. So, the words like run, ran, running, runner. Like this. So, all these are like same word. At at the root there's run only, but there's so many like few, fewer, fewest. New, newer, newest, like that. So, these words are like same, but only we want only the root word. Rest of things can be formed. And similar words can be formed by just putting those suffixes at the end. So, all these things all these the suffixes and what is happening at the end, all these things can be the end part can be truncated off and only the root can be kept. Okay? So, that's the whole idea of stemma stemming and lemmatization. Mhm? So, we'll see let's see how how it's done. So, first we'll do cleaning. So, what is So, what do we do? We remove punctuation, lower casing, convert into lower case, stop word removal, and stripping. So, let's see an example here and you'll understand. So, we have a smiley, hi, like that hello hello smiley we give and then together we will learn NLP or natural language processing within brackets and three exclamation marks. So, this is the sentence that we have. So, after cleaning, what will happen? I'll get rid of the smiley, gone, no read. This comma will be gone. This bracket will be gone. These exclamation marks, everything will be gone. Every everything will be made to small case. So, it'll become hello, learn. Okay. Hello. And stop words also will like together, we will, everything will go off. So, learn is important. Hello, learn NLP and natural language processing. So, together we will, this is not important. These are stop words. Hello is there hello. I'm just saying hello. So, these kind of thing these words will go off also. So, this is what we'll get after we have done lower casing. So, everything you see natural language processing I was I was doing NLP was in capital N L P. Everything Everything has been converted to smaller case. And just everything which is not needed, which is not relevant for it to make a meaning out of the sentence, that is gone. And now this is the cleaned up version. This is only relevant and this is what will be used for the further steps that will be fed finally into the machine learning model. So, this has been cleaned up. Now, next things. Stemming. So, there are two ways with the stemming, then we do lemmatization. Stemming means we just take the stem of the word out. So, this is like reduce words to their root form by chopping off the words at the end, suffixes, the ends. End will be the So, there's a word. So, this is the suffix. Say run ing. So, this is the root. This is the suffix. ing. Run er, runner. Like that. Run n e r. R. Okay. Just let me do once more. Run ing. Run ner. Like that. This is your root. This is your suffix. So, we just keep the root. So, sometimes the root may not be meaningful also. So, often resulting in non-dictionary root. So, we may not have like we just chop off the word at the end. So, something which are common common common, we'll just keep that and then we just just chop it off. Okay? And that may not be some meaningful word. It may be a non-dictionary word as roots. But this is very fast. We just chop off the suffix and it is very fast. You can just clean up to make it make it like a root word only. So, for example, let's see with example. Running was considered better than going to gym. And after So, there are different ways there are algorithms for stemming. So, there's a Porter stemmer, one of the algorithms. So, Porter stemmer is there. So, using this algorithm stemming algorithm, what we did what we got was So, instead of running we got run. Was was truncated to wa. Considered was only consid because consid we can consider considering. Okay? Consider, everything can be made from consider. We can make consider, considering, like that. Consideration. And the other words can also be made. So, it just chopped off to consid. Then better was better only. Than was than only. Going was converted to go. To was only to and gym is gym. So, you see here the consid and was and and then this these two have been and running was run. So, this is what it it it did. It just chopped off the word at certain things and just stemmed off. So, only the root only the root is there. This is gone. So, this is what is stemming. Similar thing we do in lemmatization, but in lemmatization, the root word is a dictionary word. It has some meaning. So, this lemma comes from the root this dictionary. So, this are this are some meaning. Okay. There's a meaning. So, meaning will be there. Reduce words to their root form that is dictionary base form, dictionary base form or lemma. Okay? So, input same thing running was considered better than going to gym. Same thing using another algorithm that is WordNet lemmatizer. So, there we saw Porter stemmer. Now there's a lemmatizer, different algorithm, WordNet. This our output will be running will be lemmatized to run. Was will go to its root form. Was will become be. Understand? Be, was, been. These are the three past past tense past participle. English revision. So, these are the forms of be. Was root form is be, like that. Considered will become consider, present tense here. Better will become good. Good, better, best. Similarly, than go to gym, like that. So, this will be all these root words will have some meaning and that will come from dictionary. Dictionary are based on that is a lemma. So, that's the difference between stemming and lemmatization. So, we use either stemming or lemmatization depending. The cleaning is common, but we use either stemming or lemmatization depending upon the use case. What are the use case? If you want accuracy, then we go for lemmatization. For example, chatbots. So, the words should be accurate. It should tell the complete thing in a under human understandable format. So, that time we'll go for lemmatization when accuracy is important. But when the speed is important, when we have a very large massive data set set and we want to do processing of that and that with that time we go for stemming. That okay, let's process this and the exact word may not be so important. I just want to get processed and get some some kind of similarity meaning of all the words. Okay? So, that in that case, in very massive data sets, we go for stemming when the speed is important because this is very fast. I don't care whether it's a root word or not. Just truncate suffix. Depending upon the algorithm. I hope this is clear the use case. So, these are the I we use either or, either lemmatization or stemming, not both. Cleaning is common, then we go for stemming or lemmatization depending upon the use case. Okay? So, this is your text pre-processing stage. Now, then if this is clear, we'll move to the next thing which is called tokenization. So, we'll convert these after this pre-processed text, we'll convert them into tokens or token numbers. We'll see why it's required, how it's done in the next video. So, that's all for this video. And bye for now.

Original Description

๐Ÿ“˜ Notes: https://robosathi.com/docs/natural_language_processing/text-pre-processing/ ๐ŸŽฅ NLP Playlist: https://www.youtube.com/playlist?list=PLnpa6KP2ZQxcDlHCeNiKbRhLWKVunQaxn ๐ŸŽฅ Evolution of LLMs: https://youtu.be/YLxYim_Kpbo โœ… This video describes the various pre-processing techniques in detail, viz., ๐Ÿ‘‰ Cleaning ๐Ÿ‘‰ Stemming ๐Ÿ‘‰ Lemmatization ๐Ÿ•” Time Stamp ๐Ÿ•˜ 00:00:00 - 00:00:28 Introduction 00:00:29 - 00:03:50 What is Text Pre-Processing ? 00:03:51 - 00:05:13 Text Cleaning 00:05:14 - 00:07:38 Stemming 00:07:39 - 00:08:58 Lemmatization 00:08:59 - 00:10:08 Stemming Vs Lemmatization - When to use what ? 00:10:09 - 00:10:27 Next: Tokenization
Watch on YouTube โ†— (saves to browser)
Sign in to unlock AI tutor explanation ยท โšก30

Related Reads

Chapters (7)

00:00:28 Introduction
0:29 00:03:50 What is Text Pre-Processing ?
3:51 00:05:13 Text Cleaning
5:14 00:07:38 Stemming
7:39 00:08:58 Lemmatization
8:59 00:10:08 Stemming Vs Lemmatization - When to use what ?
10:09 00:10:27 Next: Tokenization
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch โ†’