Pytorch Embedding Model Part 3
Key Takeaways
PyTorch embedding model training, measurement, and visualization, utilizing token IDs to access embedding layers and storing dictionaries for efficient data retrieval
Full Transcript
Now, we're going to get back to our embedding model. Today, I want to do a lot of things with it. I've got a lot I've got a plan. Got a plan at everyone. Oh, we did it again. Hey, how's it going there? Welcome on in. Jack Shiffer. What do you Good to see you. Once Oh, Torva, thank you. Torva, you got the link? You got the link. You did it before I did. Thank you, Torva. Hey, how's it going? Welcome on in. Happy Wednesday. Good to see you. We're going to continue our embedding model. We got most of it done yesterday, I think. What we've got is the nice foundation. We've got a really good foundation that is going to allow us to start training. We're going to try to train it. And I think we're going to We also need to measure it. And we need to do We need to visualize it. There's a lot of stuff to do. There's a lot of stuff to do. Wouldn't it be nice to have the AI just do that for us? >> [laughter] >> At this point, maybe to write the code for us a little bit. I think that would be really All right, what do we need to do here? So, we did the norm layer already dividing by the number of words. Linear done. Linear layer without bias, we did that. Do we need this? Do we need these things is the next question. I think the answer is we [music] do. We were told that we don't need them. However, I'm pretty sure we might need them. I think just maybe. I think maybe. We're going to find out though. What's the number of words on the output, right? What we've got defined here is actually pretty straightforward cuz there's not a lot of code here. It's only 100 100 lines. Roughly 100 lines. We could walk through it right now. Let's do a full walk-through of what we've currently got. So, that way we can figure out where we want to go from here because there's a lot of tasks that we got to get used. You should use a location-based model. Woo, Jam Buster. Or a data set. That would be pretty neat. What I'm using is this data set right here. It's a little simple sentence, right? Set wrap. You can see a very small, very easy, simple sentence. That's an easy one right there. Super easy, super simple. So, let's go through our embedding model. Congratulations you guys though who got the memberships. Lil Wayne, hey, how's it going? Good to see you. Welcome on in. All right, so where are we at right now? So, we got our embedding model. We did Thank you to Dina. We need to do some of these things here. So, here's our embedding model. Very It's fairly easy. Here's right here. This is our embedding layer. Look at that. Look how simple that is right there. That's all it is. It's just a It's just a matrix. It's a two It's 2D matrix. You Let's do a quick walk-through, then we're going to watch the video. So, here's our embedding in it. We've got the number of words and our dictionary. I'm storing these in here. It's not traditional for an embedding model to do this. I'm doing this on purpose because we want to use a lot of the metadata that's around us. I need the length of the dictionary, number of words, and the ability to look up. Right? Cuz I'm going to tokenize. I want my feed forward to include pre-vectorized data. Right? It's going to allow you to post in just any words directly. So, you can pass [music] in this entire sentence. That's why I did that here in the feed in the feed forward, the forward pass. Then, we're going to do a look-up based on all the features are going to include the index of all the IDs that we are going to need to capture the output from our embedding. The embedding's going to return rows that are assigned to each of the indexes in our IDs. That's what that is. And then, we have a whole bunch of data manipulation code here that we needed in order to get it to operate properly. That's what we had to do. We'll walk through that here. We're going to We're teaching PyTorch right now with Python. And as you can see on the screen right here, there's not a lot of PyTorch. That's because the majority of the work when it comes to AI is data management and data preparation. [music] Happy coding. Thank you. I am Harry for mentioning it. What is a tokenizer? Oh, what tokenizer are you using? My very own. I'm using my own. It's all from scratch, you guys. So, the tokenizer is actually kind of see it right here. Really simple. So, we normalize our words, which you see right here, normalized right here. Which is a simple normalization. >> [music] >> And then we capture our norm'd words. Then we throw it into our dictionary, which we pre-calculate our dictionary just by building it based on all the data [music] set of the words. So, this is a dictionary calculator, which is just a Python dictionary. It's actually a dictionary with a word followed by the word index. Right? And what I should do here is I shouldn't use the word pad. I should throw some other word in here. I should say pad like this. It's got to be caps. It's got to be in caps. Found a little bug there. Nice little bug there. We still want to get to my embedding model here cuz we got to the point Now, if you're curious, we've got the full feed-forward pass working. Now, all all we have to do is leverage our optimizer and our uh loss function to back-propagate greater gradients and apply the optimal weight updates. Default tax rate value of 19% but you can get that refunded if it's not for business. 19% is a lot. Also, how do you get that refunded if you go to the grocery store like every day? That would be a lot of paperwork. You keep telling you, Stephen, embedding is trained alongside a large language model, not separately. Yeah, so I'm going to try I'm going to try That's why I think I've got these I left these here. Right? I left these here. So, this needs to be What is that? You 256. And then this will be 256. Uh maybe then 256 128 128 and the number words. 28 So, I'm going to try Uh we're going to we're going to we're It's okay. What's our shapes here? So, I need to compare the shapes is what [music] I need to do. Features, labels, shapes. I want the out shape and the label shape here. Let's hide everything else. And then we need to a do a break. [music] Just one. Just one. Just one, you guys. Let's do just one there. All right. So, we get 23 and 32, the label. So, since we're doing all the words, we're getting a batch. Oh. How are we going to deal [music] with that? So, we need to reshape We got to do a reshape here. Something's broken here. This should do the trick. What is So, hold on. Print len dictionary. It is 23. So, then why is this 32? Right here. Why is that 32? We need to figure that out right there. We got to figure out why is it 32? So, let me I want to do I need to get the loss working and that should be the least the one last thing that we do today. I need [music] to get the loss successful. And I think I can do it right now, basically. So, this will give us number words. Let's go ahead and run that on the output once. And the number of words is 23. The output size, 23. There we go. Let's double-check this here. So, out So, let's see. out.shape Let's see the out. And then we've got our label. We got our label shape. There we go. Then we'll get the dictionary length. So, that way I that This is going to make it a little bit easier to see what's actually what what we're printing on the screen, you know? We get to see what we're printing on the screen. Let's see here, you guys. Just log the prob classic cross entropy. Yes. That's the idea. Okay. Dictionary length, 23. Perfect. Label size, 23. There we go. Output size, 32. Why? What? Hold on. That seems broken to me. So, the the Did I Did I So, I said features, features, right? So, that's fine. That part's fine. We did We're We're good there. We're good there. So, since we're inputting that number features, my edits my edit after the app after effect, [music] I just need to fix this one last thing here. I just need to see I just need to I need to at least get the loss working. The output is 32. So, why is it doing that? Uh let's see how for Oh! You should say, "Steven, it's because you didn't write you didn't write it yet." You didn't write it. You didn't write it. Size, 32. So, we need that 32 to be a 23. But, why is it 32? Huh. It's just It's just weird that it's 32. It shouldn't be. Right? Why would it be 32? Cuz we're only doing the embedding. I want So, this is going to get three. It's going to be the exact vectors that I'm looking for, which is fine. So, that gives me the features from my vectors, which is perfect. Or is it? Yes. So, I my tokenization, right? Gets me all my words, which is going to be 1 2 3. And then I do a forward pass. See, I could probably do like maybe do a dot unsque- uns- uns- s q u e e z e. Or maybe maybe we do a squeeze s q [music] u e e z e. Let's try that. Uh that did not do anything. That did literally nothing. See one. Uh expected range between -1 positive 1. Okay. Dim equals Maybe I need to unsqueeze it. Oh, it's There we go. Didn't like it. Didn't like it. Okay. Okay. So, that's the wrong direction. Though, it did it changed some things. See -1. All right, here we go. Uh no. Well, no. Still didn't give me what I'm looking for. We need to get the right shape, you guys. Because our dictionary is going to be the answer. The dictionary is going to be the answer, not the number words. Do we even need this? We don't. We just need the dictionary. So, let's fix this. Let's see, where are we getting number of words? Let's build the dictionary, and we won't return Oh, we already are returning the We'll just return dictionary. We don't need number of anymore. [music] There we go. And we will only in the embedding [music] model request the dictionary. And then we can just do a length over this for the dictionary itself. There we go. So now we've got number words. Let's see here. All right. And then this should do the trick. So that's the input shape. So uh so maybe that's why we're seeing a problem right now. That's might be what the problem is. >> [music] >> Here is my edit but the intro lagging. Okay. Whoop. Whoa. Is there anything safe for work here? Here, I'm going to mute it. I'll I'll mute it. Hey, you're doing some video editing. Nice. Where's the invite though? Acts no, the invite shape of the feature that shape. See what that looks like. All right, here we go. Uh number words all right, number of words it's going to be up to self. And then there we go. That should be I believe it should be uh 23. Should be 23. It shouldn't be anything less, shouldn't be anything more, it should be 23. All right, build dictionary. All right, how about now? Okay. Shape 3 1. Okay. And then our output size is 32. Let me go back over here. Let's undo this [music] here real quick. I would like to try something a smidge different. Let's do Let's do Let's just do one at a time. We'll do one at a time here. See, uh let's see. Uh length. Uh see what do you call this here? Mhm. I I >> [music] >> length length l e n g t h We'll say one. Let me just update these guys real quick real real real real quick here. Okay, I think that's it. So, now I can now I can customize that a little bit better. Try that now. Size one one one one looks good. I pass it through and then it gives me [music] an embedding the size of 32 to 256. 32 Where is it getting 32 from? That's what That's what my question is. If
Original Description
Today we are getting back to our embedding model. We finished most of the foundation yesterday, and now we want to train it, measure it, and visualize it. The core idea is simple: an embedding layer is just a two dimensional matrix, and we use token IDs to pick the right rows.
I also store the dictionary inside the model so we can tokenize text on the fly and pass in full sentences. Most of the work so far has not been PyTorch code, it has been data prep and shaping tensors so everything lines up. I built a basic tokenizer from scratch that normalizes words, builds a Python dictionary of word to index, and handles padding.
Now the main goal is to get training working by wiring up the loss and optimizer, but we hit a shape mismatch: the dictionary has 23 words and the labels are size 23, yet the model output is size 32. Next step is to trace where that 32 comes from, confirm the final layer outputs vocab size, and reshape the batch correctly so cross entropy can run.
More on: ML Pipelines
View skill →Related Reads
📰
📰
📰
📰
Day 15/60: The Wisdom of the Crowd — Ensemble Learning & Random Forests
Medium · AI
Day 15/60: The Wisdom of the Crowd — Ensemble Learning & Random Forests
Medium · Machine Learning
Day 15/60: The Wisdom of the Crowd — Ensemble Learning & Random Forests
Medium · Data Science
Your AI Model Has a Retirement Date
Medium · AI
🎓
Tutor Explanation
DeepCamp AI