Pytorch Embedding Model Part 2
Key Takeaways
Implements a word embedding model in PyTorch using tokenization and dictionary-based indexing
Original Description
I am pumped today because we finally have a solid plan for our word embedding model, and I want to build it in PyTorch. The goal is simple. Take words, tokenize them into dictionary indexes, pad the context, and turn inputs into vectors we can train on.
During training, we will predict the next word from the last three words, like found the bug leads to win, then the bug win leads to you. We hit a few snags, like a regex that did not clean punctuation and some tensor type errors, but we fixed them. We also replaced torch.nn.Embedding with our own embedding tensor and learned we can just index rows instead of doing full one hot matmul.
Next up is a small data iterator for batching, then cross entropy loss, then training, and finally testing word similarity. We committed the code and we will start training tomorrow.
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
More on: ML Pipelines
View skill →Related Reads
📰
📰
📰
📰
Why NIST Researchers Spent 10 Years Measuring Gravity
IEEE Spectrum
My AI Designs Rockets in 6 Simulations. A “Dumb” Baseline Nearly Kept Up.
Medium · Machine Learning
Why Trusted Defaults Become the New Scarcity
Medium · Machine Learning
Intelligence Lives in the Loop
Medium · Machine Learning
🎓
Tutor Explanation
DeepCamp AI