Paper Highlights: Grokking Structure with Transformers
Reading Structural Grokking in Vanilla Transformers by Hoogland et al.
This paper challenges the concept that your validation accuracy is what determines when you should stop training your transformer models.
https://arxiv.org/abs/2305.18741
Watch on YouTube ↗
(saves to browser)
DeepCamp AI