Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate

📰 ArXiv cs.AI

arXiv:2605.21486v1 Announce Type: cross Abstract: Hyperparameter transfer allows extrapolating optimal optimization hyperparameters from small to large scales, making it critical for training large language models (LLMs). This is done either by fitting a scaling law to the hyperparameters or by a judicious choice of parameterization, such as Maximal Update ($\mu$P), that renders optimal hyperparameters approximately scale invariant. In this paper, we first develop a framework to quantify hyperpa

Published 21 May 2026
Read full paper → ← Back to Reads