PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
📰 ArXiv cs.AI
arXiv:2605.10335v1 Announce Type: cross Abstract: Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of gradient first and second moments, incurring substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics. Motivated by steepest descent under an $\ell_p$-norm geometr
Full Article
Title: PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
Abstract:
arXiv:2605.10335v1 Announce Type: cross Abstract: Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of gradient first and second moments, incurring substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics. Motivated by steepest descent under an $\ell_p$-norm geometr
Abstract:
arXiv:2605.10335v1 Announce Type: cross Abstract: Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of gradient first and second moments, incurring substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics. Motivated by steepest descent under an $\ell_p$-norm geometr
DeepCamp AI