The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
📰 ArXiv cs.AI
arXiv:2510.09378v2 Announce Type: replace-cross Abstract: Recent efforts to accelerate LLM pretraining have focused on computationally-efficient approximations that exploit second-order structure. This raises a key question for large-scale training: how much performance is forfeited by these approximations? To probe this question, we establish a practical upper bound on iteration complexity by applying full Gauss-Newton (GN) preconditioning to transformer models of up to 150M parameters. Our exp
DeepCamp AI