Stability and Generalization in Looped Transformers
📰 ArXiv cs.AI
arXiv:2604.15259v1 Announce Type: cross Abstract: Looped transformers promise test-time compute scaling by spending more iterations on harder problems, but it remains unclear which architectural choices let them extrapolate to harder problems at test time rather than memorize training-specific solutions. We introduce a fixed-point based framework for analyzing looped architectures along three axes of stability -- reachability, input-dependence, and geometry -- and use it to characterize when fix
DeepCamp AI