OPSDL: On-Policy Self-Distillation for Long-Context Language Models
📰 ArXiv cs.AI
arXiv:2604.17535v1 Announce Type: cross Abstract: Extending the effective context length of large language models (LLMs) remains a central challenge for real-world applications. While recent post-training methods have made progress in long-context scaling, they either rely on high-quality supervision data or sparse sequence-level rewards, leading to unstable and inefficient optimization. We propose OPSDL, an On-Policy Self-Distillation method for enhancing the Long-context capabilities of LLMs.
DeepCamp AI