OPSDL: On-Policy Self-Distillation for Long-Context Language Models

📰 ArXiv cs.AI

arXiv:2604.17535v1 Announce Type: cross Abstract: Extending the effective context length of large language models (LLMs) remains a central challenge for real-world applications. While recent post-training methods have made progress in long-context scaling, they either rely on high-quality supervision data or sparse sequence-level rewards, leading to unstable and inefficient optimization. We propose OPSDL, an On-Policy Self-Distillation method for enhancing the Long-context capabilities of LLMs.

Published 21 Apr 2026
Read full paper → ← Back to Reads