When Can LLMs Learn to Reason with Weak Supervision?
📰 ArXiv cs.AI
arXiv:2604.18574v1 Announce Type: cross Abstract: Large language models have achieved significant reasoning improvements through reinforcement learning with verifiable rewards (RLVR). Yet as model capabilities grow, constructing high-quality reward signals becomes increasingly difficult, making it essential to understand when RLVR can succeed under weaker forms of supervision. We conduct a systematic empirical study across diverse model families and reasoning domains under three weak supervision
DeepCamp AI