CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models

📰 ArXiv cs.AI

arXiv:2602.17684v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based feedback from unit tests, but its scalability is fundamentally constrained by the availability and reliability of high-quality test cases. We propose CodeScaler, a reward model designed to scale both reinforcement learning training and test-time inference for code generation. CodeScaler is trained on

Published 19 May 2026
Read full paper → ← Back to Reads