Selective Off-Policy Reference Tuning with Plan Guidance

📰 ArXiv cs.AI

arXiv:2605.11505v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning. This turns all-wrong prompts into se

Published 13 May 2026
Read full paper → ← Back to Reads