Sampling for Quality: Training-Free Reward-Guided LLM Decoding via Sequential Monte Carlo

📰 ArXiv cs.AI

arXiv:2604.16453v1 Announce Type: cross Abstract: We introduce a principled probabilistic framework for reward-guided decoding in large language models, addressing the limitations of standard decoding methods that optimize token-level likelihood rather than sequence-level quality. Our method defines a reward-augmented target distribution over complete sequences by combining model transition probabilities with prefix-dependent reward potentials. Importantly, the approach is training-free: it leav

Published 21 Apr 2026
Read full paper → ← Back to Reads