BSO: Safety Alignment Is Density Ratio Matching
📰 ArXiv cs.AI
arXiv:2605.12339v1 Announce Type: cross Abstract: Aligning language models for both helpfulness and safety typically requires complex pipelines-separate reward and cost models, online reinforcement learning, and primal-dual updates. Recent direct preference optimization approaches simplify training but incorporate safety through ad-hoc modifications such as multi-stage procedures or heuristic margin terms, lacking a principled derivation. We show that the likelihood ratio of the optimal safe pol
DeepCamp AI