arXiv:2605.12339cs.LGcs.AI2026-05被引 1

用密度比匹配统一安全对齐,一步到位不依赖额外模型。

BSO: Safety Alignment Is Density Ratio Matching

论文配图:BSO: Safety Alignment Is Density Ratio Matching
图 1 · 摘自论文原文
  • 将安全对齐转化为密度比匹配问题,构造单阶段损失函数。
  • 在多个基准上提升安全与帮助性的平衡,超越现有方法。
  • 无需辅助模型,仅增一个超参数,兼容已有安全方法。

语言模型的安全性与帮助性对齐通常需要复杂的流程——包括独立的奖励与成本模型、在线强化学习和原-对偶更新。近期的直接偏好优化方法虽简化了训练,但通过多阶段过程或启发式边界项引入安全性,缺乏理论推导。我们发现最优安全策略的似然比可解析分解,使安全对齐退化为密度比匹配问题。最小化数据与模型比例间的Bregman散度,得到一类由凸生成器诱导的单阶段损失函数,即Bregman安全优化(BSO),可保证恢复最优安全策略。BSO兼具通用性与简洁性:无需辅助模型,仅增加一个超参数,且能还原现有安全感知方法作为特例。跨多个安全对齐基准的实验表明,BSO持续改善安全与帮助性的权衡。

原文摘要 · Abstract (English)

Aligning language models for both helpfulness and safety typically requires complex pipelines-separate reward and cost models, online reinforcement learning, and primal-dual updates. Recent direct preference optimization approaches simplify training but incorporate safety through ad-hoc modifications such as multi-stage procedures or heuristic margin terms, lacking a principled derivation. We show that the likelihood ratio of the optimal safe policy admits a closed-form decomposition that reduces safety alignment to a density ratio matching problem. Minimizing Bregman divergences between the data and model ratios yields Bregman Safety Optimization (BSO), a family of single-stage loss functions, each induced by a convex generator, that provably recover the optimal safe policy. BSO is both general and simple: it requires no auxiliary models, introduces only one hyperparameter beyond standard preference optimization, and recovers existing safety-aware methods as special cases. Experiments across safety alignment benchmarks show that BSO consistently improves the safety-helpfulness trade-off.

安全对齐偏好优化密度比匹配单阶段

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。