arXiv:2506.08998math.STcs.LG2025-06被引 1

发现并解释了AI对齐中偏好学习的反直觉现象

On Monotonicity in AI Alignment

  • 提出通用框架分析比较式偏好学习的单调性问题
  • 证明方法在局部配对下仍满足单调性约束
  • 提供评估工具,帮助设计更可信的对齐算法

比较式偏好学习已成为对齐人工智能模型与人类偏好的核心方法。然而,这些方法可能表现出反直觉行为:当模型偏好输出 y 胜过 z 时,其生成 y 的概率(及奖励)反而下降(此现象此前已有观察)。本文研究了此类(非)单调性的根源,针对涵盖 DPO、GPO 和 GBT 的通用比较式偏好学习框架,在弱假设下证明了方法仍满足局部配对单调性。同时,本文提出多种单调性形式化定义,并识别其保证的充分条件,构建起评估模型对单调性违反敏感度的工具箱。研究结果澄清了现有方法的局限,为开发更可信的偏好学习算法提供指导。

原文摘要 · Abstract (English)

Comparison-based preference learning has become central to the alignment of AI models with human preferences. However, these methods may behave counterintuitively. After empirically observing that, when accounting for a preference for response $y$ over $z$, the model may actually decrease the probability (and reward) of generating $y$ (an observation also made by others), this paper investigates the root causes of (non) monotonicity, for a general comparison-based preference learning framework that subsumes Direct Preference Optimization (DPO), Generalized Preference Optimization (GPO) and Generalized Bradley-Terry (GBT). Under mild assumptions, we prove that such methods still satisfy what we call local pairwise monotonicity. We also provide a bouquet of formalizations of monotonicity, and identify sufficient conditions for their guarantee, thereby providing a toolbox to evaluate how prone learning models are to monotonicity violations. These results clarify the limitations of current methods and provide guidance for developing more trustworthy preference learning algorithms.

AI对齐偏好学习单调性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。