提出新方法提升偏好优化中语义模糊的处理能力,显著改善模型对齐效果。
Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
- 基于偏好对中的语义相似内容自动重加权,减少训练时的歧义性。
- 在AlpacaEval 2上比DPO最高提升8.9分,Arena-Hard上最高提升15.0分。
- 无需增加生成长度,适配多规模模型与主流评测集,实用性强。
直接偏好优化(DPO)是广泛应用于多个领域的从人类反馈中进行强化学习(RLHF)的方法。近期研究关注标记重要性对提升DPO效果的作用。观察发现,偏好对中频繁出现相同或语义相似的内容(称为模糊内容),我们假设这些模糊内容在训练过程中可能引入歧义,从而限制对齐性能的进一步提升。通过数学分析与概念验证实验,我们揭示模糊内容可能引发歧义,进而降低性能。为此,我们提出模糊感知优化(AAO),一种简单而有效的方法,通过计算偏好对中的语义相似度,自动重加权模糊内容以减少歧义。大量实验证明,AAO在多个模型规模和主流基准数据集(包括AlpacaEval 2、MT-Bench、Arena-Hard)上持续且显著优于现有先进方法,且未明显增加响应长度。具体而言,AAO在AlpacaEval 2上相比DPO最高提升8.9分,在Arena-Hard上最高提升15.0分。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) is a widely used reinforcement learning from human feedback (RLHF) method across various domains. Recent research has increasingly focused on the role of token importance in improving DPO effectiveness. It is observed that identical or semantically similar content (defined as ambiguous content) frequently appears within the preference pairs. We hypothesize that the presence of ambiguous content during DPO training may introduce ambiguity, thereby limiting further improvements in alignment. Through mathematical analysis and proof-of-concept experiments, we reveal that ambiguous content may potentially introduce ambiguities, thereby degrading performance. To address this issue, we introduce Ambiguity Awareness Optimization (AAO), a simple yet effective approach that automatically re-weights ambiguous content to reduce ambiguities by calculating semantic similarity from preference pairs. Through extensive experiments, we demonstrate that AAO consistently and significantly surpasses state-of-the-art approaches in performance, without markedly increasing response length, across multiple model scales and widely adopted benchmark datasets, including AlpacaEval 2, MT-Bench, and Arena-Hard. Specifically, AAO outperforms DPO by up to 8.9 points on AlpacaEval 2 and achieves an improvement of by up to 15.0 points on Arena-Hard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。