arXiv:2505.18720cs.CLcs.AI2025-05ACL被引 7

用最优传输动态加权关键词,让模型更懂人类偏好

Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization

  • 基于最优传输思想,自动识别并加权语义重要词对
  • 在多个数据集上提升指令遵循能力,最高增益达12.3%
  • 适合关注模型对齐与生成质量优化的研究者

直接偏好优化(DPO)通过直接优化优选与次优响应之间的对数似然差来对齐大语言模型与人类偏好。然而,现有方法对所有词元赋予相同权重,而人类更关注有意义的部分,导致无关或噪声词元过度影响DPO损失,造成次优优化。为此,我们提出基于最优传输的词元加权方案(OTPO),通过强化语义相关词元对、弱化不相关部分,构建上下文感知的自适应加权机制,从而获得更具区分性的奖励差异估计。该方法提升了奖励稳定性,增强了可解释性,并确保优化聚焦于响应间的有意义差异。大量实验验证了OTPO在多种设置下对指令遵循能力的显著提升。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) has emerged as a promising framework for aligning Large Language Models (LLMs) with human preferences by directly optimizing the log-likelihood difference between chosen and rejected responses. However, existing methods assign equal importance to all tokens in the response, while humans focus on more meaningful parts. This leads to suboptimal preference optimization, as irrelevant or noisy tokens disproportionately influence DPO loss. To address this limitation, we propose \textbf{O}ptimal \textbf{T}ransport-based token weighting scheme for enhancing direct \textbf{P}reference \textbf{O}ptimization (OTPO). By emphasizing semantically meaningful token pairs and de-emphasizing less relevant ones, our method introduces a context-aware token weighting scheme that yields a more contrastive reward difference estimate. This adaptive weighting enhances reward stability, improves interpretability, and ensures that preference optimization focuses on meaningful differences between responses. Extensive experiments have validated OTPO's effectiveness in improving instruction-following ability across various settings\footnote{Code is available at https://github.com/Mimasss2/OTPO.}.

偏好优化词元加权最优传输大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。