arXiv:2410.04346cs.CL2024-10EMNLP被引 7

用排序指标提升大模型对齐效果,让好回答排得更靠前。

Permutative Preference Alignment from Listwise Ranking of Human Judgments

  • 基于NDCG设计可微分的列表级对齐算法
  • 在AlpacaEval等评测中超越传统方法
  • 适合追求响应质量排序的研究者

大语言模型与人类偏好对齐对于确保模型行为理想且可控至关重要。现有方法如基于人类反馈的强化学习(RLHF)和直接偏好优化(DPO)依赖于布拉德利-特瑞模型(B-T)来最大化成对选择的似然性,但当存在多个响应时,B-T模型无法保证响应的准确列表排序。为此,我们提出一种新的离线列表级方法——置换偏好对齐(PPA),将广泛使用的排序指标归一化折损累积增益(NDCG)作为替代训练目标。通过近似NDCG构造可微分代理损失,实现端到端对齐。实验表明,PPA在评估集和通用基准(如AlpacaEval)上均优于现有的成对与列表级方法。此外,我们证明基于NDCG的方法比基于B-T的方法更能有效提升排序准确性,并提供了理论解释。

原文摘要 · Abstract (English)

Aligning Large Language Models (LLMs) with human preferences is crucial in ensuring desirable and controllable model behaviors. Current methods, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on the Bradley-Terry (B-T) model to maximize the likelihood of pairwise choices. However, when multiple responses are available, the B-T model fails to guarantee an accurate list ranking of the responses. To address this issue, we propose Permutative Preference Alignment (PPA), a novel offline listwise approach that incorporates the Normalized Discounted Cumulative Gain (NDCG), a widely-used ranking metric, as an alternative training objective for LLM alignment. We develop an end-to-end alignment algorithm by approximating NDCG with a differentiable surrogate loss. Experiments demonstrate that PPA outperforms existing pairwise and listwise methods on evaluation sets and general benchmarks such as AlpacaEval. Furthermore, we show that NDCG-based approaches improve ranking accuracy more effectively than B-T-based methods and provide a theoretical explanation for this improvement.

大模型对齐排序优化NDCG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。