arXiv:2505.05465cs.CLcs.AI2025-05NeurIPS被引 17

用比较模型优化大模型偏好对齐,解决噪声数据问题。

ComPO: Preference Alignment via Comparison Oracles

  • 基于比较预言机的零阶优化,不依赖概率输出。
  • 在多个主流模型上提升对齐效果,尤其适配噪声偏好对齐。
  • 方法设计考虑了偏好对的差异性,适合实际训练场景。

直接对齐方法广泛用于将大语言模型与人类偏好对齐,但存在冗长和似然偏移问题,根源在于噪声偏好对导致优选与次选响应的似然值相近。本文提出一种新方法:基于零阶、基于比较的优化,通过比较预言机实现,并提供基本方案的收敛性保证。进一步引入启发式改进,实验验证其在使用噪声偏好对时的灵活性与兼容性。评估覆盖 Mistral-7B、Llama-3-8B、Gemma-2-9B 等多个基础与指令微调模型,采用 AlpacaEval 2、MT-Bench 与 Arena-Hard 基准。结果表明,该方法可有效克服现有直接对齐方法的局限性。研究还揭示了针对不同似然差距的偏好对设计专用方法的重要性,与 Razin 等(2025)近期发现形成互补。

原文摘要 · Abstract (English)

Direct alignment methods are increasingly used for aligning large language models (LLMs) with human preferences. However, these methods suffer from the issues of verbosity and likelihood displacement, which can be driven by the noisy preference pairs that induce similar likelihood for preferred and dispreferred responses. The contributions of this paper are two-fold. First, we propose a new preference alignment method based on zeroth-order, comparison-based optimization via comparison oracles and provide convergence guarantees for its basic scheme. Second, we improve our method using some heuristics and conduct the experiments to demonstrate the flexibility and compatibility of practical scheme in improving the performance of LLMs using noisy preference pairs. Evaluations are conducted across multiple base and instruction-tuned models (Mistral-7B, Llama-3-8B and Gemma-2-9B) with benchmarks (AlpacaEval 2, MT-Bench and Arena-Hard). Experimental results show the effectiveness of our method as an alternative to addressing the limitations of existing direct alignment methods. A highlight of our work is that we evidence the importance of designing specialized methods for preference pairs with distinct likelihood margin, which complements the recent findings in Razin et al (2025).

偏好对齐大模型训练比较优化噪声数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。