arXiv:2409.06691cs.LGcs.AI2024-09NeurIPS被引 21

用分布化软标签改进大模型对齐,避免过优化问题。

Geometric-Averaged Preference Optimization for Soft Preference Labels

  • 用加权几何平均调整损失函数,适配分布式偏好标签。
  • 在标准基准上表现更优,尤其当多数标签信心中等时提升显著。
  • 可无缝集成到现有DPO方法,适合偏好学习研究者。

许多大模型对齐算法假设人类偏好为二值且确定性,但实际偏好存在个体差异,应以分布形式表示。本文提出分布化软偏好标签,并在直接偏好优化(DPO)损失函数中引入模型输出似然的加权几何平均。该方法根据软标签动态调整损失尺度,当回应趋于同等受欢迎时,损失趋近于零。这一简单修改可应用于任意基于DPO的方法,缓解过优化与目标不匹配问题。实验通过大模型生成的AI反馈模拟软标签,结果表明几何平均在标准对齐基准上持续提升性能,尤其在多数标签为中等置信度时,生成响应更受偏好。

原文摘要 · Abstract (English)

Many algorithms for aligning LLMs with human preferences assume that human preferences are binary and deterministic. However, human preferences can vary across individuals, and therefore should be represented distributionally. In this work, we introduce the distributional soft preference labels and improve Direct Preference Optimization (DPO) with a weighted geometric average of the LLM output likelihood in the loss function. This approach adjusts the scale of learning loss based on the soft labels such that the loss would approach zero when the responses are closer to equally preferred. This simple modification can be easily applied to any DPO-based methods and mitigate over-optimization and objective mismatch, which prior works suffer from. Our experiments simulate the soft preference labels with AI feedback from LLMs and demonstrate that geometric averaging consistently improves performance on standard benchmarks for alignment research. In particular, we observe more preferable responses than binary labels and significant improvements where modestly-confident labels are in the majority.

大模型对齐偏好优化软标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。