用质量差距优化大模型对齐,提升反馈精度与稳定性。
Margin Matching Preference Optimization: Enhanced Model Alignment with Granular Feedback
- 基于布拉德利-特瑞模型设计软目标概率,融入偏好质量差距。
- 在MT-bench和RewardBench上显著超越基线,7B模型达2024年6月最优。
- 对过拟合更鲁棒,模型校准性更好,适合高精度对齐场景。
大型语言模型(LLMs)通过强化学习从人类反馈中微调,已推动当前最先进AI系统的发展。然而现有方法多依赖二值偏好标签,难以捕捉成对输出间细微的质量差异。为此,本文提出边距匹配偏好优化(MMPO),将成对偏好中的质量边距纳入优化过程,提升模型策略与奖励模型性能。具体地,基于质量边距构建符合布拉德利-特瑞模型的软目标概率,并采用标准交叉熵训练。在人类与AI反馈数据上的实验表明,MMPO在主流基准如MT-bench和RewardBench上持续优于基线方法,效果显著;其中,7B规模模型在2024年6月的RewardBench测试中达到同规模最佳表现。分析还显示,该方法更具抗过拟合能力,模型校准性更优。
原文摘要 · Abstract (English)
Large language models (LLMs) fine-tuned with alignment techniques, such as reinforcement learning from human feedback, have been instrumental in developing some of the most capable AI systems to date. Despite their success, existing methods typically rely on simple binary labels, such as those indicating preferred outputs in pairwise preferences, which fail to capture the subtle differences in relative quality between pairs. To address this limitation, we introduce an approach called Margin Matching Preference Optimization (MMPO), which incorporates relative quality margins into optimization, leading to improved LLM policies and reward models. Specifically, given quality margins in pairwise preferences, we design soft target probabilities based on the Bradley-Terry model, which are then used to train models with the standard cross-entropy objective. Experiments with both human and AI feedback data demonstrate that MMPO consistently outperforms baseline methods, often by a substantial margin, on popular benchmarks including MT-bench and RewardBench. Notably, the 7B model trained with MMPO achieves state-of-the-art performance on RewardBench as of June 2024, outperforming other models of the same scale. Our analysis also shows that MMPO is more robust to overfitting, leading to better-calibrated models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。