用胜率统一理解偏好学习,揭示方法优劣与优化关键
Preference learning made easy: Everything should be understood through win rate
- 以胜率为核心构建偏好学习理论框架
- 证明胜率优化方法具理论优势,现有方法多缺乏此特性
- 强调优化成功比目标设计更重要,指导实践改进
偏好学习(偏好比较数据对生成模型的对齐)尚未达到分类、密度估计等任务的理论成熟度。本文从成对偏好数据的采样分布出发,证明唯一同时尊重偏好与数据分布普遍性的评估方式是胜率,从而确立胜率作为理解偏好学习的核心。将现有方法分为胜率优化(WRO)与非胜率优化(non-WRO),提出多种新WRO实例,证明其具理论优势;而如DPO、SFT等常见非WRO方法则缺乏这些性质,并建议改进方向。研究还发现WRO因优化困难导致实际表现不佳,且优化成功程度比目标设计更能预测性能。分析指出现有方法最佳实践,并为未来研究提供指引:要么使非WRO方法更贴近WRO,要么提升WRO目标的优化能力。
原文摘要 · Abstract (English)
Preference learning, or the task of aligning generative models to preference comparison data, has yet to reach the conceptual maturity of classification, density estimation, etc. To close this gap, this work presents a framework to understand preference learning starting from the sampling distribution of pairwise preference data. First, we prove that the only evaluation of a generative model that respects both preferences and prevalences in the data distribution is a form of win rate, justifying win rate as the focal point to understand preference learning. We then analyze preference learning methods as win rate optimization (WRO) or non-WRO. We present novel instances of WRO beyond existing examples (RLHF, NLHF) and identify two key theoretical benefits of all such methods. We prove that common non-WRO methods like DPO and SFT on preferred samples lack these properties and suggest ways to mitigate such theoretical limitations. We also show that WRO underperforms in practice due optimization difficulties and that optimization success predicts performance better than choices which affect the objective's solution. Our analysis highlights best practices for existing methods and provides recommendations for future research, guided by the principle that one should either align non-WRO methods more closely with WRO or improve the optimization of WRO objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。