用梯度优化方法解决多目标偏好混合模型的识别难题。
MoPLEx: Estimating Plackett-Luce Mixture Models for Multi-Objective Alignment
- 通过扩充排名数据并梯度优化,提升混合模型估计效率。
- 在340亿参数模型上误差低于5%,聚类与排序准确率提升超40%和15%。
- 适合处理复杂、异质性偏好场景,如AI对齐与多目标优化。
我们研究从标注者提供的多路排序响应中学习由k个普拉克特-卢塞模型组成的混合模型,以捕捉潜在的异质偏好。该问题在人工智能对齐与偏好优化中有广泛应用。以往工作主要针对成对比较下的布拉德利-特里模型混合。然而,当混合数k超过排序长度m的一半时,多路排序模型的混合估计可能变得不可识别。为此,我们设计了一种高效算法:先将排序数据扩展至更大规模(例如通过基础模型生成新比较),再在输入嵌入空间中采用基于梯度的估计来降低推理开销。在此基础上,我们通过类似期望最大化(EM)的迭代流程拟合混合普拉克特-卢塞模型,简称MoPLEx。大量实验验证了该方法的有效性:首先,基于梯度的近似在高达340亿参数的模型上误差低于5%;其次,在UltraFeedback和PERSONA数据集上,相较于单一普拉克特-卢塞模型或布拉德利-特里模型混合,MoPLEx平均提升了43.7%的聚类准确率和15.2%的排序准确率。这些结果表明,通过梯度测量对齐,MoPLEx能有效应对遵循异质偏好的多路排序任务。
原文摘要 · Abstract (English)
We study learning a mixture of $k$ Plackett-Luce models from multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. Prior work has studied mixtures of Bradley-Terry models from pairwise comparisons. However, estimating a mixture of multi-way ranking models can become theoretically unidentifiable when $k$ exceeds $m/2$, where $m$ is the ranking length. We design an efficient algorithm to address this issue by first augmenting the rankings to a larger size (e.g., generating comparisons from a base model), followed by a gradient-based estimation to reduce inference cost (in the input embedding space). With this procedure in mind, we then fit a mixture of Plackett-Luce (PL) models via an expectation-maximization-style iteration, or MoPLEx in short. We conduct extensive experiments to verify this algorithm. First, we find that the gradient-based approximation estimates true probabilities with less than 5% error on models with up to 34 billion parameters. Second, MoPLEx improves clustering and ranking accuracy by an average of 43.7% and 15.2% over baselines using a single PL model or a mixture of Bradley-Terry models, on UltraFeedback and PERSONA datasets. These results demonstrate the effectiveness of MoPLEx for tackling multi-way rankings following heterogeneous preferences through measuring alignment via gradients.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。