让AI同时尊重群体中不同意见,避免只听主流声音。
No Preference Left Behind: Group Distributional Preference Optimization
- 用信念分布建模个体偏好差异,实现群体多元意见对齐
- 实验显示GDPO比DPO更贴近目标信念分布,误差显著降低
- 适合需要包容多样观点的场景,如政策制定、社会舆情分析
群体内的偏好并非一致,而是呈现分布特征。现有对齐方法如直接偏好优化(DPO)虽试图让模型反映人类偏好,但难以捕捉群体内部的多元分布性,常偏向主导意见而忽视少数派。为此,我们提出群体分布偏好优化(GDPO),通过引入塑造个体偏好的信念概念,利用统计估计方法校准群体信念分布,并基于信念条件化偏好对齐语言模型,构建更具包容性的对齐框架。在合成可控意见生成与真实电影评论数据集上的实验表明,DPO无法对齐目标信念分布,而GDPO在训练过程中持续缩小这一差距。评估指标显示,GDPO在群体分布偏好对齐上显著优于现有方法,推动了多元共情式对齐的发展。
原文摘要 · Abstract (English)
Preferences within a group of people are not uniform but follow a distribution. While existing alignment methods like Direct Preference Optimization (DPO) attempt to steer models to reflect human preferences, they struggle to capture the distributional pluralistic preferences within a group. These methods often skew toward dominant preferences, overlooking the diversity of opinions, especially when conflicting preferences arise. To address this issue, we propose Group Distributional Preference Optimization (GDPO), a novel framework that aligns language models with the distribution of preferences within a group by incorporating the concept of beliefs that shape individual preferences. GDPO calibrates a language model using statistical estimation of the group's belief distribution and aligns the model with belief-conditioned preferences, offering a more inclusive alignment framework than traditional methods. In experiments using both synthetic controllable opinion generation and real-world movie review datasets, we show that DPO fails to align with the targeted belief distributions, while GDPO consistently reduces this alignment gap during training. Moreover, our evaluation metrics demonstrate that GDPO outperforms existing approaches in aligning with group distributional preferences, marking a significant advance in pluralistic alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。