arXiv:2601.06180cs.LGcs.AI2026-01

让大模型更真实地学习人类偏好的强弱差异。

MixDPO: Modeling Preference Strength for Pluralistic Alignment

  • 引入混合逻辑模型,显式建模偏好强度的个体差异
  • 在三个数据集上提升对齐效果,最高增益11.2分
  • 适合需要精细捕捉人类判断差异的研究者

基于偏好的对齐目标隐含假设所有人类偏好表达强度一致,但现实中偏好强度因人而异、随情境变化。为解决这一偏差,本文提出混合逻辑直接偏好优化(MixDPO),作为直接偏好优化的推广,可建模偏好强度的分布差异。该方法使对齐目标能更准确捕捉训练样本间偏好表达强度的异质性。在三个偏好数据集上使用两个开源语言模型进行评估,结果显示,MixDPO在所有数据集上均提升整体对齐性能(如Pythia-2.8B模型提升11.2分),并有效保持子群体偏好;在推断出更高偏好异质性的场景中,增益最为显著。通过学习得到的偏好强度分布,混合理论将偏好异质性显式表达。代码已开源,支持复现。

原文摘要 · Abstract (English)

Preference based alignment objectives implicitly assume that all human preferences are expressed with equal strength. In practice, however, preference strength varies across individuals and contexts -- a phenomenon established in behavioral economics and discrete choice theory. This mismatch limits the ability of existing objectives to faithfully capture heterogeneous human judgments. Inspired by this literature, we introduce Mixed Logit Direct Preference Optimization (MixDPO), a generalization of Direct Preference Optimization that models variation in preference strength. MixDPO enables alignment objectives to capture heterogeneity in how strongly preferences are expressed across training examples. We evaluate MixDPO on three preference datasets using two open-weight language models. Across datasets, MixDPO improves aggregate alignment performance (+11.2 points on Pythia-2.8B) while preserving subgroup level preferences, with the largest gains appearing in settings with higher inferred preference heterogeneity. MixDPO makes preference heterogeneity explicit through learned strength distributions. We release our code for reproducibility.

偏好建模对齐优化异质性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。