用多个奖励函数建模多元偏好,让模型更公平地回应不同用户需求。
Pairwise Calibrated Rewards for Pluralistic Alignment
- 基于成对偏好学习多奖励函数分布,无需标注者身份或预设分组
- 训练后奖励函数的偏好比例与人工判断一致,实现配对校准
- 适合需要尊重多元价值观的应用,如跨文化对话系统
当前对齐方法假设存在单一普适的理想行为标准,但人类偏好常因用户、情境和文化而异。现有方法将分歧简化为多数意见,导致少数观点被忽视。为此,我们提出通过多个奖励函数构成的分布来反映多元偏好,每个函数诱导出不同的对齐策略。该分布直接从成对偏好数据中学习,不依赖标注者标识或预定义群体。相反,标注者的分歧被视为有信息量的软标签。核心标准是配对校准:对任意两个候选回复,偏好某一回复的奖励函数比例,应等于实际支持该回复的标注者比例。我们证明,即使小规模无异常值的集成也能准确表示多元偏好分布。实证上,我们提出并验证了一种实用的训练启发式方法,结果表明其显著提升校准度,意味着对多元价值的更忠实建模。
原文摘要 · Abstract (English)
Current alignment pipelines presume a single, universal notion of desirable behavior. However, human preferences often diverge across users, contexts, and cultures. As a result, disagreement collapses into the majority signal and minority perspectives are discounted. To address this, we propose reflecting diverse human preferences through a distribution over multiple reward functions, each inducing a distinct aligned policy. The distribution is learned directly from pairwise preference without annotator identifiers or predefined groups. Instead, annotator disagreements are treated as informative soft labels. Our central criterion is pairwise calibration: for every pair of candidate responses, the proportion of reward functions preferring one response matches the fraction of annotators with that preference. We prove that even a small outlier-free ensemble can accurately represent diverse preference distributions. Empirically, we introduce and validate a practical training heuristic to learn such ensembles, and demonstrate its effectiveness through improved calibration, implying a more faithful representation of pluralistic values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。