arXiv:2510.06391cs.CLcs.AI2025-10EMNLP被引 1

揭示奖励模型对不同群体意见的偏倚,警示其可能强化社会刻板印象。

Reward Model Perspectives: Whose Opinions Do Reward Models Reward?

  • 构建框架量化奖励模型与真实人群意见的对齐程度。
  • 发现多个群体在模型中被系统性误判,且倾向奖励有害刻板印象。
  • 提示调整提示词无法根本解决偏倚问题,需更审慎设计对齐机制。

奖励模型(RMs)在语言模型对齐中起核心作用,常作为人类偏好代理以引导下游模型行为。然而我们对其行为的理解仍有限。本文工作包括:(i) 建立衡量奖励模型捕捉观点对齐程度的框架;(ii) 探究奖励模型在社会人口学特征上的偏差程度;(iii) 研究通过提示词引导奖励偏向特定群体偏好的效果。研究聚焦争议性话题中的主观多元视角,可量化评估奖励模型在意见、态度与价值观层面的表现。结果表明,奖励模型与多个社会群体存在显著脱节,可能系统性奖励有害刻板印象,仅靠提示词引导不足以克服这些局限。研究强调在偏好学习阶段需更谨慎评估奖励模型行为,防止社会偏见在语言技术中传播。

原文摘要 · Abstract (English)

Reward models (RMs) are central to the alignment of language models (LMs). An RM often serves as a proxy for human preferences to guide downstream LM behavior. However, our understanding of RM behavior is limited. Our work (i) formalizes a framework for measuring the alignment of opinions captured by RMs, (ii) investigates the extent to which RMs demonstrate sociodemographic biases, and (iii) explores the effects of prompting to steer rewards towards the preferences of a target group. We study the subjective and diverse perspectives on controversial topics, which allows us to quantify RM perspectives in terms of their opinions, attitudes, and values. We show that RMs are poorly aligned with several demographic groups and can systematically reward harmful stereotypes, and steering alone is not enough to overcome these limitations. Our findings underscore the need for more careful consideration of RM behavior in model alignment during preference learning to prevent the propagation of unwanted social biases in the language technologies that we use.

奖励模型社会偏见模型对齐话语分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。