发现大模型奖励模型偏好不道德回答,且存在系统性偏差。
Misaligned by Reward: Socially Undesirable Preferences in LLMs
- 将社会评估数据转为成对偏好数据,检测奖励模型倾向
- 5个模型均偏好不道德选项,产生系统性偏差分布
- 避免偏见会削弱对上下文的敏感性,暴露对齐难题
奖励模型是大语言模型对齐的关键组件,在训练中作为人类偏好的代理。然而现有评估主要依赖宽泛的指令遵循基准,难以揭示模型是否捕捉了社会可接受的偏好,导致重要社会对齐失败被掩盖。我们扩展奖励模型评测至四个社会影响深远的领域:偏见、安全、道德与伦理推理。提出一种框架,将社会评估数据转化为成对偏好数据,利用真实标签或方向性偏见指标。测试发现,五种公开奖励模型及两种指令微调模型作为奖励代理时,各领域表现差异显著,无一模型整体最优。模型普遍偏好社会不当回应,其偏好导致输出分布系统性偏差。更强的偏见规避会降低对上下文的敏感度,揭示了避免偏差与保持语境忠实之间的关键对齐权衡。结果表明,标准奖励基准不足以评估社会对齐,亟需直接测量奖励模型所编码社会偏好的评估方法。
原文摘要 · Abstract (English)
Reward models are a key component of large language model alignment, serving as proxies for human preferences during training. However, existing evaluations focus primarily on broad instruction-following benchmarks, providing limited insight into whether these models capture socially desirable preferences. As a result, important failures in social alignment can remain hidden. We extend reward-model benchmarking to four socially consequential domains: bias, safety, morality, and ethical reasoning. We introduce a framework that converts social evaluation datasets into pairwise preference data, leveraging gold labels where available and directional bias indicators otherwise. This enables us to test whether reward models prefer socially undesirable responses, and whether their preferences produce systematically biased distributions over selected outputs. Across five publicly available reward models and two instruction-tuned models used as reward proxies, we find substantial variation across domains, with no single model performing best overall. The models fall well short of strong social intelligence: they often prefer socially undesirable options, and their preferences produce systematically biased distributions. Moreover, stronger bias avoidance can reduce sensitivity to context, revealing a key alignment trade-off between avoiding biased outcomes and preserving contextual faithfulness. These findings show that standard reward benchmarks are insufficient for assessing social alignment and highlight the need for evaluations that directly measure the social preferences encoded in reward models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。