通过分析奖励模型对所有单个词元的评分,揭示其内在偏见与不一致性。
Reward Model Interpretability via Optimal and Pessimal Tokens
- 系统扫描奖励模型对全部词元的评分,挖掘其决策模式。
- 发现不同模型间评分差异大,且对高频词过度偏好。
- 适合关注模型对齐风险与价值观偏差的研究者阅读。
奖励建模已成为对齐大型语言模型与人类价值观的关键组件。尽管人们广泛关注其在微调生成模型中的应用,但这些直接将提示-响应对转化为标量奖励的奖励模型本身仍缺乏深入研究。本文提出一种新方法,通过全面分析奖励模型在其整个词汇空间中的响应,揭示多个显著现象:(i) 在相似目标下训练的不同模型间存在显著异质性;(ii) 模型对高分与低分词元的编码存在系统性不对称;(iii) 对提示框架高度敏感,表现出类似人类认知偏见的特征;(iv) 显著高估高频词。我们在十种不同参数规模与架构的开源奖励模型上验证了这些效应。结果挑战了奖励模型可互换性的假设,也质疑其作为复杂、上下文依赖的人类价值观代理的可靠性。我们发现这些模型可能隐含对某些身份群体的偏见,这可能是无害性训练带来的意外后果,有风险通过下游部署于数百万用户的大型语言模型传播。
原文摘要 · Abstract (English)
Reward modeling has emerged as a crucial component in aligning large language models with human values. Significant attention has focused on using reward models as a means for fine-tuning generative models. However, the reward models themselves -- which directly encode human value judgments by turning prompt-response pairs into scalar rewards -- remain relatively understudied. We present a novel approach to reward model interpretability through exhaustive analysis of their responses across their entire vocabulary space. By examining how different reward models score every possible single-token response to value-laden prompts, we uncover several striking findings: (i) substantial heterogeneity between models trained on similar objectives, (ii) systematic asymmetries in how models encode high- vs low-scoring tokens, (iii) significant sensitivity to prompt framing that mirrors human cognitive biases, and (iv) overvaluation of more frequent tokens. We demonstrate these effects across ten recent open-source reward models of varying parameter counts and architectures. Our results challenge assumptions about the interchangeability of reward models, as well as their suitability as proxies of complex and context-dependent human values. We find that these models can encode concerning biases toward certain identity groups, which may emerge as unintended consequences of harmlessness training -- distortions that risk propagating through the downstream large language models now deployed to millions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。