arXiv:2601.20838cs.LGcs.AI2026-01被引 4

不同基础模型的奖励模型会继承不同的价值偏好,影响对齐效果。

Reward Models Inherit Value Biases from Pretraining

  • 通过分析10个开源奖励模型,发现其价值倾向与基础模型相关。
  • 相同微调条件下,Llama模型偏好'自主性',Gemma模型偏好'归属感'。
  • 这种偏差源自预训练阶段的隐式奖励信号,影响模型对齐安全。

奖励模型(RMs)在对齐大语言模型(LLMs)与人类价值观中起核心作用,但其研究远少于预训练和后训练的LLMs。由于RMs由LLMs初始化,继承了影响行为的表征,但这种影响的性质和程度尚不明确。通过对10个主流开源奖励模型的全面研究,使用经验证的心理语言学语料库,我们发现RMs在多个维度的人类价值观上表现出显著差异,这些差异与它们的基础模型相关。基于心理学中的“两大轴”(即自主性与归属感),我们发现以Llama为基础的奖励模型明显偏好‘自主性’,而以Gemma为基础的则偏好‘归属感’。这一现象即使在偏好数据和微调过程完全相同的情况下依然存在,且可追溯至相应指令微调和预训练模型的logits差异。这些对数概率差异本身可被形式化为一种隐式奖励模型;我们推导出可用的隐式奖励分数,并证明其同样表现出自主性/归属感的差异。通过对比实验,我们验证了该效应不仅可重复,而且极为持久。尽管奖励模型旨在反映人类偏好,但我们的证据表明,其输出仍受基础预训练模型的影响。这项工作强调了在预训练阶段进行安全与对齐工作的必要性,也表明开源开发者选择基础模型时,需同时考虑价值观与性能。

原文摘要 · Abstract (English)

Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pretrained and post-trained LLMs themselves. Because RMs are initialized from LLMs, they inherit representations that shape their behavior, but the nature and extent of this influence remain understudied. In a comprehensive study of 10 leading open-weight RMs using validated psycholinguistic corpora, we show that RMs exhibit significant differences along multiple dimensions of human value as a function of their base model. Using the "Big Two" psychological axes, we show a robust preference of Llama RMs for "agency" and a corresponding robust preference of Gemma RMs for "communion." This phenomenon holds even when the preference data and finetuning process are identical, and we trace it back to the logits of the respective instruction-tuned and pretrained models. These log-probability differences themselves can be formulated as an implicit RM; we derive usable implicit reward scores and show that they exhibit the very same agency/communion difference. We run experiments training RMs with ablations for preference data source and quantity, which demonstrate that this effect is not only repeatable but surprisingly durable. Despite RMs being designed to represent human preferences, our evidence shows that their outputs are influenced by the pretrained LLMs on which they are based. This work underscores the importance of safety and alignment efforts at the pretraining stage, and makes clear that open-source developers' choice of base model is as much a consideration of values as of performance.

奖励模型价值对齐偏见继承大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。