arXiv:2607.24484cs.LGcs.CL2026-07

研究奖励模型记住了什么,发现它们会偏倚记忆简单线索。

What do Reward Models Memorize?

  • 用反事实方法检测奖励模型的记忆偏差
  • 在未见配对中过度依赖长度、合规性等简单线索
  • 易样本和数据特定捷径导致判断失准,适合做安全评估

本文通过测量两个人类偏好数据集上的反事实记忆化,研究了判别式训练的奖励模型(RMs)记住了什么。结果表明:1)RMs将记忆错误分配给容易且高置信度的偏好对;2)记忆数据集特有的捷径(如模型身份、用户采样策略);3)面对未见过的偏好对时,过度泛化人类偏好的简单启发式相关特征(如文本长度、合规性)。总体而言,从人类偏好数据中判别式训练的奖励模型存在偏倚,尚无法在上下文依赖场景中准确判断回复质量。

原文摘要 · Abstract (English)

This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strategy), and 3) overgeneralize simple heuristic correlates of human preference (e.g., length, compliance) when confronted with unseen preference pairs. Overall, our findings indicate that discriminative training of RMs from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios.

奖励模型记忆偏差偏好学习AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。