奖励模型更看重答案一致性,而非真正逻辑正确性。
Reward Models Identify Consistency, Not Causality
- 通过分析发现,奖励模型依赖结构一致性而非因果推理
- 改动数值或打断推理流程会显著影响评分,但删掉问题影响很小
- 适合关注大模型对齐机制缺陷的研究者阅读
奖励模型(RMs)在对齐大语言模型(LLMs)与人类偏好、提升推理质量方面发挥关键作用。传统上,RMs 通过评估输出的正确性和连贯性进行排序。然而,本文揭示了几个颠覆常识的发现:最先进的奖励模型更注重结构一致性,而非因果正确性。具体而言,删除问题陈述对评分影响极小,而修改数值或破坏推理流程则显著改变奖励值。此外,奖励模型高度依赖完整的推理路径;若推理被截断或不完整,评分会出现显著波动,表明其主要依赖学习到的推理模式,而非对问题的实质理解。这些发现适用于多种架构、数据集和任务,提出三点核心洞见:(1)奖励模型主要评估连贯性而非真实推理质量;(2)显式问题理解在奖励分配中的作用被夸大;(3)现有奖励模型可能更擅长排序而非验证逻辑有效性。结果揭示了当前奖励建模方法的根本局限,强调需发展以因果性为导向的新型奖励模型。
原文摘要 · Abstract (English)
Reward models (RMs) play a crucial role in aligning large language models (LLMs) with human preferences and enhancing reasoning quality. Traditionally, RMs are trained to rank candidate outputs based on their correctness and coherence. However, in this work, we present several surprising findings that challenge common assumptions about RM behavior. Our analysis reveals that state-of-the-art reward models prioritize structural consistency over causal correctness. Specifically, removing the problem statement has minimal impact on reward scores, whereas altering numerical values or disrupting the reasoning flow significantly affects RM outputs. Furthermore, RMs exhibit a strong dependence on complete reasoning trajectories truncated or incomplete steps lead to significant variations in reward assignments, indicating that RMs primarily rely on learned reasoning patterns rather than explicit problem comprehension. These findings hold across multiple architectures, datasets, and tasks, leading to three key insights: (1) RMs primarily assess coherence rather than true reasoning quality; (2) The role of explicit problem comprehension in reward assignment is overstated; (3) Current RMs may be more effective at ranking responses than verifying logical validity. Our results suggest a fundamental limitation in existing reward modeling approaches, emphasizing the need for a shift toward causality-aware reward models that go beyond consistency-driven evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。