arXiv:2602.04649cs.CL2026-02ACL被引 9

提出新评估指标,识别奖励模型推理错误但答案对的问题

Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models

  • 用推理一致性衡量模型推理与人类判断的匹配度
  • 新方法在RM-Bench和JudgeBench上分别达87.1%和82%
  • 适合关注强化学习对齐安全的研究者

生成式奖励模型(GenRMs)和基于大模型的评判系统存在误导性对齐:它们能给出正确答案,但推理过程错误,因为训练和评估只关注结果准确率,这削弱了其在强化学习人类反馈(RLHF)中的泛化能力。本文提出「推理一致性」这一细粒度指标,量化模型推理过程与人类判断的一致性。对前沿模型的评估显示,该指标能有效区分顶尖模型并检测出误导性对齐,而结果准确率则无法做到。为弥合差距,我们引入结合推理一致性和结果准确性的混合信号用于GenRM训练。该方法在RM-Bench(87.1%)和JudgeBench(82%)上达到当前最优性能,平均超越仅依赖结果准确率的基线5%。使用该奖励模型进行RLHF,在Arena Hard v2上显著提升表现,尤其在创意写作任务中实现7%的改进。进一步分析表明,该方法成功避开误导性对齐陷阱,逆转了仅使用结果准确率训练时推理一致性下降的趋势。

原文摘要 · Abstract (English)

Generative Reward Models (GenRMs) and LLM-as-a-Judge exhibit deceptive alignment by producing correct judgments for incorrect reasons, as they are trained and evaluated to prioritize Outcome Accuracy, which undermines their ability to generalize during RLHF. We introduce Rationale Consistency, a fine-grained metric that quantifies the alignment between the model's reasoning process and human judgment. Our evaluation of frontier models reveals that rationale consistency effectively discriminates among state-of-the-art models and detects deceptive alignment, while outcome accuracy falls short in both respects. To mitigate this gap, we introduce a hybrid signal that combines rationale consistency with outcome accuracy for GenRM training. Our training method achieves state-of-the-art performance on RM-Bench (87.1%) and JudgeBench (82%), surpassing outcome-only baselines by an average of 5%. Using RM during RLHF, our method effectively improves performance as demonstrated on Arena Hard v2, notably yielding a 7% improvement in creative writing tasks. Further analysis confirms that our method escapes the deceptive alignment trap, effectively reversing the decline in rationale consistency observed in outcome-only training.

奖励模型对齐评估推理一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。