通过强化推理一致性提升生成式奖励模型的对齐效果
R-Align: Enhancing Generative Reward Models through Rationale-Centric Meta-Judging
- 基于推理过程与参考判断的一致性来评估奖励模型
- 发现高错误推理一致性会导致强化学习策略退化
- 新方法R-Align通过显式监督推理对齐,提升多任务表现
强化学习从人类反馈(RLHF)在主观领域中仍是大语言模型对齐的关键。现有生成式奖励模型(GenRM)虽能生成推理过程再预测偏好,但训练与评估仍仅依赖结果标签,未检验推理质量。我们发现,推理保真度——即模型偏好决策与其参考推理的一致性——比传统标签准确率更能预测下游RLHF效果。具体地,我们利用现有奖励模型基准计算虚假正确率(S-Corr),即标签正确但推理与黄金判断不一致的比例。实证表明,即使先进GenRM也存在显著S-Corr,且更高S-Corr导致优化中策略退化。为此提出基于推理中心对齐的R-Align方法,引入黄金判断并显式监督推理一致性。R-Align有效降低基准上S-Corr,并在STEM、编程、指令遵循及通用任务中持续提升主模型性能。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) remains indispensable for aligning large language models (LLMs) in subjective domains. To enhance robustness, recent work shifts toward Generative Reward Models (GenRMs) that generate rationales before predicting preferences. Yet in GenRM training and evaluation, practice remains outcome-label-only, leaving reasoning quality unchecked. We show that reasoning fidelity-the consistency between a GenRM's preference decision and reference decision rationales-is highly predictive of downstream RLHF outcomes, beyond standard label accuracy. Specifically, we repurpose existing reward-model benchmarks to compute Spurious Correctness (S-Corr)-the fraction of label-correct decisions with rationales misaligned with golden judgments. Our empirical evaluation reveals substantial S-Corr even for competitive GenRMs, and higher S-Corr is associated with policy degeneration under optimization. To improve fidelity, we propose Rationale-Centric Alignment, R-Align, which augments training with gold judgments and explicitly supervises rationale alignment. R-Align reduces S-Corr on RM benchmarks and yields consistent gains in actor performance across STEM, coding, instruction following, and general tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。