用过程奖励模型提升大模型推理对齐效果
From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment
- 引入过程奖励模型,解决结果评分与推理过程不匹配问题
- 在对话、摘要和推理任务中提升GPT-4评分3.6%~10.3%
- 无需人工标注,自动实现评分与偏好一致性
推理时对齐方法因高效性受到关注,但主流的奖励引导搜索(RGS)依赖结果奖励模型(ORM),而ORM仅针对完整输出打分,与RGS所需的中间过程评分存在粒度不匹配,导致评分不一致、对齐效果不佳。为此,本文提出过程奖励模型(PRM),并定义两个核心目标:评分一致性(确保部分与完整响应评价一致)和偏好一致性(部分序列评估符合人类偏好)。基于此,提出SP-PRM框架,通过无监督的双一致性模块实现评分一致性与偏好一致性,无需人工标注。在对话、摘要和推理任务上的大量实验表明,SP-PRM显著提升现有RGS方法性能,使GPT-4评估得分平均提升3.6%至10.3%。
原文摘要 · Abstract (English)
Inference-time alignment methods have gained significant attention for their efficiency and effectiveness in aligning large language models (LLMs) with human preferences. However, existing dominant approaches using reward-guided search (RGS) primarily rely on outcome reward models (ORMs), which suffer from a critical granularity mismatch: ORMs are designed to provide outcome rewards for complete responses, while RGS methods rely on process rewards to guide the policy, leading to inconsistent scoring and suboptimal alignment. To address this challenge, we introduce process reward models (PRMs) into RGS and argue that an ideal PRM should satisfy two objectives: Score Consistency, ensuring coherent evaluation across partial and complete responses, and Preference Consistency, aligning partial sequence assessments with human preferences. Based on these, we propose SP-PRM, a novel dual-consistency framework integrating score consistency-based and preference consistency-based partial evaluation modules without relying on human annotation. Extensive experiments on dialogue, summarization, and reasoning tasks demonstrate that SP-PRM substantially enhances existing RGS methods, achieving a 3.6%-10.3% improvement in GPT-4 evaluation scores across all tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。