arXiv:2507.17849cs.CL2025-07ACL被引 19

让大模型在复杂任务中自我评估,动态生成精准奖励。

Dynamic and Generalizable Process Reward Modeling

  • 构建奖励树,细粒度捕捉多维度评价标准。
  • 用帕累托优势筛选正负样本对,提升判断准确性。
  • 跨领域表现强,适合需要精细过程监督的任务。

过程奖励模型(PRMs)通过提供密集奖励信号,对大语言模型在复杂场景中的决策至关重要。然而,现有方法多依赖启发式策略,泛化能力差。尽管已有研究使用大模型作为评判者提供通用奖励,但大多只关注反馈结果,忽略了文本中蕴含的深层指导信息。同时,静态且粗粒度的评估标准难以适应复杂过程监督需求。为此,我们提出动态可泛化的流程奖励建模(DG-PRM),其核心是奖励树结构,用于捕获和存储细粒度、多维度的奖励标准。DG-PRM可动态选择每一步的奖励信号进行评分。为处理多面性奖励信号,我们首次引入帕累托支配估计,识别具有区分性的正负样本对。实验表明,DG-PRM在主流基准上表现卓越,显著提升各任务中密集奖励场景下的模型性能。进一步分析显示,该方法在分布外场景下也表现出良好适应性,展现出优异的泛化能力。

原文摘要 · Abstract (English)

Process Reward Models (PRMs) are crucial for guiding Large Language Models (LLMs) in complex scenarios by providing dense reward signals. However, existing PRMs primarily rely on heuristic approaches, which struggle with cross-domain generalization. While LLM-as-judge has been proposed to provide generalized rewards, current research has focused mainly on feedback results, overlooking the meaningful guidance embedded within the text. Additionally, static and coarse-grained evaluation criteria struggle to adapt to complex process supervision. To tackle these challenges, we propose Dynamic and Generalizable Process Reward Modeling (DG-PRM), which features a reward tree to capture and store fine-grained, multi-dimensional reward criteria. DG-PRM dynamically selects reward signals for step-wise reward scoring. To handle multifaceted reward signals, we pioneeringly adopt Pareto dominance estimation to identify discriminative positive and negative pairs. Experimental results show that DG-PRM achieves stunning performance on prevailing benchmarks, significantly boosting model performance across tasks with dense rewards. Further analysis reveals that DG-PRM adapts well to out-of-distribution scenarios, demonstrating exceptional generalizability.

奖励建模大模型泛化能力动态评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。