用检索增强提升数学推理评分模型的泛化能力。
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
- 通过两阶段检索,引入相似题目和解题步骤作为上下文
- 在多个真实数据集上优于现有基线模型
- 适合需要高可靠推理评估的研究者和开发者
尽管大语言模型在数学推理方面取得显著进展,过程奖励模型(PRMs)被用于评估推理步骤的逻辑有效性。然而,现有PRMs仍面临分布外(OOD)挑战:步骤级OOD源于不同模型类型与规模间的推理模式差异;问题级OOD则源于训练数据与实际问题之间的数据分布偏移。为此,本文提出检索增强型过程奖励模型(RetrievalPRM),采用两阶段检索增强机制,通过检索语义相近的问题与解题步骤作为先验信息,提升对目标步骤的评估能力,增强模型在不同模型与题型间的泛化性与推理一致性。大量实验证明,RetrievalPRM在多个真实世界数据集上均超越现有基线。开源内容包括一个检索增强数据集、一个用于PRM微调的框架及RetrievalPRM模型,为PRM性能树立新标准。
原文摘要 · Abstract (English)
While large language models (LLMs) have significantly advanced mathematical reasoning, Process Reward Models (PRMs) have been developed to evaluate the logical validity of reasoning steps. However, PRMs still struggle with out-of-distribution (OOD) challenges. This paper identifies key OOD issues, including step OOD, caused by differences in reasoning patterns across model types and sizes, and question OOD, which arises from dataset shifts between training data and real-world problems. To address these issues, we introduce Retrieval-Augmented Process Reward Model (RetrievalPRM), a novel framework designed to tackle these OOD issues. By utilizing a two-stage retrieval-enhanced mechanism, RetrievalPRM retrieves semantically similar questions and steps as a warmup, enhancing PRM's ability to evaluate target steps and improving generalization and reasoning consistency across different models and problem types. Our extensive experiments demonstrate that RetrievalPRM outperforms existing baselines across multiple real-world datasets. Our open-source contributions include a retrieval-enhanced dataset, a tuning framework for PRM training, and the RetrievalPRM model, establishing a new standard for PRM performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。