arXiv:2502.06737cs.LG2025-02ICML被引 45

VersaPRM让大模型在多领域推理中表现更优,尤其法律领域提升显著。

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

  • 用合成推理数据训练多领域奖励模型,提升泛化能力。
  • 法律领域性能比基线高7.9%,远超现有数学专用模型。
  • 开源全部数据代码模型,助力社区研究与应用。

过程奖励模型(PRMs)通过增加推理时计算量,有效提升了大语言模型的数学推理能力。然而,现有PRMs主要基于数学数据训练,其在非数学领域的泛化能力尚未得到充分验证。本文首次证明当前PRMs在其他领域表现不佳。为此,我们提出VersaPRM,一种基于新型合成推理数据生成与标注方法的多领域PRM。VersaPRM在多个领域实现一致性能提升。例如,在MMLU-Pro法律类别中,通过加权多数投票,其性能比多数投票基线高出7.9%,超过Qwen2.5-Math-PRM的1.3%提升。我们还开源了VersaPRM的所有数据、代码和模型。

原文摘要 · Abstract (English)

Process Reward Models (PRMs) have proven effective at enhancing mathematical reasoning for Large Language Models (LLMs) by leveraging increased inference-time computation. However, they are predominantly trained on mathematical data and their generalizability to non-mathematical domains has not been rigorously studied. In response, this work first shows that current PRMs have poor performance in other domains. To address this limitation, we introduce VersaPRM, a multi-domain PRM trained on synthetic reasoning data generated using our novel data generation and annotation method. VersaPRM achieves consistent performance gains across diverse domains. For instance, in the MMLU-Pro category of Law, VersaPRM via weighted majority voting, achieves a 7.9% performance gain over the majority voting baseline -- surpassing Qwen2.5-Math-PRM's gain of 1.3%. We further contribute to the community by open-sourcing all data, code and models for VersaPRM.

推理增强多领域奖励模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。