构建可扩展的数学证明评分模型,自动评估完整证明过程。
Proof-RM: A Scalable and Generalizable Reward Model for Math Proof
- 用LLM自动生成大量高质量的题目-证明-验证三元组数据。
- 在多难度、多风格、多错误类型上实现高准确率与强泛化能力。
- 适合需要提升数学推理能力的LLM研究者和开发者使用。
尽管大语言模型(LLMs)通过可验证奖励强化学习(RLVR)展现出强大的数学推理能力,但许多高级数学问题为证明题,仅靠答案匹配无法判断证明真伪。为此,本文设计了一种可扩展的数据构建流程,以极少人工投入,利用LLM生成大规模高质量的“问题-证明-验证”三元组数据。通过系统性地改变问题来源、生成方法和模型配置,构建了涵盖多种难度、语言风格和错误类型的证明对,并经分层人工评审确保标签一致性。基于这些数据,训练了一个证明检查奖励模型(RM),采用“LLM作为奖励模型的奖励模型”策略及平衡标记权重,稳定强化学习过程。实验验证了该模型在可扩展性、奖励准确性、泛化能力及测试时指导方面的优异表现,为增强LLM数学能力提供了实用方案与工具。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have demonstrated strong math reasoning abilities through Reinforcement Learning with *Verifiable Rewards* (RLVR), many advanced mathematical problems are proof-based, with no guaranteed way to determine the authenticity of a proof by simple answer matching. To enable automatic verification, a Reward Model (RM) capable of reliably evaluating full proof processes is required. In this work, we design a *scalable* data-construction pipeline that, with minimal human effort, leverages LLMs to generate a large quantity of high-quality ``**question-proof-check**'' triplet data. By systematically varying problem sources, generation methods, and model configurations, we create diverse problem-proof pairs spanning multiple difficulty levels, linguistic styles, and error types, subsequently filtered through hierarchical human review for label alignment. Utilizing these data, we train a proof-checking RM, incorporating an ``LLM-as-a-RM-for-RM'' approach and balanced token weighting to stabilize the RL process. Our experiments validate the model's scalability and strong performance from multiple perspectives, including reward accuracy, generalization ability and test-time guidance, providing important practical recipes and tools for strengthening LLM mathematical capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。