arXiv:2606.04579cs.AI2026-06KDD

用工具轨迹训练科学推理奖励模型,提升模型准确性与可验证性。

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification

论文配图:SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification
图 1 · 摘自论文原文
  • 构建包含工具使用步骤的链式推理数据集SCI-PRM70K
  • 在测试时通过Best-of-N选择显著提升模型表现
  • 适用于需要精确工具使用与事实一致性的科学推理任务

尽管过程奖励模型(PRMs)在数学推理中取得显著进展,但在生物学、化学、物理学等复杂科学领域仍鲜有应用。科学问题不仅要求逻辑严谨,还需事实一致性和领域工具的精准使用,而现有模型常出现幻觉且缺乏验证。本文首先构建了包含链式工具轨迹的大型数据集SCI-PRM70K,明确融合推理与科学工具执行过程。基于此,训练了一个高效奖励模型Sci-PRM,可在一次推理中对每一步的工具选择、执行准确性和结果解释提供细粒度监督。实验表明,Sci-PRM能显著提升基础模型性能:(1)通过测试时的Best-of-N选择实现有效扩展;(2)在强化学习中作为密集奖励信号,缓解优势消失问题,帮助模型突破现有性能瓶颈。

原文摘要 · Abstract (English)

While Process Reward Models (PRMs) have achieved remarkable success in mathematical reasoning, their application in complex scientific domains-such as biology, chemistry, and physics remains largely unexplored. Scientific problems demand not only logical rigor but also factual consistency and the precise usage of domain-specific tools, areas where current models often suffer from hallucinations and lack of verification. In this paper, we first construct SCIPRM70K, a large-scale dataset featuring Chain-of-Tool trajectories that explicitly interleave reasoning with the execution of scientific tools. Building upon this, we train an efficient reward model called Sci-PRM to provide fine-grained supervision on tool selection, execution accuracy, and result interpretation at each step in one inference. Experiments demonstrate that Sci-PRM significantly enhances foundation models in two key aspects: (1) it enables effective test-time scaling via Best-of-N selection; and (2) when integrated into Reinforcement Learning, it serves as a dense reward signal that mitigates the critical issue of advantage disappearance, allowing the model to break through existing performance ceilings.

科学推理奖励模型工具使用验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。