arXiv:2601.17223cs.CLcs.AI2026-01ACL被引 12

用规则验证中间推理步骤,让大模型决策更可信可靠。

Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning

  • 用确定性规则验证每步推理,避免神经裁判的偏见与漏洞。
  • 在医学证据评估中,推理一致性提升,F1最高比顶尖模型高20%。
  • 适合需要严格逻辑与可解释性的医疗、法律等专业领域使用。

基于可验证奖励的强化学习(RLVR)研究表明,大型语言模型可通过结果级验证信号(如代码单元测试或数学精确匹配)显著提升性能。与此同时,过程监督长期被用于引导模型的中间推理行为,但现有方法依赖神经裁判评分思维链步骤,存在透明度低、偏见和奖励欺骗等问题。为此,我们提出可验证过程奖励模型(VPRM),一种通过确定性规则验证器检查中间推理步骤的强化学习框架。我们将VPRM应用于医学证据综合中的偏倚风险评估,该领域具备指南定义的标准与规则化决策路径,支持推理轨迹的程序化验证。在多个数据集上,VPRM生成的推理高度符合领域规则,且步骤决策与最终标签间的一致性显著提高。结果显示,VPRM相比最先进模型最高提升20% F1,较可验证结果奖励提升6.5%,在证据依据性和逻辑连贯性方面均有明显改善。

原文摘要 · Abstract (English)

Recent work on reinforcement learning with verifiable rewards (RLVR) has shown that large language models (LLMs) can be substantially improved using outcome-level verification signals, such as unit tests for code or exact-match checks for mathematics. In parallel, process supervision has long been explored as a way to shape the intermediate reasoning behaviour of LLMs, but existing approaches rely on neural judges to score chain-of-thought steps, leaving them vulnerable to opacity, bias, and reward hacking. To address this gap, we introduce Verifiable Process Reward Models (VPRMs), a reinforcement-learning framework in which intermediate reasoning steps are checked by deterministic, rule-based verifiers. We apply VPRMs to risk-of-bias assessment for medical evidence synthesis, a domain where guideline-defined criteria and rule-based decision paths enable programmatic verification of reasoning traces. Across multiple datasets, we find that VPRMs generate reasoning that adheres closely to domain rules and achieve substantially higher coherence between step-level decisions and final labels. Results show that VPRMs achieve up to 20% higher F1 than state-of-the-art models and 6.5% higher than verifiable outcome rewards, with substantial gains in evidence grounding and logical coherence.

推理验证强化学习医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。