首个面向医疗推理的细粒度奖励模型评测基准,解决临床错误检测难题。
MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning

- 基于临床推理蓝图构建三阶段数据生成流程,覆盖14类细粒度错误
- 含6500道题、13000条推理链和11.3万步级标签,首创四级严重性分级
- 验证发现现有模型在医疗推理错误检测上普遍薄弱,适合医疗AI安全研究者
过程级奖励模型(PRMs)对引导大语言模型复杂推理至关重要,但现有评测基准仅覆盖数学等通用领域,未能涵盖具有安全性关键性、知识密集性和多样错误模式的医疗推理。缺乏可靠的医疗PRM评估框架,导致无法量化模型在临床推理中的错误检测能力,其在真实医疗应用中的安全性无法验证。本文提出MedPRMBench,首个面向医疗领域的过程级奖励模型评测基准。该基准基于临床推理蓝图(CRBs)的三阶段流程,从7个医学问答源系统生成高质量评估数据,涵盖3类共14种细粒度错误类型(简洁性、合理性、敏感性),并建立首个四级严重性分级体系以量化临床影响。基准包含6,500个问题、13,000条推理链和113,910个步骤级标签,另有6,879个问题用于训练。我们的医疗PRM基线达到87.1%的整体PRMScore,显著优于所有基线,并可作为即插即用的验证器,使下游医学QA准确率提升3.2–6.7个百分点。对专有前沿模型、开源推理模型及医疗专用模型的系统评估揭示了当前模型在医疗推理错误检测上的关键缺陷,为未来PRM改进指明方向。
原文摘要 · Abstract (English)
Process-Level Reward Models (PRMs) are essential for guiding complex reasoning in large language models, yet existing PRM benchmarks cover only general domains such as mathematics, failing to address medical reasoning -- which is uniquely characterized by safety criticality, knowledge intensity, and diverse error patterns. Without a reliable medical PRM evaluation framework, we cannot quantify models' error detection capabilities in clinical reasoning, leaving their safety in real-world healthcare applications unverified. We propose MedPRMBench, the first process-level reward model benchmark for the medical domain. Built through a three-phase pipeline based on Clinical Reasoning Blueprints (CRBs), MedPRMBench systematically generates high-quality evaluation data from seven medical QA sources, covering 14 fine-grained error types across three categories (Simplicity, Soundness, and Sensitivity) with the first 4-level severity grading system to quantify clinical impact. The benchmark comprises 6{,}500 questions with 13{,}000 reasoning chains and 113{,}910 step-level labels, plus 6{,}879 questions for training. Our medical PRM baseline achieves an 87.1\% overall PRMScore -- substantially surpassing all baselines -- and serves as a plug-and-play verifier that improves downstream medical QA accuracy by 3.2--6.7 percentage points. Systematic evaluation spanning proprietary frontier models, open-source reasoning models, and medical-specialized models reveals critical weaknesses in current models' medical reasoning error detection capabilities, providing clear directions for future PRM improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。