arXiv:2506.11474cs.CL2025-06EMNLP被引 29

用临床指南验证每步推理,让医学大模型更准地诊断。

Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards

  • 每步推理都用临床指南和文献检索验证,定位错误更精准。
  • 在5个医疗问答数据集上,性能提升最高达13.50%。
  • 可插即用,80亿参数小模型首次在MedQA上达80%以上准确率。

大型语言模型在临床决策中展现出潜力,但现有方法难以精确定位并修正推理过程中的具体错误。这一局限在医学领域尤为关键,因为识别和纠正推理错误对准确诊断和有效患者照护至关重要。我们提出Med-PRM,一种基于检索增强生成的流程奖励建模框架,通过从临床指南和文献中检索证据,验证每一步推理。该方法能以细粒度方式精确评估推理质量。在五个医疗问答基准和两个开放式诊断任务上的评估显示,Med-PRM将基线模型性能提升最高达13.50%。此外,我们将Med-PRM以即插即用方式集成到强大策略模型(如Meerkat)中,首次使80亿参数的小模型在MedQA上达到超过80%的准确率。代码与数据已公开于https://med-prm.github.io/。

原文摘要 · Abstract (English)

Large language models have shown promise in clinical decision making, but current approaches struggle to localize and correct errors at specific steps of the reasoning process. This limitation is critical in medicine, where identifying and addressing reasoning errors is essential for accurate diagnosis and effective patient care. We introduce Med-PRM, a process reward modeling framework that leverages retrieval-augmented generation to verify each reasoning step against established medical knowledge bases. By verifying intermediate reasoning steps with evidence retrieved from clinical guidelines and literature, our model can precisely assess the reasoning quality in a fine-grained manner. Evaluations on five medical QA benchmarks and two open-ended diagnostic tasks demonstrate that Med-PRM achieves state-of-the-art performance, with improving the performance of base models by up to 13.50% using Med-PRM. Moreover, we demonstrate the generality of Med-PRM by integrating it in a plug-and-play fashion with strong policy models such as Meerkat, achieving over 80\% accuracy on MedQA for the first time using small-scale models of 8 billion parameters. Our code and data are available at: https://med-prm.github.io/

医学推理推理验证提示工程大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。