arXiv:2604.08559cs.CLcs.AI2026-04综述被引 4

梳理医学大模型推理方法,构建真实临床数据基准

Medical Reasoning with Large Language Models: A Survey and MR-Bench

  • 基于认知理论将医学推理分为三类迭代过程
  • 提出MR-Bench新基准,发现模型在真实场景表现显著下降
  • 适合医疗AI研究者与临床决策系统开发者参考

大型语言模型(LLMs)在医学考试类任务中表现强劲,激发了其在真实临床环境中的应用兴趣。然而,临床决策具有安全敏感性、情境依赖性和证据动态演变的特点,可靠性能不仅依赖事实记忆,更需稳健的医学推理能力。本文系统综述了医学大模型推理研究,基于临床推理的认知理论,将医学推理建模为溯因、演绎和归纳的迭代过程,并将现有方法归纳为七类技术路径,涵盖训练型与免训练方法。我们在统一实验设置下对代表性模型进行跨基准评估,实现更系统可比的实证分析。为更好衡量临床真实推理能力,我们引入MR-Bench,一个源自真实医院数据的基准。在该基准上的评估揭示了模型在考试级表现与真实临床任务准确性之间存在明显差距。整体上,本综述提供了对现有医学推理方法、基准与评估实践的统一视角,指出现有模型性能与真实临床推理需求之间的关键鸿沟。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved strong performance on medical exam-style tasks, motivating growing interest in their deployment in real-world clinical settings. However, clinical decision-making is inherently safety-critical, context-dependent, and conducted under evolving evidence. In such situations, reliable LLM performance depends not on factual recall alone, but on robust medical reasoning. In this work, we present a comprehensive review of medical reasoning with LLMs. Grounded in cognitive theories of clinical reasoning, we conceptualize medical reasoning as an iterative process of abduction, deduction, and induction, and organize existing methods into seven major technical routes spanning training-based and training-free approaches. We further conduct a unified cross-benchmark evaluation of representative medical reasoning models under a consistent experimental setting, enabling a more systematic and comparable assessment of the empirical impact of existing methods. To better assess clinically grounded reasoning, we introduce MR-Bench, a benchmark derived from real-world hospital data. Evaluations on MR-Bench expose a pronounced gap between exam-level performance and accuracy on authentic clinical decision tasks. Overall, this survey provides a unified view of existing medical reasoning methods, benchmarks, and evaluation practices, and highlights key gaps between current model performance and the requirements of real-world clinical reasoning.

医学推理大模型基准测试临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。