arXiv:2607.07761cs.AI2026-07中稿 · Machine Intelligen…综述被引 5

梳理医学大模型在临床推理中的能力与需求匹配问题

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

论文配图:Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning
图 1 · 摘自论文原文
  • 按米勒金字塔构建五级临床能力框架,对应不同推理模式
  • 18个模型测试显示专科模型擅长诊断,通用模型更适于决策支持
  • 提出多层级医学推理基准数据集,助力模型评估与改进

大型语言模型(LLMs)在医疗领域展现出日益增长的临床推理潜力。本文综述了医学LLMs在推理应用与需求方面的最新进展,采用临床实践与计算方法双重视角。临床侧基于米勒金字塔建立五级能力体系,涵盖从知识回忆到动态病例管理的全过程;计算侧将演绎、归纳和溯因推理模式与常见医学目标对齐。同时,我们构建了一个覆盖五个层级医学推理能力的基准数据集,并对18个前沿模型进行评测,结果表明:医学专科模型在以诊断为中心的任务中表现优异,而通用模型在决策支持与对话任务中更具优势。最后,文章讨论了当前挑战,包括数据限制、幻觉与可解释性问题,并展望更安全、可靠且适用于临床工作流的系统发展方向。

原文摘要 · Abstract (English)

Large language models (LLMs) have emerged as important tools in healthcare, showing growing potential for clinical reasoning and patient care. This survey examines recent progress in medical LLMs, focusing on reasoning applications and requirements. We present a dual-view approach that connects clinical practice with computational methods. On the clinical side, we establish a five-level competency scheme following Miller's Pyramid, progressing from knowledge recall to dynamic case management. On the computational side, we link deductive, inductive, and abductive reasoning patterns to common medical goals and tasks. We also introduce a benchmark dataset spanning five levels of medical reasoning capability and report results on 18 state-of-the-art models, revealing that medical specialist models excel in diagnosis-centric tasks while general models lead in decision support and dialogue. We conclude by discussing current progress and open challenges, including data limitations, hallucination, and grounding issues, and outline directions toward safer, more reliable, and workflow-ready systems.

医学AI大模型临床推理综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。