arXiv:2508.00669cs.CLcs.AI2025-08综述被引 14

系统梳理医学大模型推理增强技术,揭示其在诊断与治疗中的应用潜力。

Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications

  • 按训练与测试阶段分类,归纳推理增强方法
  • 基于60项研究发现,多模态推理仍存短板
  • 适合医疗AI研发者与临床决策系统设计者参考

大型语言模型(LLMs)在医学领域的广泛应用带来了显著能力提升,但其在系统性、透明性和可验证性推理方面仍存在关键短板,而这正是临床实践的核心。这一问题推动了从单步答案生成向专为医学推理设计的模型转变。本文首次系统性综述该新兴领域,提出一种推理增强技术分类体系,涵盖训练阶段策略(如监督微调、强化学习)和测试阶段机制(如提示工程、多智能体系统)。分析这些技术在文本、图像、代码等多模态数据上的应用,以及在诊断、医学教育和治疗方案制定中的具体场景。同时,考察评估基准从单一准确率演变为对推理质量与视觉可解释性的综合评测。基于2022-2025年60篇核心研究的分析,指出当前面临的主要挑战,包括真实性与合理性之间的差距,以及对原生多模态推理的需求,并提出未来需构建高效、稳健且符合社会技术责任的医疗AI方向。

原文摘要 · Abstract (English)

The proliferation of Large Language Models (LLMs) in medicine has enabled impressive capabilities, yet a critical gap remains in their ability to perform systematic, transparent, and verifiable reasoning, a cornerstone of clinical practice. This has catalyzed a shift from single-step answer generation to the development of LLMs explicitly designed for medical reasoning. This paper provides the first systematic review of this emerging field. We propose a taxonomy of reasoning enhancement techniques, categorized into training-time strategies (e.g., supervised fine-tuning, reinforcement learning) and test-time mechanisms (e.g., prompt engineering, multi-agent systems). We analyze how these techniques are applied across different data modalities (text, image, code) and in key clinical applications such as diagnosis, education, and treatment planning. Furthermore, we survey the evolution of evaluation benchmarks from simple accuracy metrics to sophisticated assessments of reasoning quality and visual interpretability. Based on an analysis of 60 seminal studies from 2022-2025, we conclude by identifying critical challenges, including the faithfulness-plausibility gap and the need for native multimodal reasoning, and outlining future directions toward building efficient, robust, and sociotechnically responsible medical AI.

医学AI大模型推理增强系统综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。