让AI医生像人一样不断积累诊断经验,提升真实场景下的诊疗准确率。
MACD: Multi-Agent Clinical Diagnosis with Self-Learned Knowledge for LLM
- 构建多智能体系统,让LLM自动总结、优化并应用临床诊断知识。
- 在4390例真实病例上,诊断准确率平均提升11.6个百分点。
- 人机协同流程使文本病例诊断准确率比纯医生高出18.3个百分点。
大型语言模型(LLMs)在辅助医疗诊断中展现出潜力,基于提示的方法灵活且易于部署。然而,现有提示工程与多智能体方法多聚焦于单次推理优化,忽视从临床实践中积累可复用的经验,限制了其实际应用。为此,本文提出一种新型多智能体临床诊断框架(MACD),通过多智能体流程实现LLM自我学习临床知识,模拟人类医师的专业成长。进一步拓展为MACD-人机协作工作流,多个基于LLM的诊断智能体在裁判智能体和人类监督下进行迭代讨论,解决意见不一致问题。构建了包含4,390例真实患者病例的MIMIC-MACD数据集,涵盖七种疾病,其中1,314例用于知识学习,3,076例用于评估。在多种开源权重LLM上,MACD显著提升主要诊断准确率,平均优于权威知识11.6个百分点,并缩小了开源模型与顶尖LLM之间的性能差距。此外,该人机协作流程在仅含文本的病例中,相比纯医生诊断提升18.3个百分点,彰显了人机协同的潜力。本研究提出了一种可扩展的自学习范式,弥合了LLM内在知识与真实临床需求之间的差距,推动可靠、可解释、可部署的AI辅助诊断发展。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promise in supporting medical diagnosis, with prompting-based methods offering a flexible and deployable means of capability enhancement. However, existing prompt engineering and multi-agent approaches often focus on optimizing single inferences, paying less attention to the accumulation of reusable experience from clinical practice, constraining their real-world applicability. To address this, this study proposes a novel Multi-Agent Clinical Diagnosis (MACD) framework, which allows LLMs to self-learn clinical knowledge via a multi-agent pipeline that summarizes, refines, and applies diagnostic insights, mirroring the professional development of human physicians. We further extend it to a MACD-human collaborative workflow, where multiple LLM-based diagnostician agents engage in iterative consultations, supported by a judge agent and human oversight for cases where agreement is not reached. The MIMIC-MACD cohort comprising 4,390 real-world patient cases across seven diseases is constructed, including 1,314 cases for knowledge learning and 3,076 held-out cases for evaluation. Across diverse open-weight LLMs, MACD significantly improves primary diagnostic accuracy, achieving an average improvement of 11.6 percentage points over established authoritative knowledge, while narrowing the performance gap between open-weight models and state-of-the-art LLMs. Furthermore, the MACD-human workflow yields an 18.3-percentage-point improvement over physician-only diagnosis on text-only vignettes, demonstrating the synergistic potential of human-AI collaboration. This work thus presents a scalable self-learning paradigm that bridges the gap between the intrinsic knowledge of LLMs and the demands of real-world clinical practice, advancing towards a reliable, interpretable, and deployable AI-assisted diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。