arXiv:2506.06574cs.AIcs.MA2025-06被引 4

临床AI多智能体系统中,组件优化反而导致整体诊断准确率下降。

The Optimization Paradox in Clinical AI Multi-Agent Systems

  • 将诊断拆解为信息获取、解读和鉴别三阶段,用专精模型分工
  • 组件最优系统诊断准确率仅67.7%,低于顶尖多智能体系统的77.4%
  • 强调系统级整合与信息流兼容性,而非只看单个组件表现

多智能体人工智能系统在临床场景中日益普及,但其组件级优化与系统整体性能之间的关系仍不清晰。我们基于MIMIC-CDM数据集中的2,400例真实患者病例,针对四种腹部疾病(阑尾炎、胰腺炎、胆囊炎、憩室炎)评估了这一关系,将临床诊断分解为信息获取、解释与鉴别诊断三个阶段。通过综合指标(诊断结果、流程合规性、成本效率)对比单一模型系统与多智能体系统。结果显示存在悖论:尽管多智能体系统整体表现更优,但由各任务最佳组件构成的‘最佳单品’系统虽在流程上表现优异(信息准确率85.5%),诊断准确率却显著偏低(67.7%),远低于顶尖多智能体系统(77.4%)。这表明医疗AI成功集成不仅需组件优化,更需关注智能体间的信息传递与兼容性。研究强调应进行端到端系统验证,而非仅依赖组件性能指标。

原文摘要 · Abstract (English)

Multi-agent artificial intelligence systems are increasingly deployed in clinical settings, yet the relationship between component-level optimization and system-wide performance remains poorly understood. We evaluated this relationship using 2,400 real patient cases from the MIMIC-CDM dataset across four abdominal pathologies (appendicitis, pancreatitis, cholecystitis, diverticulitis), decomposing clinical diagnosis into information gathering, interpretation, and differential diagnosis. We evaluated single agent systems (one model performing all tasks) against multi-agent systems (specialized models for each task) using comprehensive metrics spanning diagnostic outcomes, process adherence, and cost efficiency. Our results reveal a paradox: while multi-agent systems generally outperformed single agents, the component-optimized or Best of Breed system with superior components and excellent process metrics (85.5% information accuracy) significantly underperformed in diagnostic accuracy (67.7% vs. 77.4% for a top multi-agent system). This finding underscores that successful integration of AI in healthcare requires not just component level optimization but also attention to information flow and compatibility between agents. Our findings highlight the need for end to end system validation rather than relying on component metrics alone.

临床AI多智能体系统优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。