arXiv:2603.04421cs.CLcs.AI2026-03中稿 · the EACL 2026 Work…被引 3

用不同厂商的AI医生协作,诊断更准更稳。

Do Mixed-Vendor Multi-Agent LLMs Improve Clinical Diagnosis?

  • 混用不同厂商的AI医生,避免共同错误。
  • 在罕见病和诊断测试中准确率、召回率均领先。
  • 适合追求高可靠性的医疗AI系统设计者。

多智能体大语言模型系统在临床诊断中展现出潜力,通过智能体协作优化医学推理。然而,现有框架多依赖同厂商团队(如同一模型家族的多个代理),易导致相关性故障模式,强化共享偏见而非纠正。本文对比了单模型、单厂商与混合厂商多智能体对话(MAC)框架。使用o4-mini、Gemini-2.5-Pro和Claude-4.5-Sonnet三名医生代理,在RareBench和DiagnosisArena上评估性能。混合厂商配置持续优于单厂商版本,实现当前最佳的召回率与准确率。重叠分析揭示其机制:异构团队融合互补归纳偏见,挖掘出单个模型或同质团队共同遗漏的正确诊断。结果表明,厂商多样性是构建鲁棒临床诊断系统的关键设计原则。

原文摘要 · Abstract (English)

Multi-agent large language model (LLM) systems have emerged as a promising approach for clinical diagnosis, leveraging collaboration among agents to refine medical reasoning. However, most existing frameworks rely on single-vendor teams (e.g., multiple agents from the same model family), which risk correlated failure modes that reinforce shared biases rather than correcting them. We investigate the impact of vendor diversity by comparing Single-LLM, Single-Vendor, and Mixed-Vendor Multi-Agent Conversation (MAC) frameworks. Using three doctor agents instantiated with o4-mini, Gemini-2.5-Pro, and Claude-4.5-Sonnet, we evaluate performance on RareBench and DiagnosisArena. Mixed-vendor configurations consistently outperform single-vendor counterparts, achieving state-of-the-art recall and accuracy. Overlap analysis reveals the underlying mechanism: mixed-vendor teams pool complementary inductive biases, surfacing correct diagnoses that individual models or homogeneous teams collectively miss. These results highlight vendor diversity as a key design principle for robust clinical diagnostic systems.

多智能体医疗AI模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。