检测医疗多智能体AI协作中的虚假共识,揭示其决策风险
Auditing medical multi-agent AI reveals risks of false consensus
- 构建流程审计框架,识别协作中的十大失效模式
- 16.63%案例存在无依据陈述,42.73%未激活专业推理
- 适合医疗AI安全评估与临床监督系统研发者参考
大型语言模型正被整合为模拟多学科会诊的医疗多智能体系统,但仅关注最终结果准确率不足以保障临床安全。本文提出MedAgentAudit——一个基于临床实践的流程审计框架,从3600个执行日志中提炼出十类典型协作失败模式,涵盖任务理解、讨论与决策合成阶段。通过部署专家验证的自动化审计工具,在14400个案例中测试六种架构、六组医学文本与视觉数据集及每模态四种大模型设置。结果显示:16.63%案例存在无依据陈述并传播至下游;98.42%讨论中智能体重复初始观点而非重审证据;42.73%未能激活专业推理;28.76%存在权威偏倚(跨轮次升至68.75%),18.53%出现自相矛盾,5.48%忽视矛盾,5.11%压制少数意见。该框架将医疗AI评估从输出评分转向过程安全与可问责性,为透明、可审计、由临床监督的智能体系统提供实践基础。
原文摘要 · Abstract (English)
Large language models are increasingly being assembled into medical multi-agent systems that emulate multidisciplinary consultation through specialist roles, peer review and consensus formation. In clinical decision support, however, apparent consensus is not enough. Clinicians also need to know whether agents checked the evidence, addressed disagreement and kept uncertainty visible. Current evaluations largely score final accuracy, leaving the safety of the collaborative process untested. Here we introduce MedAgentAudit, a clinically grounded workflow audit framework for diagnosing and quantifying collaborative failure modes in medical multi-agent systems. From 3,600 execution logs, we derive an expert-validated taxonomy of ten recurrent failures spanning task comprehension, collaborative discussion, and synthesis and decision-making. We then deploy an expert-validated automated auditor as non-interventional probes across 14,400 cases, covering six multi-agent architectures, six medical text and vision datasets, and four large language model settings per modality. Across systems, collaboration yields uneven accuracy gains and frequent process failures. Unsupported observations affect 16.63% of cases and propagate downstream. In discussion, agents repeat initial views in 98.42% of cases rather than re-examining evidence, and fail to activate specialist reasoning in 42.73%. During synthesis, final answers often substitute authority or majority count for evidence checking, showing authority bias in 28.76% (rising from 35.30% to 68.75% across rounds), self-contradiction in 18.53%, contradiction neglect in 5.48% and minority suppression in 5.11%. MedAgentAudit reframes medical AI evaluation from output scoring to process-level safety and accountability, providing a practical foundation for transparent, auditable and clinician-supervised agentic systems in medicine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。