用多角色辩论框架提升大模型对复杂病症的诊断准确率
A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support

- 设计多智能体辩论结构,模拟医生反复讨论的诊断流程
- 在罕见病和疑难病例上诊断准确率提升10.21个百分点
- 适合医疗AI研发者及临床辅助系统设计人员参考
大型语言模型在医学任务中展现潜力,但其单轮问答模式无法反映实际临床诊断过程,限制了其在复杂诊断场景中的应用。为此,我们提出基于角色分工的多智能体辩论框架(DMoA),支持迭代式诊断推理。在297个罕见病案例和1,719个疑难病例上评估了基础模型与DMoA的表现。结果显示,相较于GPT-4o基线,DMoA在两个数据集上分别将最可能诊断准确率提升10.21个百分点,安全率提升11.36个百分点。消融实验表明,性能提升不仅源于更多模型或更长输出,更得益于结构化工作流的设计。进一步分析显示,四组两成员结构、更强基础模型和更大令牌预算均有助于性能提升。研究验证了DMoA在临床任务中的潜力,提示需深入探索多智能体框架的应用价值。
原文摘要 · Abstract (English)
Large language models (LLMs) show potential for medical tasks, but their single-turn question-answer format does not reflect how clinical diagnosis is performed in practice. As a result, they remain limited in complex diagnostic settings. We developed Debate-Mixture-of-Agents (DMoA), a novel multi-agent framework that structures role-based interaction to support iterative diagnostic reasoning. Base models and DMoA were evaluated on 297 rare disease cases and 1,719 challenging cases. Across both datasets, DMoA improved most likely diagnosis accuracy by 10.21 percentage points and safety rate by 11.36 percentage points over GPT-4o baseline. Ablation experiments showed that the gains were not simply due to the use of more models or longer outputs, but also reflected the contribution of the structured workflow. Further analyses examined how framework design, base model choice, and token budget affected performance. DMoA performed better with a 4*2 structure, stronger base models, and a larger token budget. These findings demonstrate the potential of DMoA for clinical tasks and suggest further investigation of multi-agent frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。