医学多智能体协作效果如何?这个新基准给出实证答案。
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
- 构建涵盖四类临床任务的综合评测平台,覆盖文本、影像与病历数据。
- 多智能体在流程自动化中提升完整性,但多数任务不如单模型或传统方法。
- 强调按任务选AI方案,避免盲目堆叠复杂系统。
大型语言模型的快速发展激发了对多智能体协作解决复杂医疗任务的兴趣。然而,多智能体协作的实际优势尚不明确。现有评估缺乏泛化性,未覆盖真实临床中的多样化任务,且常忽略与单模型及成熟传统方法的严谨对比。为此,我们提出MedAgentBoard,一个用于系统评估多智能体、单模型与传统方法的综合性基准。该基准包含四类医疗任务:(1)医学(视觉)问答,(2)通俗摘要生成,(3)结构化电子健康记录(EHR)预测建模,(4)临床工作流自动化,覆盖文本、医学图像与结构化EHR数据。大量实验揭示:多智能体协作仅在特定场景(如临床工作流自动化)中展现优势,但在文本医学问答等任务上不如先进单模型,更普遍弱于专门的传统方法(如医学视觉问答与基于EHR的预测)。MedAgentBoard提供关键资源与洞察,强调必须采用任务导向、证据驱动的AI解决方案选择策略,提示多智能体固有的复杂性与开销需与实际性能收益权衡。所有代码、数据集、详细提示与实验结果已开源:https://medagentboard.netlify.app/。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has stimulated interest in multi-agent collaboration for addressing complex medical tasks. However, the practical advantages of multi-agent collaboration approaches remain insufficiently understood. Existing evaluations often lack generalizability, failing to cover diverse tasks reflective of real-world clinical practice, and frequently omit rigorous comparisons against both single-LLM-based and established conventional methods. To address this critical gap, we introduce MedAgentBoard, a comprehensive benchmark for the systematic evaluation of multi-agent collaboration, single-LLM, and conventional approaches. MedAgentBoard encompasses four diverse medical task categories: (1) medical (visual) question answering, (2) lay summary generation, (3) structured Electronic Health Record (EHR) predictive modeling, and (4) clinical workflow automation, across text, medical images, and structured EHR data. Our extensive experiments reveal a nuanced landscape: while multi-agent collaboration demonstrates benefits in specific scenarios, such as enhancing task completeness in clinical workflow automation, it does not consistently outperform advanced single LLMs (e.g., in textual medical QA) or, critically, specialized conventional methods that generally maintain better performance in tasks like medical VQA and EHR-based prediction. MedAgentBoard offers a vital resource and actionable insights, emphasizing the necessity of a task-specific, evidence-based approach to selecting and developing AI solutions in medicine. It underscores that the inherent complexity and overhead of multi-agent collaboration must be carefully weighed against tangible performance gains. All code, datasets, detailed prompts, and experimental results are open-sourced at https://medagentboard.netlify.app/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。