构建医疗多智能体系统统一评测框架,解决模态融合与跨专科评估难题
MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems
- 统一通信协议整合11类架构、24种医学模态
- 零样本语义评估器利用视觉语言模型验证诊断逻辑
- 覆盖11个器官系统473种疾病,揭示领域迁移脆弱性
尽管多智能体系统在复杂临床决策支持中展现潜力,但该领域仍受架构碎片化和标准化多模态整合缺失的制约。现有研究存在数据接入不统一、视觉推理评估不一致及跨专科基准缺乏等问题。为此,我们提出MedMASLab——一个面向多模态医疗多智能体系统的统一框架与评测平台。该平台引入:(1) 标准化多模态智能体通信协议,实现11类异构架构在24种医学模态下的无缝集成;(2) 自动化临床推理评估器,采用零样本语义评估范式,借助大视觉语言模型验证诊断逻辑与视觉定位;(3) 当前最全面的基准,涵盖11个器官系统与473种疾病,标准化来自11个临床基准的数据。系统性评估发现:尽管多智能体系统提升推理深度,但当前架构在跨专业子领域间表现出显著脆弱性。我们对交互机制进行了严谨消融分析,并揭示了性能与成本权衡关系,为未来自主临床系统建立新的技术基线。源代码与数据已公开:https://github.com/NUS-Project/MedMASLab/
原文摘要 · Abstract (English)
While Multi-Agent Systems (MAS) show potential for complex clinical decision support, the field remains hindered by architectural fragmentation and the lack of standardized multimodal integration. Current medical MAS research suffers from non-uniform data ingestion pipelines, inconsistent visual-reasoning evaluation, and a lack of cross-specialty benchmarking. To address these challenges, we present MedMASLab, a unified framework and benchmarking platform for multimodal medical multi-agent systems. MedMASLab introduces: (1) A standardized multimodal agent communication protocol that enables seamless integration of 11 heterogeneous MAS architectures across 24 medical modalities. (2) An automated clinical reasoning evaluator, a zero-shot semantic evaluation paradigm that overcomes the limitations of lexical string-matching by leveraging large vision-language models to verify diagnostic logic and visual grounding. (3) The most extensive benchmark to date, spanning 11 organ systems and 473 diseases, standardizing data from 11 clinical benchmarks. Our systematic evaluation reveals a critical domain-specific performance gap: while MAS improves reasoning depth, current architectures exhibit significant fragility when transitioning between specialized medical sub-domains. We provide a rigorous ablation of interaction mechanisms and cost-performance trade-offs, establishing a new technical baseline for future autonomous clinical systems. The source code and data is publicly available at: https://github.com/NUS-Project/MedMASLab/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。