arXiv:2508.05996cs.AI2025-08被引 4

用中介代理协调多个视觉语言模型,提升医疗多模态决策能力

Mediator-Guided Multi-Agent Collaboration among Open-Source Models for Medical Decision-Making

  • 引入大语言模型作中介,协调多个视觉语言模型协作
  • 无需训练,在5个医学问答数据集上表现超越单个模型
  • 适合医疗AI研发者和多模态系统设计者参考

复杂医疗决策依赖不同临床人员的协作流程。构建AI多智能体系统可加速并增强人类水平的临床决策。现有研究多聚焦纯文本任务,向多模态场景扩展仍具挑战。盲目组合多种视觉-语言模型(VLMs)可能放大错误解读。相比同规模大语言模型(LLMs),VLMs在指令遵循和自我反思方面能力较弱,制约其协作性能。本文提出MedOrch框架,通过基于LLM的中介代理,使多个基于VLM的专家代理能够交换与反思输出以实现协作。我们采用多个开源通用及领域专用VLM,而非昂贵的GPT系列模型,凸显异构模型的优势。实验表明,不同VLM代理间的协作性能超越任一单一代理。在五个医学视觉问答基准上验证,无需模型训练即实现卓越协作表现。研究强调中介引导的多智能体协作对推进医疗多模态智能的价值。

原文摘要 · Abstract (English)

Complex medical decision-making involves cooperative workflows operated by different clinicians. Designing AI multi-agent systems can expedite and augment human-level clinical decision-making. Existing multi-agent researches primarily focus on language-only tasks, yet their extension to multimodal scenarios remains challenging. A blind combination of diverse vision-language models (VLMs) can amplify an erroneous outcome interpretation. VLMs in general are less capable in instruction following and importantly self-reflection, compared to large language models (LLMs) of comparable sizes. This disparity largely constrains VLMs' ability in cooperative workflows. In this study, we propose MedOrch, a mediator-guided multi-agent collaboration framework for medical multimodal decision-making. MedOrch employs an LLM-based mediator agent that enables multiple VLM-based expert agents to exchange and reflect on their outputs towards collaboration. We utilize multiple open-source general-purpose and domain-specific VLMs instead of costly GPT-series models, revealing the strength of heterogeneous models. We show that the collaboration within distinct VLM-based agents can surpass the capabilities of any individual agent. We validate our approach on five medical vision question answering benchmarks, demonstrating superior collaboration performance without model training. Our findings underscore the value of mediator-guided multi-agent collaboration in advancing medical multimodal intelligence.

多智能体医疗AI视觉语言模型协作决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。