arXiv:2508.13754cs.AI2025-08被引 1

让多个大模型按专业分工协作,提升医疗决策准确率。

Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making

  • 根据医学领域和问题难度,动态挑选最适合的模型当专家
  • 通过自评置信度与对抗验证,融合多模型结果提升可靠性
  • 在真实医疗数据集上表现优于顶尖单模型,适合临床辅助系统

医疗决策(MDM)需整合复杂多样的临床信息,依赖深厚的专业知识。尽管大语言模型(LLMs)在支持MDM方面展现潜力,但单模型方法受限于参数知识范围和静态训练数据,难以有效融合临床信息。为此,我们提出专家感知的多模型招募与协作框架(EMRC),分两阶段运行:(i)基于公开语料构建模型能力表,按医学科室与问题难度动态选择最优模型作为专家代理;(ii)通过自评置信度与对抗验证实现多代理协作融合,提升诊断可靠性。我们在三个公开的MDM数据集上评估,结果表明EMRC优于现有单模型与多模型方法。例如,在MMLU-Pro-Health数据集上达到74.45%准确率,较最佳闭源模型GPT-4-0613提升2.69%,验证了专家感知招募策略与模型能力互补的有效性。

原文摘要 · Abstract (English)

Medical Decision-Making (MDM) is a complex process requiring substantial domain-specific expertise to effectively synthesize heterogeneous and complicated clinical information. While recent advancements in Large Language Models (LLMs) show promise in supporting MDM, single-LLM approaches are limited by their parametric knowledge constraints and static training corpora, failing to robustly integrate the clinical information. To address this challenge, we propose the Expertise-aware Multi-LLM Recruitment and Collaboration (EMRC) framework to enhance the accuracy and reliability of MDM systems. It operates in two stages: (i) expertise-aware agent recruitment and (ii) confidence- and adversarial-driven multi-agent collaboration. Specifically, in the first stage, we use a publicly available corpus to construct an LLM expertise table for capturing expertise-specific strengths of multiple LLMs across medical department categories and query difficulty levels. This table enables the subsequent dynamic selection of the optimal LLMs to act as medical expert agents for each medical query during the inference phase. In the second stage, we employ selected agents to generate responses with self-assessed confidence scores, which are then integrated through the confidence fusion and adversarial validation to improve diagnostic reliability. We evaluate our EMRC framework on three public MDM datasets, where the results demonstrate that our EMRC outperforms state-of-the-art single- and multi-LLM methods, achieving superior diagnostic performance. For instance, on the MMLU-Pro-Health dataset, our EMRC achieves 74.45% accuracy, representing a 2.69% improvement over the best-performing closed-source model GPT- 4-0613, which demonstrates the effectiveness of our expertise-aware agent recruitment strategy and the agent complementarity in leveraging each LLM's specialized capabilities.

医疗AI多模型协作大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。