用强化学习动态调度专家,提升多模态医学诊断准确率
MedRoute: RL-Based Dynamic Specialist Routing in Multi-Agent Medical Diagnosis
- 引入强化学习路由器动态选择医疗专家代理
- 在多模态数据集上诊断准确率超越现有最优方法
- 适合研究智能诊疗系统与多代理协作的学者
基于大模型的多模态医学诊断受到广泛关注,因其能结合文本与图像输入生成精准诊断。然而,这些模型普遍过于通用,在真实医疗场景中难以应对多样化病情。临床实践中,诊断由多位具有领域专长的专家协同完成。为模拟此过程,我们提出MedRoute,一种灵活的动态多代理框架,包含多个专业医疗大模型代理。此外,引入经强化学习训练的全科医生作为路由器,动态选择合适专家,并由调解者生成最终决策。该框架更贴近真实临床流程。在基于文本和图像的医疗数据集上的大量评估表明,其诊断准确率显著优于当前最先进基线。相关代码与模型已开源。
原文摘要 · Abstract (English)
Medical diagnosis using Large Multimodal Models (LMMs) has gained increasing attention due to capability of these models in providing precise diagnoses. These models generally combine medical questions with visual inputs to generate diagnoses or treatments. However, they are often overly general and unsuitable under the wide range of medical conditions in real-world healthcare. In clinical practice, diagnosis is performed by multiple specialists, each contributing domain-specific expertise. To emulate this process, a potential solution is to deploy a dynamic multi-agent LMM framework, where each agent functions as a medical specialist. Current approaches in this emerging area, typically relying on static or predefined selection of various specialists, cannot be adapted to the changing practical scenario. In this paper, we propose MedRoute, a flexible and dynamic multi-agent framework that comprises of a collaborative system of specialist LMM agents. Furthermore, we add a General Practitioner with an RL-trained router for dynamic specialist selection, and a Moderator that produces the final decision. In this way, our framework closely mirrors real clinical workflows. Extensive evaluations on text and image-based medical datasets demonstrate improved diagnostic accuracy, outperforming the state-of-the-art baselines. Our work lays a strong foundation for future research. Code and models are available at https://github.com/UCF-CRCV/MedRoute/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。