arXiv:2506.07400cs.MAcs.AI2025-06被引 13

用多个角色智能体协作,提升眼科诊断的准确与可信度。

MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models

  • 设计多智能体架构,分工处理影像与报告
  • 减少幻觉,提升医疗推理可靠性
  • 适合临床医生审阅与医学教育使用

将基于深度学习的青光眼检测与大语言模型(LLMs)结合,可自动化缓解眼科医生短缺问题并提升临床报告效率。然而,将通用大模型应用于医学影像仍面临幻觉、可解释性差及领域知识不足等挑战,可能降低临床准确性。尽管近期融合影像模型与大模型推理的方法改善了报告质量,但通常依赖单一通用智能体,难以模拟多学科团队的复杂推理过程。为此,我们提出 MedChat,一个由专业视觉模型与多个角色专用大模型智能体组成的多智能体诊断框架,由总控智能体协调。该设计提升了可靠性,降低了幻觉风险,并通过面向临床评审与教学的交互界面支持诊断报告生成。代码已开源:https://github.com/Purdue-M2/MedChat。

原文摘要 · Abstract (English)

The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologist shortages and improve clinical reporting efficiency. However, applying general LLMs to medical imaging remains challenging due to hallucinations, limited interpretability, and insufficient domain-specific medical knowledge, which can potentially reduce clinical accuracy. Although recent approaches combining imaging models with LLM reasoning have improved reporting, they typically rely on a single generalist agent, restricting their capacity to emulate the diverse and complex reasoning found in multidisciplinary medical teams. To address these limitations, we propose MedChat, a multi-agent diagnostic framework and platform that combines specialized vision models with multiple role-specific LLM agents, all coordinated by a director agent. This design enhances reliability, reduces hallucination risk, and enables interactive diagnostic reporting through an interface tailored for clinical review and educational use. Code available at https://github.com/Purdue-M2/MedChat.

多智能体医学影像大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。