arXiv:2507.05520cs.AI2025-07被引 2

用多智能体协作模拟医生会诊,提升医学影像问答准确率

Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQA

  • 设计类临床多智能体系统,融合多方视角推理与文献检索
  • 最高达70%准确率,且在未见数据上表现稳定
  • 适合需要可解释性与循证支持的医疗AI落地场景

远程皮肤科诊疗常因缺乏面对面问诊的丰富背景而受限,临床医生仅凭少量图像和简短描述进行诊断,缺少体格检查、第二意见或参考文献支持。尽管许多医疗AI系统通过领域微调弥补这一缺陷,本研究提出:模仿临床推理过程可能更有效。实验在六种配置下测试了七种视觉语言模型:基础模型、微调版本,以及分别引入多视角推理层(类比同行会诊)或检索增强生成(类比查阅文献)的改进版本。结果发现,微调使七种模型中有四种性能下降,平均降幅达30%;基础模型在测试集上严重崩溃。而受临床启发的架构实现了最高70%的准确率,在未见数据上保持稳定,并生成可解释、基于文献的输出,对临床采纳至关重要。研究表明,医疗AI的成功在于重建临床诊断中协作与循证的核心实践。

原文摘要 · Abstract (English)

Dermatological care via telemedicine often lacks the rich context of in-person visits. Clinicians must make diagnoses based on a handful of images and brief descriptions, without the benefit of physical exams, second opinions, or reference materials. While many medical AI systems attempt to bridge these gaps with domain-specific fine-tuning, this work hypothesized that mimicking clinical reasoning processes could offer a more effective path forward. This study tested seven vision-language models on medical visual question answering across six configurations: baseline models, fine-tuned variants, and both augmented with either reasoning layers that combine multiple model perspectives, analogous to peer consultation, or retrieval-augmented generation that incorporates medical literature at inference time, serving a role similar to reference-checking. While fine-tuning degraded performance in four of seven models with an average 30% decrease, baseline models collapsed on test data. Clinical-inspired architectures, meanwhile, achieved up to 70% accuracy, maintaining performance on unseen data while generating explainable, literature-grounded outputs critical for clinical adoption. These findings demonstrate that medical AI succeeds by reconstructing the collaborative and evidence-based practices fundamental to clinical diagnosis.

医学AI多智能体视觉问答可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。