arXiv:2608.22323cs.CV2026-08

评测大模型在临床多模态诊断合成能力,发现多数模型准确率低于50%。

MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

论文配图:MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis
图 1 · 摘自论文原文
  • 构建包含625例专家验证病例的多模态诊断基准,融合病史与医学影像
  • 平均每例含2.79张图像,共1042个基于ICD-11编码的诊断标注
  • 揭示医疗知识、图像理解与证据整合能力对诊断性能的关键影响

大型语言模型(LLM)在诊断决策中的应用日益受到关注。然而,现有基准主要聚焦于文本推理或孤立的视觉问答任务,缺乏对临床叙事与医学影像的全面整合,无法评估专家级临床判断所依赖的多模态诊断合成能力。为弥补这一缺口,我们提出MedReaMM,一个专门用于评估模型将详细病史与多张医学影像融合,生成准确鉴别诊断能力的基准。该数据集源自顶级医学期刊和精选临床案例库,包含625个专家验证病例,平均每个病例有2.79张医学影像,共计1042个标准化诊断,以ICD-11编码标注。这些病例主要涵盖罕见、非典型或多系统表现,需超越常规模式识别的专家级证据整合。我们评估了23个大型多模态模型(LMMs),发现多数模型诊断准确率低于50%,凸显其在多模态诊断合成方面存在显著差距。进一步分析表明,医疗知识掌握度、医学图像理解力以及证据整合能力均与诊断表现高度相关。

原文摘要 · Abstract (English)

The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus on textual reasoning or isolated visual question-answering (VQA) tasks, lacking holistic integration of clinical narratives and medical imaging, and thus failing to assess the multimodal diagnostic synthesis capability central to expert clinical judgment. To bridge this gap, we introduce MedReaMM, a benchmark specifically designed to evaluate models' ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm. Constructed from case reports sourced from top-tier medical journals and curated clinical case databases, MedReaMM comprises 625 expert-validated cases with an average of 2.79 medical images per case and a total of 1,042 standardized diagnoses annotated with ICD-11 codes. These cases predominantly represent rare, atypical, or multi-system presentations that demand expert-level evidence integration beyond routine pattern recognition. We evaluate 23 Large Multimodal Models (LMMs) and find that most achieve diagnostic accuracy scores below 50%, underscoring a substantial gap in multimodal diagnostic synthesis capability. Further analysis reveals that medical knowledge proficiency, medical image understanding, and evidence integration are all highly correlated with diagnostic performance.

多模态模型医学诊断大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。