构建首个大规模多轮多模态临床诊断评测基准,真实还原诊疗推理过程。
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

- 设计多轮交互式评测框架,模拟逐步披露信息的临床诊断流程。
- 15个模型在3760张医学图像上平均诊断准确率不足一半,完全正确率更低。
- 揭示五大典型错误模式,为提升模型临床推理能力指明方向。
临床诊断评估不仅应考察模型能否给出正确诊断,还应反映临床实践的真实情境,包括多模态信息的逐步披露、诊断假设的动态更新以及临床推理的持续优化。然而,现有对多模态大语言模型(MLLMs)的评估通常依赖单轮或孤立任务,难以充分捕捉现实临床诊断的复杂性。为弥合这一差距,我们构建了迄今为止规模最大的多轮多模态临床诊断评估基准 ClinMM-Bench,包含1,089个具有挑战性的真实临床案例及3,760张跨八个专科的医学影像。我们采用两级评估框架系统评估了15个代表性MLLMs,分别衡量诊断准确率与诊断推理质量。结果表明,专有模型整体诊断准确率最高,但所有模型中完全正确的诊断比例仍有限。在诊断推理质量方面,当前模型虽能识别合理的诊断方向,但在生成可靠推理路径方面仍存在显著局限。误差分析进一步识别出五种典型失败模式:信息整合失败、知识映射错误、感知错误、过早闭合和视觉幻觉。
原文摘要 · Abstract (English)
Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single-turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinical diagnosis. To bridge this gap, we developed ClinMM-Bench, the largest multi-turn multimodal clinical diagnostic evaluation benchmark to date. ClinMM-Bench contains 1,089 challenging real-world clinical cases and 3,760 medical images across eight specialties. We systematically evaluated 15 representative MLLMs using a two-level evaluation framework that assessed both diagnostic accuracy and diagnostic reasoning quality. Results showed that proprietary models achieved the highest overall diagnostic accuracy, but the proportion of completely correct diagnoses remained limited across all models. In terms of diagnostic reasoning quality, current models can identify plausible diagnostic directions but still have considerable limitations in generating reliable diagnostic reasoning. Error analysis further identified five representative failure modes: information synthesis failure, knowledge mapping error, perception error, premature closure, and visual hallucination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。