arXiv:2602.21950cs.CL2026-02ACL

评测大模型在复杂临床病例中的多证据综合能力,发现其诊断合成存在明显短板。

MEDSYN: Benchmarking Multi-EviDence SYNthesis in Complex Clinical Cases for Multimodal Large Language Models

  • 构建包含7类视觉证据的多模态临床案例基准
  • 模型在鉴别诊断上表现接近专家,但最终诊断准确率差距大
  • 揭示模型依赖文本证据、跨模态利用不均的问题,可指导优化

多模态大语言模型在医疗应用中潜力巨大,但现有评估基准未能充分反映真实临床复杂性。我们提出MEDSYN,一个支持多语言、多模态的复杂临床案例基准,每例最多包含7种不同类型的视觉临床证据(CE)。参照临床工作流程,我们在18个MLLM上评估其生成鉴别诊断(DDx)和选择最终诊断(FDx)的能力。尽管顶尖模型在生成DDx方面常与甚至超过人类专家,但所有模型的DDx-FDx性能差距远大于专家,表明其在整合异质性临床证据方面存在失效模式。消融分析表明该问题源于:(i) 过度依赖判别力较弱的文本证据(如病史),(ii) 跨模态证据利用不均衡。我们引入证据敏感性(Evidence Sensitivity)量化后者,并证明较小的差距与更高诊断准确率相关。最后展示其如何用于指导模型改进。我们将开源该基准与代码。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have shown great potential in medical applications, yet existing benchmarks inadequately capture real-world clinical complexity. We introduce MEDSYN, a multilingual, multimodal benchmark of highly complex clinical cases with up to 7 distinct visual clinical evidence (CE) types per case. Mirroring clinical workflow, we evaluate 18 MLLMs on differential diagnosis (DDx) generation and final diagnosis (FDx) selection. While top models often match or even outperform human experts on DDx generation, all MLLMs exhibit a much larger DDx--FDx performance gap compared to expert clinicians, indicating a failure mode in synthesis of heterogeneous CE types. Ablations attribute this failure to (i) overreliance on less discriminative textual CE ($\it{e.g.}$, medical history) and (ii) a cross-modal CE utilization gap. We introduce Evidence Sensitivity to quantify the latter and show that a smaller gap correlates with higher diagnostic accuracy. Finally, we demonstrate how it can be used to guide interventions to improve model performance. We will open-source our benchmark and code.

多模态临床推理大模型评测医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。