首个评估视觉语言模型是否像医生一样推理的医学影像诊断基准
DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?
- 构建涵盖20类任务、5种影像模态的多阶段临床推理评估框架
- 19个模型在复杂推理任务中性能显著下降,多数依赖表面关联而非真实理解
- 适合关注医疗AI可解释性与临床可信度的研究者使用
视觉语言模型(VLMs)在自然图像上表现出强大的零样本泛化能力,并在可解释的医学图像分析中展现出早期潜力。然而,现有基准未能系统评估这些模型是否真正像临床医生一样进行推理,还是仅模仿表层模式。为此,我们提出DrVD-Bench,首个针对临床视觉推理的多模态基准。该基准包含三个模块:视觉证据理解、推理轨迹评估和报告生成评估,共7,789对图像-问题数据,覆盖20种任务类型、17种诊断类别及五种成像模态——CT、MRI、超声、放射摄影和病理学。基准设计严格反映从模态识别到病灶定位再到诊断的临床推理流程。我们评估了19个VLMs,包括通用型与医学专用型、开源与专有模型,发现随着推理复杂度提升,性能急剧下降。尽管部分模型初现人类式推理痕迹,但仍常依赖捷径关联而非基于视觉的扎实理解。DrVD-Bench为开发临床可信的VLMs提供了严谨且结构化的评估框架。
原文摘要 · Abstract (English)
Vision-language models (VLMs) exhibit strong zero-shot generalization on natural images and show early promise in interpretable medical image analysis. However, existing benchmarks do not systematically evaluate whether these models truly reason like human clinicians or merely imitate superficial patterns. To address this gap, we propose DrVD-Bench, the first multimodal benchmark for clinical visual reasoning. DrVD-Bench consists of three modules: Visual Evidence Comprehension, Reasoning Trajectory Assessment, and Report Generation Evaluation, comprising a total of 7,789 image-question pairs. Our benchmark covers 20 task types, 17 diagnostic categories, and five imaging modalities-CT, MRI, ultrasound, radiography, and pathology. DrVD-Bench is explicitly structured to reflect the clinical reasoning workflow from modality recognition to lesion identification and diagnosis. We benchmark 19 VLMs, including general-purpose and medical-specific, open-source and proprietary models, and observe that performance drops sharply as reasoning complexity increases. While some models begin to exhibit traces of human-like reasoning, they often still rely on shortcut correlations rather than grounded visual understanding. DrVD-Bench offers a rigorous and structured evaluation framework to guide the development of clinically trustworthy VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。