构建脑肿瘤MRI问答数据集,评估视觉语言模型在医学影像中的表现
UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation

- 基于473例胶质瘤MRI构建2387个问答对,覆盖多序列3D影像
- 现有模型在多序列3D MRI上表现差,过度依赖语言先验导致模态崩溃
- 为神经肿瘤影像的智能辅助诊断提供评测基准,适合医疗AI研究者
脑肿瘤诊断严重依赖磁共振成像(MRI)评估,需放射科医师综合分析数千张跨多维序列和纵向研究的图像,该过程要求高水平神经放射学训练,认知负荷大且耗时。随着放射科需求增加,该专业人才难以规模化,加剧了当前医疗系统的压力。视觉-语言模型(VLMs)有望通过半自动化、交互式方式减轻负担。然而,由于缺乏针对神经肿瘤领域的专用评测基准,其应用仍受限。本文提出一个临床相关的视觉问答(VQA)基准——UCSF-PDGM-VQA数据集,包含来自公共UCSF-PDGM数据集的473例胶质瘤相关MRI研究中的2,387个问答对。同时,我们为六种先进的视觉语言模型(VLMs)和一种大型语言模型建立了性能基线。结果显示,当前模型无法有效处理多序列、三维MRI扫描,导致视觉特征被抑制,过度依赖语言先验,引发模态崩溃。这些发现凸显了现有模型在临床场景中的可靠性与安全性缺陷,亟需开发鲁棒的领域特定视觉语言模型。
原文摘要 · Abstract (English)
Brain tumor diagnosis is largely dependent on Magnetic Resonance Imaging (MRI) evaluation, which requires radiologists to synthesize thousands of images across multiple 3D sequences and longitudinal studies. This process requires advanced neuro-radiology training, poses substantial cognitive load, and is highly time-consuming. Despite increasing demands in radiology, this expertise is difficult to scale, straining the current health systems. Vision-Language Models (VLMs) provide an opportunity to reduce this burden through a semi-automated, interactive interpretation of complex brain MRIs. However, they are currently underutilized in neuro-oncology due to a lack of specialized benchmarks for evaluating them. We introduce a clinically relevant visual question answering (VQA) benchmark -- the UCSF-PDGM-VQA dataset -- consisting of 2,387 QA pairs from 473 glioma-related MRI studies in the public UCSF-PDGM dataset. We further establish a performance baseline for six state-of-the-art vision-language models (VLMs) and one large language model on this dataset. We find that current models are incapable of effectively processing multi-sequence, 3-dimensional MRI scans, thus resulting in a suppression of visual features and over-reliance on language priors, causing modality collapse. These findings underscore a critical deficiency in current model reliability and safety within clinical settings, necessitating the development of robust, domain-specific VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。