构建脑部MRI罕见病异常检测新基准,测试模型对未知病症的识别与推理能力。
NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI
- 提出面向罕见病的脑部MRI评估基准NOVA,涵盖281种罕见病理。
- 在真实临床数据上,主流视觉语言模型性能显著下降,暴露其泛化缺陷。
- 适合关注医疗AI鲁棒性、开放世界识别的研究者使用。
在实际应用中,部署模型常遇到训练时未见的数据分布。已有方法多聚焦于已知异常类型,难以评估模型对真正未知病症的应对能力。为此,我们提出NOVA——一个包含约900例脑部MRI扫描的评估基准,覆盖281种罕见病灶及多样化的成像协议。每例均配有详尽临床描述和双盲专家标注的边界框。该数据集支持异常定位、图像描述生成与诊断推理的联合评估。由于不用于训练,NOVA成为极端条件下的分布外泛化测试平台:模型需跨越样本外观与语义空间双重差异。基线实验显示GPT-4o、Gemini 2.0 Flash、Qwen2.5-VL-72B等主流模型在各项任务中表现大幅下滑,验证了其作为严苛测试基准的有效性。
原文摘要 · Abstract (English)
In many real-world applications, deployed models encounter inputs that differ from the data seen during training. Out-of-distribution detection identifies whether an input stems from an unseen distribution, while open-world recognition flags such inputs to ensure the system remains robust as ever-emerging, previously $unknown$ categories appear and must be addressed without retraining. Foundation and vision-language models are pre-trained on large and diverse datasets with the expectation of broad generalization across domains, including medical imaging. However, benchmarking these models on test sets with only a few common outlier types silently collapses the evaluation back to a closed-set problem, masking failures on rare or truly novel conditions encountered in clinical use. We therefore present $NOVA$, a challenging, real-life $evaluation-only$ benchmark of $\sim$900 brain MRI scans that span 281 rare pathologies and heterogeneous acquisition protocols. Each case includes rich clinical narratives and double-blinded expert bounding-box annotations. Together, these enable joint assessment of anomaly localisation, visual captioning, and diagnostic reasoning. Because NOVA is never used for training, it serves as an $extreme$ stress-test of out-of-distribution generalisation: models must bridge a distribution gap both in sample appearance and in semantic space. Baseline results with leading vision-language models (GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B) reveal substantial performance drops across all tasks, establishing NOVA as a rigorous testbed for advancing models that can detect, localize, and reason about truly unknown anomalies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。