arXiv:2505.14064eess.IVcs.AI2025-05被引 16

构建脑部MRI罕见病异常检测新基准,测试模型对未知病症的识别与推理能力。

NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI

  • 提出面向罕见病的脑部MRI评估基准NOVA,涵盖281种罕见病理。
  • 在真实临床数据上,主流视觉语言模型性能显著下降,暴露其泛化缺陷。
  • 适合关注医疗AI鲁棒性、开放世界识别的研究者使用。

在实际应用中,部署模型常遇到训练时未见的数据分布。已有方法多聚焦于已知异常类型,难以评估模型对真正未知病症的应对能力。为此,我们提出NOVA——一个包含约900例脑部MRI扫描的评估基准,覆盖281种罕见病灶及多样化的成像协议。每例均配有详尽临床描述和双盲专家标注的边界框。该数据集支持异常定位、图像描述生成与诊断推理的联合评估。由于不用于训练,NOVA成为极端条件下的分布外泛化测试平台:模型需跨越样本外观与语义空间双重差异。基线实验显示GPT-4o、Gemini 2.0 Flash、Qwen2.5-VL-72B等主流模型在各项任务中表现大幅下滑,验证了其作为严苛测试基准的有效性。

原文摘要 · Abstract (English)

In many real-world applications, deployed models encounter inputs that differ from the data seen during training. Out-of-distribution detection identifies whether an input stems from an unseen distribution, while open-world recognition flags such inputs to ensure the system remains robust as ever-emerging, previously $unknown$ categories appear and must be addressed without retraining. Foundation and vision-language models are pre-trained on large and diverse datasets with the expectation of broad generalization across domains, including medical imaging. However, benchmarking these models on test sets with only a few common outlier types silently collapses the evaluation back to a closed-set problem, masking failures on rare or truly novel conditions encountered in clinical use. We therefore present $NOVA$, a challenging, real-life $evaluation-only$ benchmark of $\sim$900 brain MRI scans that span 281 rare pathologies and heterogeneous acquisition protocols. Each case includes rich clinical narratives and double-blinded expert bounding-box annotations. Together, these enable joint assessment of anomaly localisation, visual captioning, and diagnostic reasoning. Because NOVA is never used for training, it serves as an $extreme$ stress-test of out-of-distribution generalisation: models must bridge a distribution gap both in sample appearance and in semantic space. Baseline results with leading vision-language models (GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B) reveal substantial performance drops across all tasks, establishing NOVA as a rigorous testbed for advancing models that can detect, localize, and reason about truly unknown anomalies.

医学影像异常检测开放世界视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。