首次系统评估医学视觉语言模型的分布外检测能力,提升医疗影像安全可靠性。
Delving into Out-of-Distribution Detection with Medical Vision-Language Models
- 提出分层提示机制,增强模型对异常医学影像的识别能力。
- 在跨模态评测中验证模型在语义与数据分布漂移下的鲁棒性。
- 适用于医疗AI系统安全监控,尤其适合临床部署前的可靠性评估。
近期医学视觉语言模型(VLMs)在图像分类任务中表现出色,主要得益于其强大的零样本泛化能力。然而,由于医学影像数据固有的高度变异性与复杂性,这些模型在分布外(OOD)数据检测方面的潜力尚未被充分探索。本文首次系统研究医学VLMs的OOD检测能力,评估了多种先进方法在涵盖通用与专用领域的多类医学VLM上的表现。为更真实反映实际场景挑战,我们设计了一套跨模态评估流程,全面测试模型在语义偏移与协变量偏移双重压力下的稳健性。此外,我们提出一种新型分层提示方法,显著提升了检测性能。大量实验验证了该方法的有效性,代码已开源。
原文摘要 · Abstract (English)
Recent advances in medical vision-language models (VLMs) demonstrate impressive performance in image classification tasks, driven by their strong zero-shot generalization capabilities. However, given the high variability and complexity inherent in medical imaging data, the ability of these models to detect out-of-distribution (OOD) data in this domain remains underexplored. In this work, we conduct the first systematic investigation into the OOD detection potential of medical VLMs. We evaluate state-of-the-art VLM-based OOD detection methods across a diverse set of medical VLMs, including both general and domain-specific purposes. To accurately reflect real-world challenges, we introduce a cross-modality evaluation pipeline for benchmarking full-spectrum OOD detection, rigorously assessing model robustness against both semantic shifts and covariate shifts. Furthermore, we propose a novel hierarchical prompt-based method that significantly enhances OOD detection performance. Extensive experiments are conducted to validate the effectiveness of our approach. The codes are available at https://github.com/PyJulie/Medical-VLMs-OOD-Detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。