测试医学视觉语言模型在噪声下的表现,发现其泛化能力大打折扣。
On the Robustness of Medical Vision-Language Models: Are they Truly Generalizable?
- 构建新基准MediMeta-C,系统评估模型在多种医学图像畸变下的表现
- 5个主流模型在噪声下性能显著下降,跨模态泛化能力受限
- 提出RobustMedCLIP,通过少量样本微调提升鲁棒性,适合临床应用
医学视觉语言模型(MVLMs)在医学图像分析中表现出卓越的泛化能力,但其在噪声和失真条件下的表现仍缺乏充分检验。临床影像固有易受采集伪影和噪声影响,而现有评估多基于洁净数据集,忽视了模型在真实场景中的鲁棒性——即面对现实畸变时的稳定表现。为此,我们首先提出MediMeta-C,一个系统性地对多个医学影像数据集施加多种扰动的畸变基准。结合已有的MedMNIST-C,构建了全面的MVLM鲁棒性评估框架。我们进一步提出RobustMedCLIP,一种针对预训练MVLM的视觉编码器适应方法,通过少量样本微调增强对各类畸变的抗性。在5种医学影像模态上对5个主流MVLM进行广泛实验,结果表明现有模型在畸变下性能严重退化,且在跨域-跨模态间存在权衡。研究揭示了多样化训练与鲁棒适应策略的必要性,证明高效低秩适配结合少量样本微调,可在保持跨模态泛化能力的同时显著提升鲁棒性。
原文摘要 · Abstract (English)
Medical Vision-Language Models (MVLMs) have achieved par excellence generalization in medical image analysis, yet their performance under noisy, corrupted conditions remains largely untested. Clinical imaging is inherently susceptible to acquisition artifacts and noise; however, existing evaluations predominantly assess generally clean datasets, overlooking robustness -- i.e., the model's ability to perform under real-world distortions. To address this gap, we first introduce MediMeta-C, a corruption benchmark that systematically applies several perturbations across multiple medical imaging datasets. Combined with MedMNIST-C, this establishes a comprehensive robustness evaluation framework for MVLMs. We further propose RobustMedCLIP, a visual encoder adaptation of a pretrained MVLM that incorporates few-shot tuning to enhance resilience against corruptions. Through extensive experiments, we benchmark 5 major MVLMs across 5 medical imaging modalities, revealing that existing models exhibit severe degradation under corruption and struggle with domain-modality tradeoffs. Our findings highlight the necessity of diverse training and robust adaptation strategies, demonstrating that efficient low-rank adaptation when paired with few-shot tuning, improves robustness while preserving generalization across modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。