评测15个大模型在乳腺影像域偏移下的鲁棒性,发现专用模型表现最好但非唯一因素。
Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift

- 统一冻结主干线性探测协议,跨12个外部数据集评估模型性能
- 专用视觉-语言模型在跨域任务中平均表现最优,但训练数据不决定一切
- 强调需按数据集级别评估泛化能力,避免单一指标误导
基础模型在乳腺影像中日益用作图像特征提取器,但其在外部域偏移下的鲁棒性尚不明确。我们采用统一的冻结主干线性探测协议,在3个源数据集上训练,并在12个任务兼容的域外(OOD)数据集上评估15个基础模型主干,所有数据经过标签标准化处理。乳腺影像专用的视觉-语言模型(Mammo-FM 和 MaMA)在跨域平均表现最佳,但鲁棒性无法仅由乳腺影像暴露程度解释。DINOv3 作为纯视觉基线仍具竞争力,且乳腺影像适配预训练并未持续提升泛化能力。数据集级分析显示,即使顶尖模型在不同数据集上表现也存在异质性。特征空间分析表明,有效表示可保留临床信号同时保持数据集和采集结构特征。这些发现强调了以数据集级别进行域外评估是衡量乳腺影像表示能力的核心标准。代码已公开:https://github.com/biomedia-mira/mammo-ood。
原文摘要 · Abstract (English)
Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and cancer status using a unified frozen-backbone linear-probe protocol, training on 3 source datasets and evaluating on 12 task-compatible out-of-distribution (OOD) datasets after label harmonization. Mammography-specific vision-language models (Mammo-FM and MaMA) provide the strongest mean OOD performance, but robustness is not explained by mammography exposure alone. DINOv3 remains a competitive vision-only baseline, and mammography-adapted pretraining does not consistently improve generalization. Dataset-level analysis further shows that even leading models show heterogeneous performance across datasets. Feature-space inspection reveals that useful representations can preserve clinical signal while retaining dataset and acquisition structure. These findings highlight dataset-level OOD evaluation as a central criterion for assessing mammography representations. Our code is publicly available: https://github.com/biomedia-mira/mammo-ood.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。