arXiv:2607.15416cs.CV2026-07被引 1

混合异常富集数据反而降低乳腺筛查AI性能,因数据集差异影响模型泛化。

Dataset-Origin Signatures and Shortcut Learning in Screening Mammography AI: A Cross-Dataset Case Study

论文配图:Dataset-Origin Signatures and Shortcut Learning in Screening Mammography AI: A Cross-Dataset Case Study
图 1 · 摘自论文原文
  • 用真实筛查数据训练,避免引入外部异常富集数据
  • 加入外部阳性样本后模型表现下降,AUC最低至0.620
  • 不同数据集特征差异大,适合做领域自适应研究

可靠的乳腺筛查人工智能需要代表低癌症发病率和细微异常的训练数据。我们检验了在真实筛查数据中加入来自异常富集外部数据集的活检确诊病例是否能提升性能。基于纽芬兰与拉布拉多乳腺筛查数据集(NLBSD)联合CBIS-DDSM和CMMD,采用预训练的Mammo-CLIP权重初始化EfficientNet-B5作为冻结线性探测器,在一致预处理和患者级划分下进行评估。仅使用NLBSD的模型达到AUC-ROC 0.737(95%置信区间[0.686, 0.785])。加入外部正例后所有配置性能均下降(AUC-ROC = 0.620–0.644;DeLong检验,霍尔姆校正p < 0.05),且随着新增数据源增多,性能恶化加剧。仅当训练与测试域匹配时,领域匹配评估才带来微弱提升,无一配置超过仅用NLBSD的模型。作为诊断任务,我们将问题重构为预测每张影像的数据集来源,结果显示即使经相同预处理,各数据集仍几乎完美分离,表明数据集特有特征强烈影响学习表征。这些发现表明,简单合并异常富集的乳腺影像数据集会引入超出额外阳性样本收益的领域偏移。采集方式、强度映射及数据构建差异在归一化后依然存在,提示应采用领域感知策略整合异构乳腺影像数据。

原文摘要 · Abstract (English)

Reliable AI for screening mammography requires training data representative of the low cancer prevalence and subtle abnormalities found in screening populations. We examined whether supplementing such data with biopsy-confirmed cases from abnormal-enriched external datasets improves performance. Using the Newfoundland and Labrador Breast Screening Dataset (NLBSD) alongside CBIS-DDSM and CMMD, we evaluated an EfficientNet-B5 encoder initialized with Mammo-CLIP weights as a frozen linear probe under consistent preprocessing and patient-level splits. The NLBSD-only model achieved an AUC-ROC of 0.737 (95% CI [0.686, 0.785]). Adding external positive cases reduced performance in every configuration (AUC-ROC = 0.620--0.644; DeLong test, Holm-corrected $p < 0.05$), with degradation increasing as additional sources were introduced. Domain-matched evaluation produced modest gains only when the training and test domains coincided, and no configuration surpassed the NLBSD-only model. As a diagnostic, we reframed the task as predicting each examination's dataset of origin. The datasets were separated almost perfectly despite identical preprocessing, indicating that dataset-specific characteristics strongly influence the learned representation. These findings show that naïvely pooling abnormal-enriched mammography datasets can introduce domain shift that outweighs the benefit of additional positive cases. Differences in acquisition, intensity mapping, and dataset construction persist after normalization, motivating domain-aware strategies for combining heterogeneous mammography datasets.

医学影像域偏移乳腺筛查数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。