通过自监督学习减少跨医院胸片的站点泄露,提升模型泛化能力。
Feature-level Site Leakage Reduction for Cross-Hospital Chest X-ray Transfer via Self-Supervised Learning

- 用多站点自监督学习和特征级对抗混淆减少图像中的医院来源信息。
- 多站点自监督使测试集AUC从0.67提升至0.78,显著改善跨院迁移效果。
- 发现传统方法假设的不变性不成立,测量泄露能改变对迁移策略的判断。
跨医院胸片模型失败常归因于领域偏移,但多数研究假设不变性却未验证。本文直接测量站点泄露,并探讨其如何改变对迁移方法的结论。研究采用多站点自监督学习(SSL)与特征级对抗站点混淆,先在NIH和CheXpert上无病灶标签预训练ResNet-18,再冻结编码器,在NIH上训练线性肺炎分类器并评估在RSNA上的迁移表现。通过后验线性探针量化来自冻结主干特征 $f$ 与投影特征 $z$ 的采集站点预测精度。三次随机种子下,多站点SSL将RSNA AUC从ImageNet初始化的0.6736 ± 0.0148提升至0.7804 ± 0.0197。在 $f$ 上引入对抗混淆后,站点探测准确率从0.9890 ± 0.0021降至0.8504 ± 0.0051(随机猜测为0.50),在 $z$ 上从0.8912 ± 0.0092降至0.7810 ± 0.0250。结果表明,测量泄露可重塑对迁移方法的解读:多站点自监督驱动迁移,而对抗混淆揭示了不变性假设的局限性。
原文摘要 · Abstract (English)
Cross-hospital failure in chest X-ray models is often attributed to domain shift, yet most work assumes invariance without measuring it. This paper studies how to measure site leakage directly and how that measurement changes conclusions about transfer methods. We study multi-site self-supervised learning (SSL) and feature-level adversarial site confusion for cross-hospital transfer. We pretrain a ResNet-18 on NIH and CheXpert without pathology labels. We then freeze the encoder and train a linear pneumonia classifier on NIH only, evaluating transfer to RSNA. We quantify site leakage using a post hoc linear probe that predicts acquisition site from frozen backbone features $f$ and projection features $z$. Across 3 random seeds, multi-site SSL improves RSNA AUC from 0.6736 $\pm$ 0.0148 (ImageNet initialization) to 0.7804 $\pm$ 0.0197. Adding adversarial site confusion on $f$ reduces measured leakage but does not reliably improve AUC and increases variance. On $f$, site probe accuracy drops from 0.9890 $\pm$ 0.0021 (SSL-only) to 0.8504 $\pm$ 0.0051 (CanonicalF), where chance is 0.50. On $z$, probe accuracy drops from 0.8912 $\pm$ 0.0092 to 0.7810 $\pm$ 0.0250. These results show that measuring leakage changes how transfer methods should be interpreted: multi-site SSL drives transfer, while adversarial confusion exposes the limits of invariance assumptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。