用反事实生成提升医学影像模型对设备差异的鲁棒性
Robust image representations with counterfactual contrastive learning
- 通过因果图像合成生成更真实的域差异正样本
- 在5个数据集上显著提升对设备差异的泛化性能
- 尤其改善低频设备图像表现,减少性别差异偏差
对比学习能显著提升模型泛化能力与下游性能,但其表征质量高度依赖数据增强策略。正样本对应保留语义信息,同时去除与数据采集域相关的无关变化。传统对比学习采用预定义通用图像变换模拟域偏移,但难以真实反映医学影像中的设备差异(如扫描仪不同)。为此,本文提出反事实对比学习框架,利用因果图像合成技术生成能准确捕捉相关域变化的正样本对。在涵盖胸部X光和乳腺钼靶的五个数据集上,针对SimCLR与DINO-v2两种对比目标进行评估,该方法在应对采集域偏移时表现出更强鲁棒性。尤其在训练集中代表性不足的扫描仪图像上,下游性能显著优于标准对比学习。进一步实验表明,该框架不仅适用于采集域偏移,还能降低按生物性别分组的性能差距。
原文摘要 · Abstract (English)
Contrastive pretraining can substantially increase model generalisation and downstream performance. However, the quality of the learned representations is highly dependent on the data augmentation strategy applied to generate positive pairs. Positive contrastive pairs should preserve semantic meaning while discarding unwanted variations related to the data acquisition domain. Traditional contrastive pipelines attempt to simulate domain shifts through pre-defined generic image transformations. However, these do not always mimic realistic and relevant domain variations for medical imaging, such as scanner differences. To tackle this issue, we herein introduce counterfactual contrastive learning, a novel framework leveraging recent advances in causal image synthesis to create contrastive positive pairs that faithfully capture relevant domain variations. Our method, evaluated across five datasets encompassing both chest radiography and mammography data, for two established contrastive objectives (SimCLR and DINO-v2), outperforms standard contrastive learning in terms of robustness to acquisition shift. Notably, counterfactual contrastive learning achieves superior downstream performance on both in-distribution and external datasets, especially for images acquired with scanners under-represented in the training set. Further experiments show that the proposed framework extends beyond acquisition shifts, with models trained with counterfactual contrastive learning reducing subgroup disparities across biological sex.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。