多中心乳腺钼靶数据中,简单模型反而比自洽学习更有效。
Revisiting Invariant Learning for Out-of-Domain Generalization on Multi-Site Mammogram Datasets
- 用多国数据训练,不依赖复杂不变性算法
- 标准ERM在未见群体上准确率更高
- 适合关注医疗AI公平性的研究者
实现人工智能健康公平需要诊断模型在不同人群间保持可靠性。然而,乳腺癌筛查系统常因域过拟合导致在不同人口群体中性能下降。尽管不变性学习算法旨在消除机构特异性关联,其在医学影像中的有效性仍待验证。本研究全面评估了乳腺钼靶图像的域泛化技术。构建了涵盖美国(CBIS-DDSM、EMBED)、葡萄牙(INbreast、BCDR)和塞浦路斯(BMCD)的数据集训练环境,并在埃及(CDD-CESM)和瑞典(CSAW-CC)的未见队列上评估全局泛化能力。对比了不变风险最小化(IRM)与方差风险外推(VREx)与经过严格优化的经验风险最小化(ERM)基线。结果表明,标准ERM在域外测试中始终优于专门设计的不变性机制。尽管VREx在稳定注意力图方面展现潜力,但不变性目标不稳定且易欠拟合。结论是,当前构建公平AI的最佳方式是最大化跨国数据多样性,而非依赖复杂的算法不变性。
原文摘要 · Abstract (English)
Achieving health equity in Artificial Intelligence (AI) requires diagnostic models that maintain reliability across diverse populations. However, breast cancer screening systems frequently suffer from domain overfitting, degrading significantly when deployed to varying demographics. While Invariant Learning algorithms aim to mitigate this by suppressing site-specific correlations, their efficacy in medical imaging remains underexplored. This study comprehensively evaluates domain generalization techniques for mammography. We constructed a multi-source training environment aggregating datasets from the United States (CBIS-DDSM, EMBED), Portugal (INbreast, BCDR), and Cyprus (BMCD). To assess global generalizability, we evaluated performance on unseen cohorts from Egypt (CDD-CESM) and Sweden (CSAW-CC). We benchmarked Invariant Risk Minimization (IRM) and Variance Risk Extrapolation (VREx) against a rigorously optimized Empirical Risk Minimization (ERM) baseline. Contrary to expectations, standard ERM consistently outperformed specialized invariant mechanisms on out-of-domain testing. While VREx showed potential in stabilizing attention maps, invariant objectives proved unstable and prone to underfitting. We conclude that engineering equitable AI is currently best served by maximizing multi-national data diversity rather than relying on complex algorithmic invariance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。