arXiv:2504.19888cs.CV2025-04被引 14

用自监督学习和混合模型提升乳腺癌筛查影像检测准确率

Enhancing breast cancer detection on screening mammogram using self-supervised learning and a hybrid deep model of Swin Transformer and Convolutional Neural Network

  • 用自监督预训练+Transformer与卷积网络融合提升特征提取
  • 在两个数据集上分别达到0.864和0.889的AUC值
  • 适合缺乏标注数据的医学影像AI研究者参考

目前高质量标注医疗训练数据稀缺,限制了人工智能在乳腺癌诊断中的应用。深度模型进行乳腺钼靶分析与病灶检测需大量标注图像,而获取成本高、耗时长。为此,我们提出一种新方法,结合自监督学习(SSL)与混合深度模型HybMNet,该模型融合局部自注意力与细粒度特征提取能力,以增强乳腺癌检测性能。方法采用两阶段训练:(1) 自监督预训练:使用EsViT技术,基于少量钼靶图像对Swin Transformer(Swin-T)进行预训练;(2) 下游训练:提出的HybMNet将预训练的Swin-T作为主干,与基于CNN的网络及新型融合策略结合。Swin-T利用局部自注意力识别高分辨率图像中关键区域,而CNN从选定区域提取细粒度局部特征。融合模块整合两者全局与局部信息,生成鲁棒预测结果。整个HybMNet端到端训练,损失函数综合两模块输出以优化特征提取与分类表现。结果显示,该方法在区分良性(正常)与恶性钼靶图像方面表现优异,在CMMD数据集上AUC达0.864(95% CI: 0.852, 0.875),在INbreast数据集上达0.889(95% CI: 0.875, 0.903),验证了其有效性。

原文摘要 · Abstract (English)

Purpose: The scarcity of high-quality curated labeled medical training data remains one of the major limitations in applying artificial intelligence (AI) systems to breast cancer diagnosis. Deep models for mammogram analysis and mass (or micro-calcification) detection require training with a large volume of labeled images, which are often expensive and time-consuming to collect. To reduce this challenge, we proposed a novel method that leverages self-supervised learning (SSL) and a deep hybrid model, named \textbf{HybMNet}, which combines local self-attention and fine-grained feature extraction to enhance breast cancer detection on screening mammograms. Approach: Our method employs a two-stage learning process: (1) SSL Pretraining: We utilize EsViT, a SSL technique, to pretrain a Swin Transformer (Swin-T) using a limited set of mammograms. The pretrained Swin-T then serves as the backbone for the downstream task. (2) Downstream Training: The proposed HybMNet combines the Swin-T backbone with a CNN-based network and a novel fusion strategy. The Swin-T employs local self-attention to identify informative patch regions from the high-resolution mammogram, while the CNN-based network extracts fine-grained local features from the selected patches. A fusion module then integrates global and local information from both networks to generate robust predictions. The HybMNet is trained end-to-end, with the loss function combining the outputs of the Swin-T and CNN modules to optimize feature extraction and classification performance. Results: The proposed method was evaluated for its ability to detect breast cancer by distinguishing between benign (normal) and malignant mammograms. Leveraging SSL pretraining and the HybMNet model, it achieved AUC of 0.864 (95% CI: 0.852, 0.875) on the CMMD dataset and 0.889 (95% CI: 0.875, 0.903) on the INbreast dataset, highlighting its effectiveness.

乳腺癌检测自监督学习混合模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。