arXiv:2503.15008eess.IVcs.AI2025-03被引 5

融合CNN与Transformer,提升乳腺超声癌变检测精度

A Novel Channel Boosted Residual CNN-Transformer with Regional-Boundary Learning for Breast Cancer Detection

  • 设计新型混合模型RBCMT,结合局部纹理与全局上下文特征
  • 在BUSI数据集上达95.63%准确率,优于现有方法
  • 适合医学影像分析、肿瘤检测研究者参考

基于深度学习的乳腺超声图像(BUSI)肿瘤检测近年取得显著进展。尽管深度卷积神经网络(CNN)和视觉变压器(ViT)各自表现良好,但模型复杂度及对比度、纹理、肿瘤形态变化带来的不确定性仍影响效果。本文提出新型混合框架CB-Res-RBCMT,结合定制残差CNN与新型ViT组件,用于精细分析BUSI图像。RBCMT采用茎干卷积块与CNN-Transformer(CMT)块,并引入区域与边界(RB)特征提取机制,捕捉对比度与形态差异。CMT块通过多头注意力实现全局上下文交互,兼具轻量化与高效性;定制逆残差与茎干CNN有效提取局部纹理并缓解梯度消失。此外,通道增强(CB)策略融合原始RBCMT通道与迁移学习生成的残差特征图,丰富有限数据的特征多样性,经空间注意力模块优化像素选择,减少冗余,增强微小对比与纹理差异的识别能力。在标准谐调严格版BUSI数据集上,该模型达到F1-score 95.57%、准确率95.63%、敏感度96.42%、精确率94.79%,优于现有ViT与CNN方法,验证了其在多特征捕捉与乳腺癌诊断中的优越性能。

原文摘要 · Abstract (English)

Recent advancements in detecting tumors using deep learning on breast ultrasound images (BUSI) have demonstrated significant success. Deep CNNs and vision-transformers (ViTs) have demonstrated individually promising initial performance. However, challenges related to model complexity and contrast, texture, and tumor morphology variations introduce uncertainties that hinder the effectiveness of current methods. This study introduces a novel hybrid framework, CB-Res-RBCMT, combining customized residual CNNs and new ViT components for detailed BUSI cancer analysis. The proposed RBCMT uses stem convolution blocks with CNN Meet Transformer (CMT) blocks, followed by new Regional and boundary (RB) feature extraction operations for capturing contrast and morphological variations. Moreover, the CMT block incorporates global contextual interactions through multi-head attention, enhancing computational efficiency with a lightweight design. Additionally, the customized inverse residual and stem CNNs within the CMT effectively extract local texture information and handle vanishing gradients. Finally, the new channel-boosted (CB) strategy enriches the feature diversity of the limited dataset by combining the original RBCMT channels with transfer learning-based residual CNN-generated maps. These diverse channels are processed through a spatial attention block for optimal pixel selection, reducing redundancy and improving the discrimination of minor contrast and texture variations. The proposed CB-Res-RBCMT achieves an F1-score of 95.57%, accuracy of 95.63%, sensitivity of 96.42%, and precision of 94.79% on the standard harmonized stringent BUSI dataset, outperforming existing ViT and CNN methods. These results demonstrate the versatility of our integrated CNN-Transformer framework in capturing diverse features and delivering superior performance in BUSI cancer diagnosis.

乳腺癌检测CNN-Transformer医学图像分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。