arXiv:2509.05343cs.CV2025-09被引 1

在5种CNN中系统加入注意力机制,提升医学影像诊断准确性和泛化能力。

Systematic Integration of Attention Modules into CNNs for Accurate and Generalizable Medical Image Diagnosis

  • 在VGG16、ResNet18等模型中嵌入注意力模块,自适应聚焦关键区域。
  • EfficientNetB5加混合注意力后,在脑肿瘤和胎盘组织数据集上均表现最优。
  • 增强特征定位能力,适合构建可解释的临床辅助诊断系统。

深度学习已成为医学图像分析的强大工具;然而,传统卷积神经网络(CNN)常难以捕捉对精准诊断至关重要的细粒度与复杂特征。为解决此问题,本文系统地将注意力机制集成至五种广泛应用的CNN架构——VGG16、ResNet18、InceptionV3、DenseNet121和EfficientNetB5,以增强其对显著区域的关注能力并提升判别性能。具体而言,每个基线模型均通过引入挤压-激励模块或混合卷积块注意力模块进行增强,实现通道与空间特征表示的自适应重校准。所提模型在两个不同医学影像数据集上评估:包含多种肿瘤亚型的脑肿瘤MRI数据集,以及含四种组织类别的胎盘组织病理数据集。实验结果表明,注意力增强的CNN在所有指标上均优于基线架构。特别是,搭载混合注意力的EfficientNetB5在两个数据集上均取得最高性能,表现显著提升。除分类准确率提高外,注意力机制还增强了特征定位能力,从而提升跨异构成像模态的泛化性。本工作提供了一个系统化的注意力模块嵌入框架,严格评估其在多种医学影像任务中的影响,为开发鲁棒、可解释且临床可用的深度学习决策支持系统提供了实践指导。

原文摘要 · Abstract (English)

Deep learning has become a powerful tool for medical image analysis; however, conventional Convolutional Neural Networks (CNNs) often fail to capture the fine-grained and complex features critical for accurate diagnosis. To address this limitation, we systematically integrate attention mechanisms into five widely adopted CNN architectures, namely, VGG16, ResNet18, InceptionV3, DenseNet121, and EfficientNetB5, to enhance their ability to focus on salient regions and improve discriminative performance. Specifically, each baseline model is augmented with either a Squeeze and Excitation block or a hybrid Convolutional Block Attention Module, allowing adaptive recalibration of channel and spatial feature representations. The proposed models are evaluated on two distinct medical imaging datasets, a brain tumor MRI dataset comprising multiple tumor subtypes, and a Products of Conception histopathological dataset containing four tissue categories. Experimental results demonstrate that attention augmented CNNs consistently outperform baseline architectures across all metrics. In particular, EfficientNetB5 with hybrid attention achieves the highest overall performance, delivering substantial gains on both datasets. Beyond improved classification accuracy, attention mechanisms enhance feature localization, leading to better generalization across heterogeneous imaging modalities. This work contributes a systematic comparative framework for embedding attention modules in diverse CNN architectures and rigorously assesses their impact across multiple medical imaging tasks. The findings provide practical insights for the development of robust, interpretable, and clinically applicable deep learning based decision support systems.

医学影像注意力机制CNN改进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。